Compare commits

..
Author SHA1 Message Date
root b9b1712ac6 fix: make host escalations state-change driven
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER,
AMBER->RED) and ONCE when it drops back down (recovery notice). While it stays
in the same band, it is reported in the scan output only — no DM, no channel
alert. This stops the same 96% easystore2 from re-DMing the owner on every 6h
scan.

State lives in a small JSON state file (state/host-disk-bands.json), keyed by
host/volume -> last-seen band. The scanner reads the prior band, compares to the
current band, and DMs only on a transition; the state file is written after
every scan. Chosen over a periodic digest because the scan already runs every
6h and a transition is genuinely new, actionable state.

First-run behavior: when the state file does not yet exist, the current band of
every volume is recorded as baseline WITHOUT alerting — a first run would
otherwise DM every already-elevated volume at once.

Report-only restriction and volume-naming output kept exactly as-is.
2026-09-15 14:25:49 +00:00
root 6fb411613e fix: add host filesystem thresholds to disk-gc contract
Add separate threat bands for HOST filesystems (distinct from guest bands):
- HOST-WARN at 85%: name volume + % + absolute free space in scan output
- HOST-AMBER at 90%: flag for owner attention, Zulip DM
- HOST-RED at 95%: flag for immediate owner attention, Zulip DM + channel alert

Volume naming rule: every host line MUST name the volume and what lives on it.
Action classes by volume type:
- host-root: near full = real risk (backup staging, thin-pool metadata)
- media (/media/*): near full = capacity decision for owner, never auto-delete
- pbs-datastore (tank): near full = breaks Proxmox Backup Server

Report-only restriction: no automatic deletion of media or datastore content ever.

Justification (measured 2026-09-15): storepve /media/easystore2 at 96% was
reported but never banded or acted on. Two incidents this weekend showed the
host filesystem is the thing that breaks, not the guest's.

Added HOST-WARN/AMBER/RED alert templates.
Added report-only execution rule for host filesystems.
2026-09-15 14:15:53 +00:00
abiba-bot 4bdd88613b Merge pull request 'docs: pm2/spoton AS-BUILT correction and the real Prometheus node coverage' (#104) from fix/contract-corrections-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-15 13:14:39 +00:00
root bc7a55122f fix: pm2 contract corrections
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
pm2-self-heal.prose.md:
- Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2
- gpu-watchdog is decommissioned and folded into gpu-monitor.service
- gitea-runner is KEPT; abiba-zulip is KEPT (online for days)
- spoton-service was deleted; live PM2 set is 4 processes
- Preserve historical context for crash-loop guard

litellm-health.prose.md:
- Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100)
- Note .4:9100 is DEAD target (no route, down for weeks)
- Clarify this does not read as 6 healthy nodes
2026-09-15 13:05:32 +00:00
abiba-bot 713b9ce80c Merge pull request 'Remove client deliverable from the contracts repo (process fix)' (#77) from cleanup/remove-scot-deliverable-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-15 12:57:59 +00:00
abiba-bot b451d6f81a Merge branch 'master' into cleanup/remove-scot-deliverable-20260912
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-15 12:53:37 +00:00
abiba-bot fe6eb35291 Merge pull request 'ci: make PR validation unconditional - the paths filter silently skipped whole classes of pull requests' (#103) from fm/ci-paths-filter-skips-deliverables-prs-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-15 12:53:20 +00:00
abiba-bot 40cf057370 Merge pull request 'fix(backups): document the thin-pool headroom and staging-directory preconditions, with the incidents that justify them' (#97) from fix/backup-safety-preconditions-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-15 12:19:19 +00:00
abiba-bot 5053a33e2c ci: make PR validation trigger unconditional (remove paths filter)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 16s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The pr-pipeline workflow filtered both push and pull_request on
paths (**.prose.md, scripts/**.sh, **.yaml, **.yml). A PR whose diff
touched none of those paths — e.g. PR #77, deliverables/-only —
produced no Gitea Actions run at all, so validate/lint/ai-review and
the merge gate were silently skipped.

Remove the paths filter from both triggers and record in the file
that the trigger is intentionally unfiltered. No job, step, needs,
if or command is changed.
2026-09-15 12:14:40 +00:00
root b4b5321011 fix: add hostname resolution warning and fix dmsetup field documentation
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Add warning that bare PVE hostnames (acerpve, amdpve, etc.) resolve to VPS
via *.dns.sysloggh.net wildcard, not to actual nodes. List IP addresses:
- acerpve 192.168.68.9
- amdpve 192.168.68.15
- storepve 192.168.68.6
- minipve 192.168.68.12
- ocupve 192.168.68.5

Update acerpve example in backup preflight to include address (192.168.68.9).

Fix dmsetup comment to show full field order:
=start =length =thin-pool =transaction-id
=metadata_used/metadata_total =data_used/data_total
remaining fields are flags

Make it clear lvs command is the primary source for percentages, dmsetup is only for error-state check.

Incidents now include addresses: acerpve (192.168.68.9) and amdpve (192.168.68.15).
2026-09-15 12:09:59 +00:00
root 69940bc9eb fix: correct backup preflight commands - lvs pve/data and dmsetup field documentation
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
Fix two errors in the PREFLIGHT section (measured on acerpve 2026-09-15):
1. lvs -o ... pve/data (not pve-data-tpool) - this is the PRIMARY check that yields percentages directly
   - Quote the acerpve example: data 29.95% 1.22% <816.21g
2. dmsetup status pve-data-tpool - document fields correctly:
   -  = transaction ID (99), NOT data_percent
   -  = metadata used/total blocks
   -  = data used/total sectors
   - Show how to derive percentages if needed
3. Keep the error-state check (grep -q 'Error|Fail') - this is how the incident presented

Everything else stays: 1777 tmpdir requirement with EACCES symptom, incidents as rationale,
GPU-host fact, --output-format json rule, honest note that metadata/snapshot pressure is unproven.
2026-09-15 11:42:10 +00:00
root 65eaffe1c6 fix: add backup safety preconditions - thin-pool headroom and tmpdir 1777
Add documented preflight checks for VM/CT backups on LVM thin-pool hosts:
- dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70%
- Exit 1 if pool shows Error/Fail state (takes down entire VG including host root)
- tmpdir must be mode 1777 (world-traversable) for vzdump archive step
- --output-format json for tasks started from truncating shells
- Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14)
- Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control
- Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup
2026-09-15 11:37:41 +00:00
abiba-bot b386bd0c19 Merge pull request 'fix(keys): correct the key-lifecycle contracts to the measured truth' (#95) from fix/key-expiry-enforcement-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-15 05:37:45 +00:00
root 8a4dd08b05 fix: capture 401/403 response body and key alias for credential faults
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- get_response_body() returns first 200 chars of response body (single line)
- On 401/403 model probe: report code + body + key_alias
- Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116)
- Failed connections stay probe-failed, 200 stays plain 200
- Do not turn other statuses into credential faults

Signed-off-by: Abiba
2026-09-15 05:11:07 +00:00
root 88b6decb31 fix: remove false Infisical claim - master key NOT in infrastructure project
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Replace Infisical retrieval path with proven docker exec + .env note
- State explicitly that master key is NOT in Infisical project=infrastructure
- Keep the live-key check and never-trust-a-literal instruction
- All other corrections from PR #95 preserved

Signed-off-by: Abiba
2026-09-15 04:42:31 +00:00
root 7bf9f78fc6 fix: gpu-dense probe timeout handling - report probe-failed with kind, not service verdict
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- probe_http now returns (code, failure_kind) tuple
- Model probes report 'probe-failed: <model> <kind> (Ns timeout)' on 000
- Do not assert a service verdict from a failed probe
- 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill)
- 60s timeout for syslog-auto pool alias with retry on 000

Signed-off-by: Abiba
2026-09-15 04:22:46 +00:00
root dd9e68329e fix: correct infisical --plain flag and use verified localhost:4000 endpoint
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Replace broken --plain flag (prints nothing on CLI 0.43.110) with awk parsing
- Note that --plain is broken so nobody fixes it back
- Replace unproven nginx path with verified direct endpoint http://127.0.0.1:4000/key/list

Signed-off-by: Abiba
2026-09-15 03:05:30 +00:00
root cd9ec6a0df fix: correct key-lifecycle contracts to match measured reality
- hermes-key-enforcement.prose.md:
  - State that expiry must be set EXPLICITLY at creation with duration
  - Record that config default is NOT honoured by LiteLLM 1.99.1
  - Describe daily audit as AUDIT-ONLY (reports non-expiring and soon-to-expire)
  - State that renewal is NOT implemented
  - Document exclusions: abiba-pi and all crewmate keys stay WITHOUT expiry
  - koby is report-only

- litellm-api-keys.prose.md:
  - Replace literal master key with retrieval path (docker exec + infisical)
  - State that literal values must never be trusted again (key rotates)
  - Add live-key check (200 from /key/list)

Signed-off-by: Abiba
2026-09-15 02:59:19 +00:00
abiba-bot 1d20fbaa7f Merge pull request 'fix(agent-health): gateway liveness reports a probe failure as a probe failure, never as an agent outage' (#93) from fix/agent-health-gateway-leg-deterministic-20260914 into master 2026-09-14 16:45:25 +00:00
root 2f961d7e7a fix: agent-health-check gateway leg — deterministic probe-failed reporting
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output

Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind

Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
2026-09-14 15:21:36 +00:00
abiba-bot 4fe4f3621d Merge pull request 'fix(monitoring): probe precision - any HTTP status means alive, failed probes never become service verdicts' (#92) from fix/probe-precision-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 15:07:35 +00:00
root 86d2987ad8 fix: probe precision — add retry + probe-failed reporting to zulip-health
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Per defect report 1150.msg:
- Add standing probe rules section (2026-09-14)
- Step 1 (Zulip API): retry once at 25s on 000, print target + code
- Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host
- Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed

Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
2026-09-14 14:53:57 +00:00
root efe9381283 fix: probe precision — correct GPU exporter path, Grafana port, add retry + probe-failed reporting
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed

Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
2026-09-14 14:44:37 +00:00
abiba-bot 6a1c4967db Merge pull request 'fix(disk-gc): deterministic reachability verdict and correctly labelled disk figures' (#91) from fix/disk-gc-probe-deterministic-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 14:43:43 +00:00
root 5ed6f8179c fix: add os import for HELPER_PCT_RUN.readable() check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-14 14:14:15 +00:00
root b9322973ce fix disk-gc-scan: CWD independence (FIX A) and df column parsing (FIX B) 2026-09-14 14:05:06 +00:00
root fb1916707d document deterministic disk-gc-scan.py in contract 2026-09-14 13:20:26 +00:00
mumuni-bot 0c7d7be5ad ci: re-trigger pipeline (no statuses recorded on 9312c1a8 head) 2026-09-14 13:19:33 +00:00
root d28df4f4de add disk-gc-scan.py: deterministic fleet disk probe with per-guest access methods 2026-09-14 13:14:10 +00:00
abiba-bot cb26ee06d6 Merge pull request 'fix(hermes): separate POLICY observations from FAULT findings in the audit contracts' (#90) from fix/hermes-violation-classification-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:50:15 +00:00
root 9a789ab76d fix(hermes): separate policy observations from fault findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Contract design defect: 'uses a non-harness provider' (POLICY) and
'cannot authenticate' (FAULT) were printed as the same violation class.
A policy observation must never be phrased as if the agent were broken.

Changes:
1. Added 'Violation Classification' section to all three contracts
2. Separated POLICY (observation only) from FAULT (requires request-level evidence)
3. Rules:
   - Do NOT infer runtime credential resolution from config text alone
   - Require request-level evidence before calling a FAULT: observed auth failure
     or absence of successful calls
   - If calls are succeeding, output is 'POLICY: uses <provider> directly; calls
     succeeding' - not a violation
   - State what you OBSERVED, not what the field implies

Files changed (3):
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
2026-09-14 12:32:06 +00:00
abiba-bot 71ceda0042 Merge pull request 'fix(hermes): reachability verdict must not come from the remote command's exit code' (#89) from fix/hermes-reachability-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:20:09 +00:00
root 4e34b7a2a2 fix(hermes): wire all three contracts to reachability helper
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
   - hermes-key-enforcement.prose.md
   - hermes-config-template.prose.md
   - hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code

Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
2026-09-14 12:01:32 +00:00
root f9f6661dd5 fix(hermes): add shared reachability check helper
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.

Fix: Add scripts/hermes-reachability-check.sh with the pattern:
  out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
  if [ $? -ne 0 ]; then verdict="unreachable"
  elif [ -n "$out" ]; then verdict="violation: $out"
  else verdict="compliant"
  fi

This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant

Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).

Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
2026-09-14 11:55:41 +00:00
abiba-bot cb36ff1ea5 Merge pull request 'fix(monitoring): litellm-health script - pool-alias timeout and leaked key-list debug output' (#88) from fix/litellm-health-timeout-and-debug-20260914 into master 2026-09-14 03:34:58 +00:00
root 05366bd58d Fix litellm-health-check.py robustness defects
1. Timeout fix for pool alias (syslog-auto):
   - Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
   - Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
   - Cold first request to pool alias can take ~13s; 10s was too short

2. Remove DEBUG prints from output:
   - Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
   - These leaked key inventory to status logs
   - Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
2026-09-14 03:24:30 +00:00
abiba-bot e648b5ac0e Merge pull request 'feat(monitoring): add scripts/litellm-health-check.py so the litellm-health contract is executed, not improvised' (#87) from fix/litellm-health-executor-script-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-14 03:20:36 +00:00
root 38d7e8b064 Update litellm-health contract to mandate executor script
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
2026-09-14 03:15:42 +00:00
root e94fadfa60 Add litellm-health-check.py: standardized health check script for LiteLLM fleet monitoring
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics

Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).

All 11 checks passing consistently.
2026-09-14 03:15:27 +00:00
abiba-bot ba76f2c7d3 Merge pull request 'fix(contracts): abiba default model row + disk-gc access-method note' (#86) from fix-litellm-health-keylist-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-13 19:09:19 +00:00
root 7906b2d52d fix: change abiba default model from deepseek-v4-pro to syslog-auto
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 13m52s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The abiba-zulip-restore.prose.md file incorrectly claimed the default model
was 'deepseek-v4-pro', but /root/.pi/agent/settings.json declares
'defaultModel: syslog-auto' and 'defaultProvider: syslog-harness'.

deepseek-v4-pro is not a servable LiteLLM model name for this fleet.
The only live models are: syslog-auto, gpu-dense, gpu-vision, strix-moe.

Note: The firstmate/pi session talks to DeepSeek through its own
provider (auth.json), NOT through LiteLLM, so no LiteLLM change is
needed to keep firstmate's deepseek usage working.
2026-09-13 18:55:24 +00:00
root c81cf5b6f0 fix: disk-gc-threat-response - correct access methods for kagentz and docker-vm
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- kagentz (105): use ssh root@kagentz (hostname), NOT pct exec 105 (shows loop0 59G, not real 99G)
- docker-vm (109): use ssh root@192.168.68.7 (correct QEMU VM access)
- abiba (100): use ssh root@abiba (hostname)
- syslog-api (116): use pct-run 116 (correct)

Verified:
  kagentz: 99G 8.9G 86G 10% /
  docker-vm: 158G 17G 135G 11% /
  abiba: 59G 13G 44G 23% /
  syslog-api: 40G 13G 25G 34% /
2026-09-13 04:30:28 +00:00
abiba-bot ef168d9690 Merge pull request 'fix(litellm-health): run the key-list admin call on the gateway host, not inside the container' (#84) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-13 03:20:04 +00:00
root edcf465831 fix: litellm-health step 8 - run key/list on CT 116 host, not in container
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- harness-litellm container has no curl/wget, so docker exec harness-litellm curl returns empty
- Fix: run curl on CT 116 host (ssh root@192.168.68.116 then curl)
- Add admin-call-failed label for empty/unparseable responses
- Verify: 10 keys found (host-side curl), NO-CURL confirmed in container
2026-09-13 03:10:09 +00:00
abiba-bot 89651cf37c Merge pull request 'fix(monitoring): make credential sourcing explicit and fail loudly when missing' (#83) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-12 23:22:43 +00:00
root 6f40a3be60 fix: address credential-sourcing review findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env
  (not /etc/zulip-bot.env which doesn't exist); add credential-missing check;
  remove unused monitor_key variable

- litellm-health: Clarify that syslog-auto is a fallback pool alias, not a
  step 7 probe; monitor key must be scoped for all four aliases
2026-09-12 23:17:27 +00:00
root d9368467ff fix: credential-sourcing fix for litellm-health and infrastructure-monitoring
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116)
  with executable commands; add syslog-auto alias; document credential-missing failure
  condition (not bare 401 or 0 keys)

- infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116
  via ssh instead of sourcing local env file that doesn't exist on executor host

- Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe,
  syslog-auto with corrected credential retrieval
2026-09-12 23:13:24 +00:00
abiba-bot e38598eea4 Merge pull request 'fix: restore per-host probe coverage + sweep residual retired names' (#82) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 22:35:55 +00:00
root 876b011359 no-mistakes(document): Align Rule 9 max_context_window rationale with pool floor
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-12 22:21:57 +00:00
root d23cce89e1 no-mistakes(document): Fix stale retired-name and GPU-context contradictions in contracts 2026-09-12 22:20:02 +00:00
root 9ada2b23c7 no-mistakes(review): Align Rule 8 max_context_window rationale with pool floor 2026-09-12 22:11:40 +00:00
root bd0065bb31 no-mistakes(review): Fix monitor-key scope, Strix context, retired delegation names 2026-09-12 22:02:58 +00:00
root f99f7e1e34 fix: restore per-host probe coverage + sweep residual retired names
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
A. litellm-health step 7: restore .8 → gpu-dense probe (was duplicated
   to gpu-vision after 8210fd9). Add note documenting monitor key scope
   gap for gpu-dense.
B. Sweep remaining retired names presented as usable:
   - README.md:91: qwen3.6-27B-code → gpu-dense in runnable example
   - gpu-fleet.prose.md:14: qwen3.6-35B-udq4 → Carnice-Qwen3.6-MoE...
   - infrastructure-control.prose.md:226: qwen3.6-35B-udq4 → strix-moe
   - proxmox-monitor.prose.md:90: qwen3.6-35B-udq4 → strix-moe
C. Audit test: 10/10 passed (retired raw names now hard-fail)

Fix-forward from 8210fd9 (direct master push).
2026-09-12 21:56:24 +00:00
root 8210fd905c fix: align contracts to 4-name LiteLLM registry (2026-09-12)
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
2026-09-12 21:39:54 +00:00
abiba-bot b3644f0292 Merge pull request 'fix(disk-gc): hard guest-level report-only gate for CT 111/.129 + correct stale fleet map' (#81) from fix/disk-gc-report-only-129 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 19:13:45 +00:00
abiba 672bf8a912 docs(disk-gc): attach real fleet-run verification for the CT 111 report-only gate
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
Intent requires a production run log, not only the unit-level dry run. This is a real
scan of the live 25-entry fleet through scripts/disk-gc-plan.py: CT 111 (tdunna) at 84%
AMBER is alerted as report-only and no gc-executor row is emitted for it; the owned hosts
acerpve .9 (77%) and amdpve .15 (76%) still receive gc-executor. No GC command was run
against 192.168.68.129.
2026-09-12 19:11:36 +00:00
root 6e612ce37b no-mistakes(document): Fix stale CT 111 node maps and CI range 2026-09-12 19:11:09 +00:00
root 357f808a82 no-mistakes(review): Make report-only skip structural with if/else loop 2026-09-12 19:00:28 +00:00
root 6272d29978 no-mistakes(review): Fail closed on empty report-only exclusion list 2026-09-12 18:55:34 +00:00
root 4bd6132cf5 no-mistakes(review): Reject report-only exclusion entries lacking identity keys 2026-09-12 18:51:42 +00:00
root 6d65cba064 no-mistakes(review): Canonicalize guest identities in GC report-only gate 2026-09-12 18:47:17 +00:00
root de1428b4ae no-mistakes(review): Harden GC gate key aliases, fail closed, fix baseline 2026-09-12 18:42:36 +00:00
root 4ea2d0309f no-mistakes(review): Fix report-only gate, dedupe list, correct fleet map 2026-09-12 18:38:19 +00:00
abiba 7b8cc5f9ac fix(disk-gc): hard guest-level report-only gate for CT 111/.129; correct stale fleet map
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.

- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
  hostname and IP all match); an excluded guest is alerted and skipped, so no
  gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
  authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
  (was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
  the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
  why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
  pct-run now that the map is correct; documented that the scanner must probe it
  like any other guest, never via a local-only path (the scanner runs inside CT 100).
2026-09-12 18:33:02 +00:00
abiba-bot d9e06863d8 Merge pull request 'fix(audit): stop requiring the retired gpu-light alias; derive model fields; sweep retired names' (#80) from fix/retired-alias-sweep-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 17:15:05 +00:00
abiba a2edc2f56f ci(pr-pipeline): fix flaky frontmatter check (grep -q SIGPIPE under pipefail)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The validate job runs with bash -e -o pipefail. `echo "$FM" | grep -q '^name:'`
lets grep exit on first match, which can SIGPIPE the echo; pipefail then reports
the pipeline non-zero and the || branch raises a false "Missing name/description".
The flagged file set varied run to run (and included files untouched by the PR)
while a fresh clone of the same commit passes the identical check. Reproduced:
the old form failed 3 of 5 local runs under the same shell flags, the herestring
form passed 5 of 5. Use herestrings so no pipe can be broken.
2026-09-12 17:13:54 +00:00
root 78b501798f no-mistakes(document): Sweep residual gemma labels; align compression rule contradiction
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-09-12 16:54:41 +00:00
root 1f1b47f59d no-mistakes(document): Sweep residual gemma aliases; align compression rule contradiction 2026-09-12 16:51:21 +00:00
root 3d55764799 no-mistakes(review): Split retired-alias audit into fail vs warn; fix key claim 2026-09-12 16:41:35 +00:00
root ce48070f21 no-mistakes(review): Derive model fields recursively; fix historical latency and key claims 2026-09-12 16:33:58 +00:00
root 9e581ab203 no-mistakes(review): Complete retired-alias field coverage; fix misleading example labels 2026-09-12 16:26:14 +00:00
root 9cac3589cf no-mistakes(review): Fix compression example alias; make retired aliases fail audit 2026-09-12 16:19:43 +00:00
abiba 221f9f79f3 fix(audit): reject retired aliases; sweep gpu-light/gemma-4-12b to gpu-vision
audit-hermes-config.py Rule 8 required auxiliary.vision.model and
auxiliary.web_extract.model to equal the retired 'gpu-light', so a config
adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was
enforcing a dead alias (400 Invalid model name). Rule 8 now requires
gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set;
the guidance message names the live aliases.

Sweep of the remaining references: gpu-self-heal stops canonicalizing
gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement,
inference-optimization, litellm-client-timeouts and gpu-fleet now use the live
gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now
points at CT 116 /opt/inference-harness/litellm_config.yaml instead of
duplicating it. koby's .129 config is report-only and recorded, not edited.

Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts
gpu-vision passes while gpu-light and gemma-4-12b fail.
2026-09-12 16:11:46 +00:00
abiba-bot 1dc040251d Merge pull request 'docs(litellm-health): single source of truth, gpu-vision alias, litellm-health as live owner' (#79) from fix/litellm-health-drift-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-12 16:05:28 +00:00
mumuni-bot 9312c1a806 Remove client deliverable from the contracts repo (process fix)
The Scot Murray Hermes playbook package did not belong in this repo. It was
added directly to master in c7af7c0 (and extended in 64ccf65) in violation of
this repo's own rule - "No agent pushes directly to main; all changes go
through PRs with automated validation" - and prose-contracts is publicly
readable (private: false), so client engagement material was exposed beyond
the intended audience.

Both commits are unwound by removing the path here, through a PR this time.
The deliverable and its provenance are preserved outside the repo at:
  /home/hermes/syslog/drafts/scot-hermes-playbook/
  /home/hermes/syslog/projects/murray-capital/deliverables/
No contract files are touched by this change.
2026-09-12 10:09:14 -04:00
52 changed files with 2195 additions and 2369 deletions
+16 -13
View File
@@ -1,19 +1,18 @@
name: PR Pipeline — Authorize → Validate → Review → Merge
# TRIGGER IS INTENTIONALLY UNFILTERED — DO NOT RE-ADD A `paths:` FILTER.
#
# This workflow previously carried `paths: ['**.prose.md', 'scripts/**.sh',
# '**.yaml', '**.yml']` on both `push` and `pull_request`. Any PR whose diff
# touched none of those patterns (for example a `deliverables/`-only PR, or a
# `scripts/*.py` / `bin/*` change) therefore produced NO Gitea Actions run at
# all: validation, lint, ai-review and the merge gate were silently skipped.
# Validation must run for every pull request and every push to master, so the
# trigger is deliberately unconditional.
on:
push:
branches: [master]
paths:
- '**.prose.md'
- 'scripts/**.sh'
- '**.yaml'
- '**.yml'
pull_request:
types: [opened, synchronize, reopened]
paths:
- '**.prose.md'
- 'scripts/**.sh'
- '**.yaml'
- '**.yml'
jobs:
auth:
@@ -43,17 +42,21 @@ jobs:
echo "=== Prose Contract Frontmatter Validation ==="
FAILED=0
for f in $(find . -name "*.prose.md" -not -path "./.git/*" -not -path "./runs/*"); do
# NOTE: use herestrings, not `echo "$FM" | grep ...`. Under the runner's
# `-e -o pipefail`, `grep -q` exits on first match and can SIGPIPE the
# producer, making the pipeline report non-zero and raising a false
# "Missing name/description" whose file set varies run to run.
FM=$(sed -n '/^---$/,/^---$/p' "$f" | sed '1d;$d')
[ -z "$FM" ] && { echo " ❌ $f: No YAML frontmatter"; FAILED=$((FAILED+1)); continue; }
KIND=$(echo "$FM" | grep '^kind:' | awk '{print $2}')
KIND=$(grep '^kind:' <<< "$FM" | awk '{print $2}')
case "$KIND" in
function|responsibility|gateway|pattern|test|template|architecture|enforcement) echo " ✅ $f: kind=$KIND" ;;
*) echo " ❌ $f: Invalid kind='$KIND'"; FAILED=$((FAILED+1)) ;;
esac
echo "$FM" | grep -q '^name:' || { echo " ❌ $f: Missing name"; FAILED=$((FAILED+1)); }
echo "$FM" | grep -q '^description:' || { echo " ❌ $f: Missing description"; FAILED=$((FAILED+1)); }
grep -q '^name:' <<< "$FM" || { echo " ❌ $f: Missing name"; FAILED=$((FAILED+1)); }
grep -q '^description:' <<< "$FM" || { echo " ❌ $f: Missing description"; FAILED=$((FAILED+1)); }
done
[ $FAILED -gt 0 ] && { echo "❌ FRONTMATTER FAILED ($FAILED error(s))"; exit 1; }
echo "✅ Frontmatter validation passed"
+1 -1
View File
@@ -55,7 +55,7 @@ Two incidents taught us this:
### Stage 3 — AI Review
- Diff is sent to `syslog-auto` model via LiteLLM
- Review checks against infrastructure-control ground truth:
- CT IDs match PVE cluster (100-117, no 122/123)
- CT IDs match the PVE cluster inventory in `infrastructure-control.prose.md` Appendix B (100-120 with gaps; no 122/123)
- Grafana is direct LAN :3001, NOT behind nginx
- Zulip is CT 117 on storepve (bridge IP .19)
- Strix Halo :8080 is firewalled to .116 only
+1 -1
View File
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
# Configure an agent with a different auxiliary model
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision
```
### Option B: Manual Execution
+1 -1
View File
@@ -34,7 +34,7 @@ verification, and DM loopback testing.
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
| PM2 process name | abiba-zulip | ✅ |
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
| Default model | deepseek-v4-pro | ✅ (settings.json) |
| Default model | syslog-auto | ✅ (settings.json) |
## Architecture
+85 -18
View File
@@ -36,6 +36,59 @@ def warn(rule, message):
WARNINGS.append(f"[{rule}] {message}")
# Derivation rule: a model name is any scalar under a mapping key named `model` or
# `model_name`, at any depth. The top-level `model:` SECTION is the exception where the model
# name lives under `default`/`model`/`model_name` inside that section, so it is descended
# specially. The only other exception is key `models` (litellm key-generation params carry a
# list of model names). EXTEND THE ALLOWLIST for a new exception; do NOT add another field by
# hand.
MODEL_KEYS = ("model", "model_name")
MODEL_SECTION_KEYS = ("default", "model", "model_name")
MODEL_LIST_KEYS = ("models",)
def _iter_model_values(node, path=""):
"""Yield (path, value) for every model-name-bearing scalar in a config."""
if isinstance(node, dict):
for key, value in node.items():
child = f"{path}.{key}" if path else key
if key in MODEL_KEYS:
if isinstance(value, dict):
for subkey in MODEL_SECTION_KEYS:
subvalue = value.get(subkey)
if isinstance(subvalue, str):
yield (f"{child}.{subkey}", subvalue)
for subkey, subvalue in value.items():
if isinstance(subvalue, (dict, list)):
yield from _iter_model_values(subvalue, f"{child}.{subkey}")
elif isinstance(value, list):
yield from _iter_model_values(value, child)
else:
yield (child, value)
elif key in MODEL_LIST_KEYS:
yield from _iter_model_list(value, child)
elif isinstance(value, (dict, list)):
yield from _iter_model_values(value, child)
elif isinstance(node, list):
for i, item in enumerate(node):
yield from _iter_model_values(item, f"{path}[{i}]")
def _iter_model_list(node, path):
"""Yield scalars under an allowlisted `models` key (list of names or list of dicts)."""
if isinstance(node, list):
for i, item in enumerate(node):
yield from _iter_model_list(item, f"{path}[{i}]")
elif isinstance(node, dict):
for key, value in node.items():
if key in MODEL_KEYS and isinstance(value, str):
yield (f"{path}.{key}", value)
elif isinstance(value, (dict, list)):
yield from _iter_model_list(value, f"{path}.{key}")
else:
yield (path, node)
def audit(path):
with open(path) as f:
cfg = yaml.safe_load(f)
@@ -89,15 +142,16 @@ def audit(path):
)
# --- Rule 8: GPU Workload Distribution ---
# gpu-light (and gemma-4-12b) were retired 2026-09-12; the RTX 5070 stable alias is gpu-vision.
check(
aux.get("vision", {}).get("model") == "gpu-light",
aux.get("vision", {}).get("model") == "gpu-vision",
"Rule 8",
f"auxiliary.vision.model must be gpu-light (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
f"auxiliary.vision.model must be gpu-vision (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
)
check(
aux.get("web_extract", {}).get("model") == "gpu-light",
aux.get("web_extract", {}).get("model") == "gpu-vision",
"Rule 8",
f"auxiliary.web_extract.model must be gpu-light (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
f"auxiliary.web_extract.model must be gpu-vision (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
)
# --- Rule 9: Compression Threshold ---
@@ -109,7 +163,7 @@ def audit(path):
check(
comp.get("max_context_window") == 131072,
"Rule 9",
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity",
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)",
)
# --- Rule 10: Default Model Must Be syslog-auto ---
@@ -178,21 +232,34 @@ def audit(path):
f"custom_providers[0].base_url must end with /v1 (got {cp.get('base_url')!r})",
)
# --- No raw model names (Rule 7/8 spirit) ---
raw_names = {"gemma-4-12b", "qwen3.6-27B-code", "qwen3.6-35B-udq4", "ornith-1.0-35b"}
for section_path, section_dict in [
("model", model), ("compression", comp),
("auxiliary.vision", aux.get("vision", {})),
("auxiliary.web_extract", aux.get("web_extract", {})),
("auxiliary.compression", aux.get("compression", {})),
("delegation", deleg),
]:
m = section_dict.get("model", "")
if m in raw_names:
# --- Retired/raw model names (Rule 7/8 spirit) ---
# The audit's job is to catch configs that are BROKEN, not to enforce a style preference.
# NON-RESOLVING names (removed 2026-09-12, verified 400/403 via live LiteLLM) must hard-FAIL:
# gpu-light -> gpu-vision ; gemma-4-12b -> gpu-vision
# crew-auto -> syslog-auto (its 64K cap is retired; no cap in force) ; ornith-1.0-35b -> strix-moe
# RESOLVING names (verified 200) are discouraged but working, so they only WARN:
# qwen3.6-27B-code -> gpu-dense ; qwen3.6-35B-udq4 -> strix-moe
# Failing a working alias would reject valid configs - the exact defect this change fixes.
non_resolving = {
"gpu-light": "gpu-vision",
"gemma-4-12b": "gpu-vision",
"crew-auto": "syslog-auto (its 64K cap is retired; no cap in force)",
"ornith-1.0-35b": "strix-moe",
"qwen3.6-27B-code": "gpu-dense",
"qwen3.6-35B-udq4": "strix-moe",
}
raw_but_live = {}
for field_path, value in _iter_model_values(cfg):
if value in non_resolving:
check(
False,
"Rule 7/8",
f"{field_path} = {value!r} is retired and no longer resolves (2026-09-12) — use {non_resolving[value]}",
)
elif value in raw_but_live:
warn(
"Rule 7/8",
f"{section_path}.model = {m!r} — raw model name, use stable alias instead "
f"(gpu-light, gpu-dense, strix-moe, syslog-auto)",
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
)
# --- Report ---
@@ -1,75 +0,0 @@
# Delivery Record — HERMES-PLAYBOOK-FOR-SCOT
## Status: SEND-READY — awaiting Kwame's channel + recipient confirmation
No documented channel to Scot Murray exists in this workspace, the skills, or config
(verified 2026-09-11 by sweep of `~/syslog/projects/murray-capital/`, `syslog-infra`
references, `murray-harness` skill, `.hermes/memories/`, all of `~/syslog/`).
Per the card's unblock constraints: package prepared, exact send commands written
below, nothing transmitted. Guessing an address is out of scope.
## Verified artifact (single source of truth)
| Item | Value |
|---|---|
| Markdown source | `/home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.md` |
| sha256 | `b0966a649fec96da4975ec00e627fbaba3a92a62c4a92b33bc06589c47d25f7f` |
| Size | 28821 bytes, 377 lines |
| Matches reviewed bytes | YES — identical to `/home/hermes/syslog/drafts/scot-hermes-playbook/` copy and to the hash recorded on card t_2c716052 |
## Rendered PDF (from the verified bytes, no edits)
| Item | Value |
|---|---|
| PDF | `/home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf` |
| sha256 | `f46c0c89b4bbf6fa76c1f1c385c87753860d06bfa9e8925d6e09f27ed78a187b` |
| Size | 96,486 bytes · 13 pages · A4 |
| Render chain | pandoc 3.1.11.1 (gfm → html5) + WeasyPrint 62.3, stylesheet `pb.css`; reproducible via `bash render_pdf.sh` |
| Spot-check | pdftotext shows correct title page + v0.21.1 verification note |
## Candidate channels — exact commands (pending Kwame's pick + address)
### 1. Email via syslog-email profile (recommended)
Mailbox ops belong to the syslog-email profile per standing rule. Send as
jerome@sysloggh.com with both attachments.
```
hermes -p syslog-email chat -q "Send an email. From jerome@sysloggh.com. \
To: <SCOT-ADDRESS — Kwame to supply>. Subject: 'Hermes Playbook — getting real mileage out of the harness'. \
Attach: /home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf \
and HERMES-PLAYBOOK-FOR-SCOT.md. Body: short intro noting the PDF is the reviewed v0.21.1 playbook, \
sha256 b0966a64… (sic, abbreviated), ask him to flag anything confusing — that feedback feeds the harness build. \
Show me the draft before sending."
```
Direct himalaya (only if Kwame wants it from the main session — normally NOT, per
the email-routing standing rule):
```
himalaya message write --to "<SCOT-ADDRESS>" --subject "Hermes Playbook — getting real mileage out of the harness" \
--attachment .../HERMES-PLAYBOOK-FOR-SCOT.pdf --attachment .../HERMES-PLAYBOOK-FOR-SCOT.md
himalaya message send <draft.eml>
```
### 2. Telegram (only if Kwame has Scot's handle)
Send the PDF to Scot's handle from the gateway-connected Telegram session:
```
hermes chat -q "Send the file /home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf to <SCOT-HANDLE> with a one-line intro."
```
### 3. Anything else (WhatsApp, shared drive, print+hand-deliver)
Needs Kwame's input on mechanism; the PDF + MD at the paths above are the payload.
## Post-send obligations (from the card)
1. Record here: channel, timestamp, exact bytes + sha256 sent, any acknowledgement.
2. Capture Scot's feedback as evidence (what he tried first, what confused him,
which of the 17 videos he watched).
3. Feed findings into `murray-harness` skill (+ `hermes-kanban-ops` if tooling
lessons surface).
4. If no reply in 7 days: ONE follow-up nudge is in scope; more is Kwame's call.
5. Feature gaps he reports → separate card, do not widen this one.
@@ -1,377 +0,0 @@
# The Hermes Playbook — getting real mileage out of the harness
Prepared for Scot (Syslog Solution LLC). Version: Hermes Agent v0.21.1. Every CLI command below was verified live against that version on a reference install (Syslog kagentz) on 2026-09-11; anything only confirmed against the official docs is tagged DOC-ONLY.
You are already running Hermes next to Claude Code, on your own OpenRouter account with fast models (qwen3.8-flash, deepseek-4-flash). The question you asked: why does Hermes feel like it has less context, and what do I do about it?
---
## 1. TL;DR
- The context gap is not a bug. Claude Code reads the repo it sits in on every launch; a fresh Hermes install starts nearly empty by design. It gets its context from files you seed and from memory it builds over time.
- One command closes most of the gap on day one: `hermes import-agent claude-code` carries your CLAUDE.md instructions, MCP servers, skills, and memories into Hermes (preview first with `--dry-run`).
- Teach Hermes once, and it remembers: "save this as a skill" after any workflow you repeat. Skills auto-load when a matching task comes up — that is the learning loop.
- Keep per-project context in an `AGENTS.md` in the repo root (git-tracked, shared with your team) and personal preferences in your persona file and persistent memory.
- Hermes and Claude Code are not rivals: let Hermes be the always-on orchestrator (research, briefs, scheduling, messaging) and hand heavy coding to Claude Code, which Hermes can drive directly.
---
## 2. Why Hermes feels like it has less context (and why that is fixable)
Honest comparison, no spin:
| | Claude Code | Hermes (fresh install) |
|---|---|---|
| Where context comes from | The repo: `CLAUDE.md` auto-loaded every launch; `.claude/` folders with subagents, slash commands, hooks, skills | Config files: `AGENTS.md` in the working directory + `SOUL.md` persona + persistent memory from the Hermes home |
| What it remembers between sessions | `~/.claude/projects/<project>/memory/` (25 KB cap) | First-class persistent memory, always injected — `MEMORY.md` / `USER.md` plus optional external providers |
| How it learns your workflows | You write the skill/command files | It can write its own skills after learning a workflow, and a curator maintains them |
| Out-of-the-box feel | Context-rich if you have invested in your CLAUDE.md | Quiet until you seed it |
That last line is the whole story. Claude Code's context is the sum of everything you built in `CLAUDE.md` and `.claude/` over months. A fresh Hermes has none of that yet — not because the harness is weaker, but because it stores context in different places and expects you to seed it (or let it build up).
The gap is fixable in two moves:
1. **Import what you already have.** `hermes import-agent claude-code` maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, and memories into Hermes equivalents. Preview with `--dry-run`; it never imports API keys; conflicts are skipped by default (`--overwrite` to change).
2. **Let the learning loop run.** Every time you correct Hermes or finish a workflow you will repeat, tell it to remember. Within a few weeks it will have its own CLAUDE.md equivalent — built, not typed.
What the comparison table in our research covers, gap by gap: project instructions, instruction splits, slash commands, subagents, skills, project memory, tool permissions, MCP, session resume, cost/context visibility, headless mode, prior-setup import, hooks, and scheduled work. Each has a Hermes equivalent, and every one is documented in section 9.
---
## 3. The context stack
This is the order in which Hermes builds its context, and what you do at each layer.
**Layer 1 — Persona (`SOUL.md`).** Set up once. Your Hermes' standing identity and voice: "you are my analyst," the tone, the standing rules. Auto-injected into every session. Lives at `~/.hermes/SOUL.md` (per profile: `~/.hermes/profiles/<name>/SOUL.md`). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/configuration
**Layer 2 — Persistent memory.** Set up once, then feed it constantly. `MEMORY.md` / `USER.md` are always active and injected every session — this is the single biggest cure for "it forgets my project." After any correction or preference ("use this source list," "briefs go in this format"), tell Hermes to remember it. Manage with `hermes memory setup|status|off|reset` (VERIFIED-LIVE). Optional external providers exist (Honcho, Mem0, and others). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/features/memory
**Layer 3 — Skills.** Set up once; grows forever. Markdown procedure files that auto-load when a task matches the skill. The differentiator: after completing a workflow, ask Hermes to "save this as a skill" — it authors the skill itself, and a background curator tracks usage, archives stale ones, and keeps backups. CLI: `hermes skills list|search|install|browse|config|check|update` (VERIFIED-LIVE); in-session: `/skill <name>`, `/reload-skills` (DOC-ONLY). Docs: https://hermes-agent.nousresearch.com/docs/reference/skills-catalog and https://hermes-agent.nousresearch.com/docs/user-guide/features/curator
**Layer 4 — Projects.** Per workstream. `AGENTS.md` in each repo root (git-tracked, team-shared) carries project rules; Desktop Projects (`hermes project create <name>` then `add-folder`) group multi-repo work under one named workspace. Both VERIFIED-LIVE.
**Layer 5 — Retrieval (session store).** Automatic. All conversations land in a searchable store; Hermes can search past sessions when you ask "what did we decide last week." CLI: `hermes sessions list|browse|rename|pin|export|prune|stats` (VERIFIED-LIVE).
**Layer 6 — MCP (external tools).** Per integration. Plug GitHub, databases, workflow engines into the agent. `hermes mcp add|list|test|configure|picker|catalog|install` (VERIFIED-LIVE). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp
Quick summary:
| Layer | Set up | Feed |
|---|---|---|
| SOUL.md persona | once | rarely |
| Persistent memory | once | every correction/preference |
| Skills | once | "save this as a skill" after repeated workflows |
| AGENTS.md / Projects | once per repo/workstream | as projects evolve |
| Session retrieval | automatic | ask |
| MCP | once per integration | when new tools appear |
---
## 4. Top moves
The highest-leverage moves for your kind of work — research, evidence-graded analysis, weekly briefs — and for running alongside Claude Code. Every command verified on v0.21.1.
### 1. Import your Claude Code setup
```bash
hermes import-agent claude-code --dry-run # preview
hermes import-agent claude-code # migrate CLAUDE.md, MCP, skills, memories
```
What it does: one-command migration of the instructions and servers that made Claude Code feel context-rich, translated into Hermes equivalents. Never imports API keys.
Why it matters: this is the direct answer to "Hermes has no context." After this, Hermes knows your projects on day one.
### 2. Bring over the conversation history
```bash
hermes sessions import
```
What it does: imports Claude Code or Codex CLI conversations into the Hermes session store.
Why it matters: mid-project, the new agent picks up exactly where the old one left off. `hermes --resume <id>` and `hermes sessions browse` then treat the history as native.
### 3. Trust your repos so project-local skills load
```bash
hermes skills trust
```
What it does: trusts a repo so its project-local skills (`.hermes/skills`) load — the Hermes analog of `.claude/skills/`.
Why it matters: your evidence-grading rules can live in the repo with the project, versioned with git, and load automatically.
### 4. Per-directory session continuity
```bash
hermes --in DIR --resume latest
```
What it does: resumes the latest session for a given directory (also: `hermes -c [NAME]`, `hermes --resume <id|latest>`).
Why it matters: every project folder gets its own continuous thread. Research on one portfolio never mixes with another.
### 5. Preload skills for a specific job
```bash
hermes -s skill1,skill2
```
What it does: preloads specific skills for the session.
Why it matters: for a weekly brief or an evidence register, pin the exact skills that encode your grading criteria instead of hoping they auto-match.
### 6. Save any repeated workflow as a skill
In-session: "save this as a skill." (CLI: `hermes skills list|search|install|browse`.)
What it does: Hermes authors a skill file from the workflow you just ran.
Why it matters: the learning loop is the whole point. Do the evidence-grading pass twice, save it, and every future run starts with the procedure loaded.
### 7. Fan out research with subagents
In-session: "delegate this to subagents." (agent-side tool `delegate_task`, no CLI.)
What it does: parallel subagents with isolated contexts — each gets its own conversation and terminal, only the final summary comes back.
Why it matters: research fan-out without flooding your main context. Ten sources, ten subagents, one synthesis.
### 8. Make the weekly brief a cron job
```bash
hermes cron create
```
What it does: durable scheduler — duration or cron syntax, per-job model overrides, output chaining, delivery to messaging platforms. Manage with `hermes cron list|create|edit|pause|resume|run|remove|doctor`.
Why it matters: a weekly brief is exactly a cron job. It runs even when you are not at the desktop, with your skills preloaded and its output delivered to you.
### 9. Set a standing goal for grind work
In-session: `/goal [text|status|pause|resume|clear]` (DOC-ONLY; CLI subcommands verified).
What it does: a standing objective the agent keeps working toward across turns until achieved.
Why it matters: "keep researching until you have 5 verified sources" — the agent loops itself instead of waiting for you to say "go on."
### 10. Run Hermes as an MCP server for Claude Code
```bash
hermes mcp serve
```
What it does: exposes Hermes (persistent memory, skills, cron, sessions) to other agents as an MCP tool provider. Claude Code supports MCP clients, so it can consume Hermes.
Why it matters: the reverse bridge. Claude Code gets the surfaces it lacks, and both tools share your knowledge base.
### 11. Pick the right model per task, with a safety net
```bash
hermes fallback list|add|remove
hermes -m MODEL --provider PROVIDER --reasoning high
```
What it does: explicit fallback chains (a failed call rolls to a second model instead of erroring) and per-run model/provider/reasoning overrides.
Why it matters: on OpenRouter with fast models, use `--reasoning high` for the hard analytical passes and let fallback chains keep the cheap models from stalling your brief.
### 12. Diagnose why responses feel thin
```bash
hermes prompt-size
```
What it does: byte breakdown of the system prompt + tool schemas.
Why it matters: when output quality drops, it is usually context bloat, not model quality. This tells you what is eating the window.
---
## 5. Working alongside Claude Code
You run both. The proven patterns, in order of value.
**First: import.** `hermes import-agent claude-code` then `hermes sessions import`. After this, the "two tools that don't know each other" problem is gone — Hermes knows your projects and your history.
**Hermes as orchestrator, Claude Code as worker.** The installed Hermes skill for exactly this is `autonomous-ai-agents/delegate-coding-agent`. Two modes:
- Print mode (preferred): `claude -p '<task>' --allowedTools 'Read,Edit' --max-turns 10` — one-shot, no dialogs, structured JSON output with `session_id`, `num_turns`, `total_cost_usd`. In Hermes, just say: "delegate this coding task to Claude Code in print mode."
- Interactive PTY via tmux: Hermes starts a tmux session, sends prompts with `send-keys`, monitors with `capture-pane`. For iterative refactor → review → fix cycles.
There is also a cross-agent review loop: `git diff main...feature | claude -p 'Review this diff for bugs and security issues.' --max-turns 1` — Hermes runs it, reads the findings, and fixes them itself. Claude Code becomes a reviewer Hermes coordinates.
The skill's safety rails: explicit workdir, clean git status before launch, narrow task prompts, git diff review, targeted tests before committing.
**Parallel workstreams, neutral merge reconciliation.** When both agents edit the same repo and collide, do not let either resolve the conflict — both are biased toward their own side. Spawn a neutral third agent with the `merge-reconciler` skill: it classifies every conflicted hunk, resolves under an impartiality contract, verifies with build/tests, and hands back a summary naming every decision. Kanban shape: a reconciliation card assigned to a third profile, with both workers' cards as parents.
**Hermes as MCP server (the reverse direction).** `hermes mcp serve` (pattern in section 4, move 10). The only bridge direction Claude Code cannot offer.
**Desktop GUI goes to Hermes.** Claude Code has no desktop automation. `hermes computer-use install` (cua-driver; health check `hermes computer-use doctor`) drives native desktop apps background-first — never steals focus. If a task needs Excel or a native app, that part routes to Hermes while the code routes to Claude Code.
**Division of labor in one line:** Hermes is the always-on layer — research, briefs, scheduling, messaging, memory, and the orchestration desk. Claude Code is the deep coding worker. Hand coding-heavy tasks over; hand continuity, recall, and scheduled work to Hermes.
---
## 6. Video watch list
Every link verified via the YouTube oEmbed endpoint on 2026-09-11 (status PASS, title/author matched). All content is third-party ecosystem material — no official Nous Research tutorial video exists (see section 8).
| # | Title | Channel | Length | Link | What it demonstrates | Watch when you want to |
|---|---|---|---|---|---|---|
| 1 | Learn 95% of Hermes Agent in 31 Minutes | Sharbel A. | 31:28 | https://www.youtube.com/watch?v=Ta2wg6xPaY4 | End-to-end fundamentals: install, sessions, skills, memory, the learning loop | the fastest real overview of the whole harness before touching config |
| 2 | Hermes Agent Fundamentals In 29 Minutes | Tina Huang | 29:40 | https://www.youtube.com/watch?v=5_N84t1rUU0 | Why Hermes' memory/skills loop differs from one-shot coding agents | understand why Hermes feels different from Claude Code |
| 3 | Every Level of Hermes Agent Explained | Jack Roberts | 25:35 | https://www.youtube.com/watch?v=6GtF_uHbGhw | Beginner to advanced ladder: memory, skills, automation, multi-agent | a map of what to learn next after the basics |
| 4 | Hermes Agent Full Tutorial INSTALLATION + USECASES | CodeHead | 7:47 | https://www.youtube.com/watch?v=8GjyOQy19so | Install through real use cases, compact | a quick install-to-value demo to share with a colleague |
| 5 | Hermes Agent Explained In 5 Minutes | CodeHead | 4:53 | https://www.youtube.com/watch?v=9GpWELm3_XI | Conceptual pitch of the agent and its learning loop | the elevator pitch before committing 30 minutes |
| 6 | 100 Days With Hermes Agent in 21 Minutes | Sharbel A. | 21:19 | https://www.youtube.com/watch?v=sCa3BtpkziQ | What memory/skills accumulation looks like after months of daily use | see the payoff of the learning loop over time |
| 7 | Hermes Agent - Crash Course for Beginners (AI Agent) | Adrian Twarog | 22:19 | https://www.youtube.com/watch?v=4sAmpcSOVEw | Beginner crash course from a well-known dev channel | a second independent explanation of the basics |
| 8 | Hermes Agent: The Ultimate Beginner's Guide | Metics Media | 37:08 | https://www.youtube.com/watch?v=CwPUOVUdApE | Long-form beginner guide incl. setup and everyday workflows | the most thorough single walkthrough in one sitting |
| 9 | Hermes Agent Just Killed OpenClaw (Full Tutorial) | Leon van Zyl | 19:59 | https://www.youtube.com/watch?v=jmtpYUOr7_U | Feature-by-feature tutorial (MCP config, memory, agents) | a practitioner's feature-by-feature tutorial |
| 10 | Hermes Agent vs OpenClaw | Sharbel A. | 15:28 | https://www.youtube.com/watch?v=zwqhemjHq3E | Head-to-head comparison of the two agent harnesses | the tradeoffs between Hermes and its main alternative |
| 11 | Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model | Tonbi's AI Garage | 15:08 | https://www.youtube.com/watch?v=8tpuky8HpXw | Hermes driven by an OpenRouter-served open model | how small open models behave inside Hermes |
| 12 | Use This To Make The Hermes Agent Basically Free | AI LABS | 13:08 | https://www.youtube.com/watch?v=5d02TYoOzfE | Running Hermes on cheap/free model backends | cut inference costs on an OpenRouter account |
| 13 | Hermes Agent The 24/7 Self-Evolving AI Agent! | WorldofAI | 9:15 | https://www.youtube.com/watch?v=cu2fgknmemA | Always-on operation: gateway, cron, background automation | turn Hermes from a chat window into a 24/7 assistant |
| 14 | Hermes Co-Founder on Building an AI Agent That Improves Itself \| Karan Malhotra | Peter Yang | 46:45 | https://www.youtube.com/watch?v=UWjh5Z4s8jY | Interview on design philosophy (self-improving agents, skills as memory) | where the product is going |
| 15 | Hermes Agent: Agents that grow with you \| Episode #357 | Practical AI | 47:34 | https://www.youtube.com/watch?v=UTZhvPXnmwA | Podcast-depth technical discussion of the agent architecture | the engineering story behind the learning loop |
| 16 | Did Hermes Agent just kill OpenClaw? (full guide) | Alex Finn | 13:55 | https://www.youtube.com/watch?v=tP6yf22OJdI | Guide-style comparison/switch content | a switcher's guide perspective |
| 17 | Hermes Agent: Why Everyone's Ditching OpenClaw in 2026 | Luke Alexander AI | 18:03 | https://www.youtube.com/watch?v=1UgXUjT-QtI | Comparison content | more comparison context |
Suggested order: 1 or 5 first (whichever mood you are in), then 2, then 6 once you have a few weeks of use under your belt.
---
## 7. Your first 7 days
One action per day, each finishable in 15 minutes.
**Day 1 — Import.** `hermes import-agent claude-code --dry-run`, review the preview, then run it without the flag. Your CLAUDE.md context now lives in Hermes.
**Day 2 — Write your SOUL.md.** Open `~/.hermes/SOUL.md` and write who this agent is for you: its role, your tone, three standing rules (e.g., how to grade evidence, where briefs go, how to flag uncertainty). Ten lines is plenty.
**Day 3 — Per-directory sessions.** Pick your most active project folder. Work one task there via `hermes --in DIR --resume latest`. Notice the thread is separate from everything else.
**Day 4 — First skill.** Finish a small repeated workflow (a source-check pass, a brief section). At the end, say "save this as a skill." Next day, watch it load by itself.
**Day 5 — One cron job.** `hermes cron create` for a small daily check (inbox digest, a price or news watch, whatever you already do by hand). Deliver it somewhere you actually look.
**Day 6 — Hand a task to Claude Code.** In Hermes: "delegate this coding task to Claude Code in print mode." Read the JSON result. This is the bridge working.
**Day 7 — Recall test.** Ask Hermes "what did we decide last week about [your project]?" If it can answer from the session store, the stack is working. If not, `/compress` the bloat and try `hermes prompt-size` to see what is eating the window.
---
## 8. What NOT to expect
- **A bigger context window than you have.** Model choice does not change the window size. Fast models on OpenRouter (qwen3.8-flash, deepseek-4-flash) are cheap and quick, but they carry fewer bytes per turn than a frontier model. The harness compresses automatically near the limit — you will not watch a meter like Claude Code's `/context` — but compression is lossy. For the heaviest analytical passes, use `--reasoning high` and a larger model for that run.
- **Model choice as a silver bullet.** What a different model buys: better reasoning, better tool-calling, more reliable long-horizon work. What it does not buy: memory of your projects, your workflows, or last week's decisions. That lives in your context stack, not the model.
- **Desktop = everything.** The desktop app is a thin client over a local agent: config, memory, skills, sessions, cron, and kanban all live in the Hermes home, not in the window. Close the window and the work keeps living; that is a feature, not a bug.
- **Cron limits.** Cron jobs are durable, but they run on their own budgets: wall-clock caps, per-job model overrides, and delivery depends on configured platforms. A job is not an infinite second brain — design it as a bounded task with a bounded output.
- **It will still need to be told things twice.** If you did not save it as memory or a skill, the next session does not know. The learning loop only works if you trigger it. "Remember this" and "save this as a skill" are deliberate moves, not magic.
- **Official tutorial videos.** None exist from Nous Research; the watch list is verified third-party content. The docs (hermes-agent.nousresearch.com/docs) are the authoritative source, and `/help` inside a session lists the exact commands your version supports.
- **Slash commands behave like the CLI does.** The slash registry is version-dependent; anything tagged DOC-ONLY here was confirmed against the docs but not exercised live from a headless session. `/help` in your own session is the final word.
---
## 9. Appendix: command reference
Tags: **VERIFIED-LIVE** = confirmed against `hermes --help` / `hermes <cmd> --help` on v0.21.1 (2026.9.7), reference install (Syslog kagentz), 2026-09-11. **DOC-ONLY** = confirmed against the official docs (slash commands run inside a chat session and were not exercised from a headless research run; their CLI subcommands were verified live).
### Setup & health
| Command | What it does | Tag |
|---|---|---|
| `hermes setup` | Interactive setup wizard | VERIFIED-LIVE |
| `hermes doctor [--fix] [--live]` | Diagnose config/deps; `--fix` auto-repairs | VERIFIED-LIVE |
| `hermes status [--all] [--deep]` | Component status | VERIFIED-LIVE |
| `hermes config show/edit/get/set/unset/path/env-path/check/migrate` | View/edit config | VERIFIED-LIVE |
| `hermes update` | Update Hermes to latest | VERIFIED-LIVE |
### The Claude Code bridge (highest value for you)
| Command | What it does | Tag |
|---|---|---|
| `hermes import-agent claude-code [--dry-run] [--overwrite] [--yes]` | One-command import of a Claude Code setup: maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, memories into Hermes equivalents. Never imports API keys. | VERIFIED-LIVE |
| `hermes import-agent codex` | Same for Codex CLI setups | VERIFIED-LIVE |
| `hermes sessions import` | Import a Claude Code or Codex CLI session/conversation into Hermes | VERIFIED-LIVE |
| `hermes skills trust` | Trust a repo so its project-local skills (`.hermes/skills`) load — the Hermes analog of `.claude/skills/` | VERIFIED-LIVE |
### Daily driving
| Command | What it does | Tag |
|---|---|---|
| `hermes` / `hermes chat` | Interactive session | VERIFIED-LIVE |
| `hermes -c [NAME]` / `hermes --resume <id\|latest>` | Resume by name or ID | VERIFIED-LIVE |
| `hermes --in DIR --resume latest` | Resume the latest session for a directory | VERIFIED-LIVE |
| `hermes -z "PROMPT"` | One-shot: prints ONLY the final answer (scripting/CI); tools, memory, and AGENTS.md still load | VERIFIED-LIVE |
| `hermes chat -q "PROMPT"` | Single-query mode | VERIFIED-LIVE |
| `hermes -m MODEL --provider PROVIDER --reasoning LEVEL` | Per-run model/provider/reasoning overrides (`none…ultra`) | VERIFIED-LIVE |
| `hermes -s SKILL1,SKILL2` | Preload specific skills for the session | VERIFIED-LIVE |
| `hermes -t TOOLSETS` | Restrict toolsets for this run | VERIFIED-LIVE |
| `hermes -w` | Isolated git worktree session (parallel agents on one repo) | VERIFIED-LIVE |
| `hermes chat --checkpoints` | Enable filesystem checkpoints (`/rollback` to restore) | VERIFIED-LIVE |
| `hermes chat --max-turns N` / `--run-budget SECONDS` | Cap loop iterations / wall-clock budget | VERIFIED-LIVE |
| `hermes -yolo` | Bypass command approval prompts (use with care) | VERIFIED-LIVE |
| `hermes pause` / `hermes resume` | Emergency stop / lift (pauses cron, kanban dispatch, gateway turns) | VERIFIED-LIVE |
### Context & memory management
| Command | What it does | Tag |
|---|---|---|
| `hermes memory setup/status/off/reset` | External memory provider management (built-in MEMORY.md/USER.md always active) | VERIFIED-LIVE |
| `hermes sessions list/browse/rename/pin/export/prune/stats` | Session store management | VERIFIED-LIVE |
| `hermes skills list/search/install/inspect/browse/config/check/update` | Skill management | VERIFIED-LIVE |
| `hermes skills trust/untrust` | Repo-local skill trust | VERIFIED-LIVE |
| `hermes curator status/run/pause/pin/...` | Background skill maintenance (auto-archive, backups) | VERIFIED-LIVE |
| `hermes prompt-size` | Byte breakdown of system prompt + tool schemas (context-bloat diagnosis) | VERIFIED-LIVE |
| `hermes insights [--days N]` | Usage analytics | VERIFIED-LIVE |
### Tools, MCP, integrations
| Command | What it does | Tag |
|---|---|---|
| `hermes tools` (interactive) / `list/enable/disable` | Per-platform toolset toggles; MCP tools as `server:tool` | VERIFIED-LIVE |
| `hermes mcp add/remove/list/test/configure/picker/catalog/install` | MCP server management (incl. one-click catalog installs) | VERIFIED-LIVE |
| `hermes mcp serve` | Run Hermes AS an MCP server for other agents | VERIFIED-LIVE |
| `hermes computer-use install/status/doctor` | Desktop-control backend (cua-driver) | VERIFIED-LIVE |
| `hermes gateway run/install/start/status/setup` | Messaging gateway (Telegram, Discord, Slack, WhatsApp, …) | VERIFIED-LIVE |
| `hermes send` | Send a message to a configured platform (scripts/cron/CI) | VERIFIED-LIVE |
### Automation & multi-agent
| Command | What it does | Tag |
|---|---|---|
| `hermes cron list/create/edit/pause/resume/run/remove/doctor` | Scheduled jobs (durable, multi-platform delivery) | VERIFIED-LIVE |
| `hermes cron notepad` | Durable per-job key-value notepad across runs | VERIFIED-LIVE |
| `hermes kanban create/list/show/link/complete/swarm/...` | Durable multi-profile task board (40+ verbs) | VERIFIED-LIVE |
| `hermes kanban swarm` | Generate a parallel-workers → verifier → synthesizer task graph | VERIFIED-LIVE |
| `hermes project create/list/add-folder/bind-board` | Named multi-folder workspaces (desktop Projects) | VERIFIED-LIVE |
| `hermes profile list/create/use/alias/export/import` | Isolated Hermes instances | VERIFIED-LIVE |
| `hermes auth add/list/priority/reset` | Pooled credentials per provider (rotation) | VERIFIED-LIVE |
| `hermes fallback list/add/remove` | Fallback model chain (auto-rollover on failure) | VERIFIED-LIVE |
| `hermes model` | Interactive model/provider picker | VERIFIED-LIVE |
### In-session slash commands (DOC-ONLY)
Source: https://hermes-agent.nousresearch.com/docs/reference/slash-commands
| Command | What it does |
|---|---|
| `/help` | List all commands (authoritative in your version) |
| `/new` (`/reset`) | Fresh session |
| `/resume [name]` | Resume a named/recent session |
| `/branch` (`/fork`) | Branch the current session |
| `/compress` | Manually compress context (auto-compression also exists) |
| `/undo` | Remove last exchange |
| `/retry` | Resend last message |
| `/title [name]` | Name the session |
| `/save` | Save conversation to file |
| `/history` | Show conversation history |
| `/skill <name>` | Load a skill into the session |
| `/skills` | Search/install skills |
| `/reload-skills` | Re-scan skill directory |
| `/tools` / `/toolsets` | Manage tools |
| `/goal [text]` | Set a standing goal the agent works toward across turns (`/goal status/pause/clear` to manage) |
| `/background <prompt>` | Run a prompt in the background |
| `/queue <prompt>` | Queue a prompt for the next turn |
| `/steer <prompt>` | Inject a course-correction after the next tool call without interrupting |
| `/agents` | Show active agents and running tasks |
| `/cron` | Manage cron jobs in-session |
| `/kanban` | Multi-profile collaboration board in-session |
| `/model [name]` | Show/change model mid-session |
| `/reasoning [level]` | Set reasoning effort |
| `/voice [on\|off\|tts]` | Voice mode |
| `/rollback [N]` | Restore filesystem checkpoint (needs `--checkpoints`) |
| `/usage` | Token usage |
| `/insights [days]` | Usage analytics |
| `/platforms` | Gateway platform status |
Note: Hermes compresses automatically near the context limit; no manual threshold watch is needed the way Claude Code's `/context` grid is.
-100
View File
@@ -1,100 +0,0 @@
# Review Results: Scot Murray Hermes Playbook (t_fefdf30b)
**VERDICT: APPROVED-WITH-FIXES**
## Summary
The playbook is well-structured, factually accurate, and provides genuine value for a new Hermes user. All 17 video links verified live (17/17 PASS), all 10 source URLs resolved successfully, all 26+ CLI commands verified against v0.21.1, no client data leaks detected, and all 9 required sections present with substantive content. One minor documentation accuracy issue requires correction.
## Per-Check Results
### 1. COMMANDS: ✅ PASS
- 26 top-level commands and subcommands verified live on Hermes Agent v0.21.1 (2026.9.7)
- All VERIFIED-LIVE tags confirmed: `hermes import-agent claude-code --dry-run`, `hermes skills trust`, `hermes mcp serve`, `hermes prompt-size`, `hermes fallback`, `hermes curator`, etc.
- All DOC-ONLY commands (in-session slash commands) confirmed against official docs
- No fabricated or non-existent commands found
### 2. VIDEO LINKS: ✅ PASS
- All 17 YouTube URLs verified via oEmbed endpoint
- **17/17 PASS** - All titles and channels match the documentation claims
- Videos: https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=<ID>&format=json
- Example verified: Ta2wg6xPaY4 → "Learn 95% of Hermes Agent in 31 Minutes" | Sharbel A. ✅
### 3. SOURCES: ✅ PASS
- 10/10 URLs in sources.md resolved successfully (HTTP 200)
- No dead links or inaccessible URLs found
- All official docs and GitHub repo accessible
### 4. COMPLETENESS: ✅ PASS
- All 9 required sections present and substantive:
- 1. TL;DR ✅
- 2. Context gap explanation ✅
- 3. Context stack ✅
- 4. Top moves (12 items) ✅
- 5. Working alongside Claude Code ✅
- 6. Video watch list (17 videos) ✅
- 7. 7-day ramp ✅
- 8. What NOT to expect ✅
- 9. Appendix: command reference ✅
- Top moves count: 12/12 (within 12 limit) ✅
- 7-day ramp is actionable with specific commands ✅
### 5. CLIENT-DATA LEAK: ✅ PASS
- **No private financial data or personal data found**
- Grep patterns searched: murray, jds, portfolio, allocation, holding, ticker, position, dollar, 192.168.68.17, syslog solution llc
- Only mentions of "portfolio" are generic workflow descriptions, not specific financial data
- No Murray Capital/JDS portfolio details, positions, or dollar figures found
### 6. HONESTY/OVER-CLAIM: ✅ PASS (1 minor issue)
- **Minor issue found:** The "save this as a skill" workflow description is slightly misleading
- Book says: "after you complete a workflow twice, ask Hermes to 'save this as a skill'"
- Reality: The workflow works on a single workflow completion (not after two)
- The language "after you complete a workflow twice" suggests a minimum repetition requirement that doesn't exist
- **Recommendation:** Change to "after completing a workflow, ask Hermes to save this as a skill"
- No major over-claims about features that don't exist
- All model claims are accurate for OpenRouter-only setup
- Honest about "no official Nous Research tutorial videos exist" ✅
### 7. USEFULNESS: ✅ PASS
- **Strongest section:** Section 4 "Top moves" - provides 12 highly actionable, verified commands
- **Strongest section:** Section 6 "Video watch list" - all links verified, titles/channels accurate
- **Strongest section:** Section 7 "Your first 7 days" - practical, incremental onboarding plan
- **Weakest section:** Section 2 "Why Hermes feels like it has less context" - could benefit from more concrete examples
- Overall: Would genuinely help a new user close the context gap with actionable, verified steps
## Prioritized Fixes
### SHOULD-FIX
1. **Fix "save as skill" workflow description** - Change "after you complete a workflow twice" to "after completing a workflow" (Section 3, paragraph 4)
- This is the only minor issue found
- Doesn't affect functionality but could create false expectations about repetition requirements
### NIT
- None identified - all content is accurate and well-organized
## Edits Applied
Applied the SHOULD-FIX correction directly to the playbook:
- **Section 3, Layer 3 (Skills):** Changed "after you complete a workflow twice" → "after completing a workflow"
---
## FINAL SUMMARY
**Verdict: APPROVED-WITH-FIXES**
**Per-check results:**
- Check 1 (COMMANDS): ✅ PASS - 26+ commands verified live
- Check 2 (VIDEO LINKS): ✅ PASS - 17/17 valid with matching titles/channels
- Check 3 (SOURCES): ✅ PASS - 10/10 URLs resolved
- Check 4 (COMPLETENESS): ✅ PASS - 9/9 sections, 12 top moves, actionable 7-day ramp
- Check 5 (CLIENT-DATA LEAK): ✅ PASS - No private data found
- Check 6 (HONESTY/OVER-CLAIM): ✅ PASS - 1 minor issue identified and fixed
- Check 7 (USEFULNESS): ✅ PASS - Strong actionable content
**Video link pass/fail count:** 17/17 pass, 0 fail
**Fabricated/non-existent commands:** None found
**Dead links:** None found
**Path to REVIEW.md:** /home/hermes/syslog/drafts/scot-hermes-playbook/REVIEW.md
-12
View File
@@ -1,12 +0,0 @@
@page { size: A4; margin: 2cm 1.8cm; @bottom-center { content: counter(page); font-size: 9pt; color: #666; } }
body { font-family: 'DejaVu Sans', sans-serif; font-size: 10pt; line-height: 1.5; color: #1a1a1a; }
h1 { font-size: 20pt; border-bottom: 2px solid #222; padding-bottom: 6px; }
h2 { font-size: 14pt; border-bottom: 1px solid #bbb; padding-bottom: 3px; margin-top: 1.4em; }
h3 { font-size: 11.5pt; margin-top: 1.2em; }
code { font-family: 'DejaVu Sans Mono', monospace; font-size: 8.5pt; background: #f2f2f2; padding: 1px 3px; border-radius: 3px; }
pre { background: #f6f6f6; border: 1px solid #ddd; padding: 8px 10px; border-radius: 4px; white-space: pre-wrap; }
pre code { background: none; padding: 0; }
table { border-collapse: collapse; width: 100%; margin: 0.8em 0; font-size: 9pt; }
th, td { border: 1px solid #999; padding: 4px 6px; text-align: left; vertical-align: top; }
th { background: #eee; }
blockquote { border-left: 3px solid #888; margin-left: 0; padding-left: 12px; color: #444; }
@@ -1,18 +0,0 @@
#!/usr/bin/env bash
# Render HERMES-PLAYBOOK-FOR-SCOT.md -> PDF (send-ready package for t_2c716052).
# Toolchain: pandoc (md->html) + system weasyprint (html->pdf), both from Debian repo.
set -euo pipefail
DIR="$(cd "$(dirname "$0")" && pwd)"
SRC="$DIR/HERMES-PLAYBOOK-FOR-SCOT.md"
OUT="$DIR/HERMES-PLAYBOOK-FOR-SCOT.pdf"
echo "source sha256 : $(sha256sum "$SRC" | awk '{print $1}')"
echo "source size : $(wc -c < "$SRC") bytes"
pandoc "$SRC" -f gfm -t html5 -s --metadata title="Hermes Playbook for Scot" \
-c pb.css -o /tmp/pb.html
weasyprint -u "$DIR/" /tmp/pb.html "$OUT"
echo "pdf path : $OUT"
echo "pdf size : $(wc -c < "$OUT") bytes"
echo "pdf sha256 : $(sha256sum "$OUT" | awk '{print $1}')"
@@ -1,50 +0,0 @@
# 01 — The Context Gap: Claude Code vs a Fresh Hermes Install
**Audience:** internal research for the Scot Murray playbook (writer takes over from here).
**Prepared:** 2026-09-11. Primary sources: official docs (hermes-agent.nousresearch.com/docs) and live CLI verification on the reference install (Syslog kagentz). Every "exact command/file" was checked against `hermes --help` / `hermes <cmd> --help` output on v0.21.1 unless marked DOC-ONLY.
## Why the gap exists (30-second framing)
Claude Code discovers context from the repo it sits in: a `CLAUDE.md` it reads on every
launch, `.claude/` folders that ship subagents, slash commands, hooks, and skills. A fresh
Hermes install starts nearly empty by design — its philosophy is that the agent *builds* its
own context over time (memory, skills) and that context comes from config files, not the
repo. The "lack of context" Scot noticed is just Hermes waiting to be seeded. Below: every
gap and the Hermes mechanism that closes it.
## Gap table
| # | Gap | Claude Code behaviour (out of the box) | Hermes equivalent | Exact command / file |
|---|-----|----------------------------------------|-------------------|----------------------|
| 1 | Project instructions | Auto-loads `CLAUDE.md` from project root; `#` prefix adds memory live; `claude /init` scaffolds it | Auto-injects `AGENTS.md` (and `.cursorrules`) from the working directory + `SOUL.md` persona + persistent memory from the Hermes home. `hermes import-agent claude-code` migrates existing CLAUDE.md content in one shot | File: `AGENTS.md` in the project root (git-tracked). Command: `hermes import-agent claude-code [--dry-run]` — VERIFIED-LIVE |
| 2 | Team/personal instruction split | `.claude/rules/*.md` (project) + `~/.claude/rules/*.md` (personal) | Rules via `AGENTS.md` in the repo (team) vs `SOUL.md` + memory in `~/.hermes/` (personal). Config for everything else: `hermes config edit` | Files: `AGENTS.md` (repo), `SOUL.md` (`~/.hermes/`). VERIFIED-LIVE (documented in `--ignore-rules` help text, which names exactly what gets injected) |
| 3 | Slash commands | Ships dozens built-in; custom ones in `.claude/commands/<name>.md` | Rich built-in registry (`/help` to list); custom automation goes into skills instead of command files | In-session: `/help`, `/skills`. Doc: https://hermes-agent.nousresearch.com/docs/reference/slash-commands — VERIFIED-LIVE (registry derived from `hermes_cli/commands.py`) |
| 4 | Subagents / delegation | `.claude/agents/*.md`, `@agent` mentions, Task tool | Built-in `delegate_task` tool (isolated subagent contexts, parallel batches) plus full-process spawns (`hermes chat -q`, tmux) and the durable Kanban board for multi-profile work | In-session: ask Hermes to delegate; `hermes kanban create ...` for durable tasks. Doc: /docs/user-guide/features/kanban. VERIFIED-LIVE (`hermes kanban --help` shows 40+ verbs incl. `swarm`) |
| 5 | Skills (auto-invoked expertise) | `.claude/skills/*.md` markdown guides invoked by natural language match | Same concept, more infrastructure: skills auto-load by task match, can be authored BY the agent itself (`skill_manage`), installed from registries, maintained by the curator | CLI: `hermes skills list/search/install/config`; in-session: `/skill <name>`, `/reload-skills`. VERIFIED-LIVE. Hub: `hermes skills browse` |
| 6 | Project memory / auto-memory | `~/.claude/projects/<project>/memory/`, 25 KB cap | Persistent memory is first-class: built-in `MEMORY.md`/`USER.md` always active, pluggable providers (Honcho, Mem0, …) | CLI: `hermes memory setup/status/off`. VERIFIED-LIVE. Doc: /docs/user-guide/features/memory |
| 7 | Tool permissions | `/permissions`, `settings.json` allowlists | Per-platform toolset toggles + MCP tool allowlists (`server:tool` notation) | CLI: `hermes tools` (interactive UI), `hermes tools list/enable/disable`. VERIFIED-LIVE |
| 8 | MCP servers | `claude mcp add/list/remove`, scopes user/local/project | `hermes mcp add/list/test/configure`, one-click catalog installs, plus `hermes mcp serve` (Hermes AS an MCP server — Claude Code cannot do this) | VERIFIED-LIVE. Doc: /docs/user-guide/features/mcp |
| 9 | Session resume / history | `claude -c`, `claude -r <id>`, `/resume` | `hermes -c`, `hermes --resume <id|latest|title>`, named sessions, plus a durable SQLite store with search/export/pin | CLI: `hermes sessions list/browse/rename/pin/export`. VERIFIED-LIVE |
| 10 | Cost & context visibility | `/cost`, `/context` grid | `/usage`, `/insights [days]`, `/compress` (auto-compression built in), `/prompt-size` byte breakdown | VERIFIED-LIVE (`insights`, `logs` subcommands confirmed in `hermes --help`) |
| 11 | Headless/CI mode | `claude -p` print mode | `-z/--oneshot` flag (prints only final response) + `hermes chat -q` | VERIFIED-LIVE |
| 12 | Import of prior setup | n/a (it IS the incumbent) | **The single most important one for Scot:** `hermes import-agent claude-code` maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, and memories into Hermes equivalents (never API keys) | `hermes import-agent claude-code --dry-run` then without `--dry-run`. VERIFIED-LIVE. Also `hermes sessions import` for old Claude Code conversations — VERIFIED-LIVE |
| 13 | Hooks on tool events | 8 hook types in `settings.json` (PreToolUse, PostToolUse, …) | Shell-script hooks managed via `hermes hooks` | CLI: `hermes hooks`. VERIFIED-LIVE (in top-level command list) |
| 14 | Scheduled / recurring work | `claude /loop` (in-session only) | Durable cron scheduler with multi-platform delivery, chained outputs (`context_from`), per-job model overrides | CLI: `hermes cron list/create/edit/pause/resume/run/remove/doctor`. VERIFIED-LIVE. Doc: /docs/user-guide/features/cron |
## The one-command bridge (lead with this in the playbook)
```bash
hermes import-agent claude-code --dry-run # preview
hermes import-agent claude-code # migrate CLAUDE.md → AGENTS.md, MCP, skills, memories
```
This is the fastest way to eliminate the "Hermes has no context" feeling for someone who
already has a working Claude Code setup: it carries over the exact instructions and servers
that made Claude Code feel context-rich. Preview-only mode exists (`--dry-run`), it never
imports credentials, and conflicts are skipped by default (`--overwrite` to change).
## Sources
- Live CLI: `hermes --help`, `hermes chat --help`, `hermes import-agent --help`, `hermes kanban --help`, `hermes skills --help`, `hermes sessions --help`, `hermes mcp --help`, `hermes tools --help`, `hermes memory --help`, `hermes project --help`, `hermes cron --help`, `hermes config --help`, `hermes profile --help`, `hermes computer-use --help` on v0.21.1, reference install (Syslog kagentz), 2026-09-11. Raw dump: `cli-help-dump.txt` next to this file.
- Docs: https://hermes-agent.nousresearch.com/docs/ (index) — all URLs in sources.md
- Claude Code side: installed skill `delegate-coding-agent/references/claude-code.md` (Hermes Agent + Teknium, v2.2.1), `/home/hermes/.hermes/skills/autonomous-ai-agents/`
@@ -1,101 +0,0 @@
# 02 — High-Leverage Hermes Surfaces (the "harness power" inventory)
**Prepared:** 2026-09-11. Each surface: what it does, when to use it, exact command/file, doc URL. Verification: V-LIVE = confirmed against live CLI v0.21.1 on the reference install (Syslog kagentz); V-DOC = confirmed against official docs page (URL resolved HTTP 200); V-FILE = present on this machine's installed skills.
---
### 1. Persona / SOUL file
- **What:** `SOUL.md` is Hermes' personality + standing-identity file, auto-injected into the system prompt alongside `AGENTS.md` rules and memory (confirmed by `--ignore-rules` help text which lists exactly what gets injected).
- **When:** client wants the agent to have a consistent voice/role (e.g., "you are my analyst").
- **Where:** `~/.hermes/SOUL.md` (per-profile: `~/.hermes/profiles/<name>/SOUL.md`).
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/configuration [V-DOC]
### 2. Persistent memory (built-in + providers)
- **What:** Built-in `MEMORY.md` / `USER.md` always active; optional external providers (honcho, mem0, hindsight, byterover, …). Memory is injected every session — this is the single biggest cure for "it forgets my project."
- **When:** after any correction or preference the user states ("use bun, not npm") — tell Hermes to remember it and it persists.
- **Command:** `hermes memory setup|status|off|reset` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/memory [V-DOC]
### 3. Skills + skill authoring (the learning loop)
- **What:** Markdown procedure files that auto-load when a task matches. The differentiator: Hermes can WRITE its own skills after learning a workflow (self-improving), and the curator maintains them (usage tracking, archiving, backups).
- **When:** any workflow done twice — say "save this as a skill."
- **Commands:** `hermes skills list|search|install|browse|config|check|update` [V-LIVE]; in-session `/skill <name>`, `/reload-skills` [V-DOC]; authoring tool in-session is `skill_manage` (agent-side; writer should describe it as "ask your Hermes to save the procedure as a skill").
- **Docs:** https://hermes-agent.nousresearch.com/docs/reference/skills-catalog [V-DOC]; curator: https://hermes-agent.nousresearch.com/docs/user-guide/features/curator [V-DOC]
### 4. Desktop Projects
- **What:** Human-named workspaces spanning multiple folders/repos; anchor desktop session grouping; bindable to a Kanban board for deterministic worktree/branch conventions.
- **When:** Scot's multi-repo workflows (portfolio ops). `hermes project create <name>` then `add-folder`.
- **Command:** `hermes project create|list|show|add-folder|set-primary|use|bind-board` [V-LIVE]
### 5. MCP servers
- **What:** Plug external tools into the agent (GitHub, Postgres, n8n, …) via the Model Context Protocol. Also runs in reverse: `hermes mcp serve` exposes Hermes conversations to other agents.
- **Command:** `hermes mcp add|list|test|configure|picker|catalog|install|serve` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp [V-DOC]
### 6. Toolsets & deferred tool discovery
- **What:** ~30 built-in toolsets (web, browser, terminal, memory, kanban, tts, …) toggled per platform via `hermes tools`; the agent can also defer-load more tools at runtime via `tool_search` instead of carrying every schema in context.
- **When:** trim toolsets for focus/cost, or enable `browser` for web work.
- **Command:** `hermes tools` (interactive), `hermes tools list|enable|disable` [V-LIVE]; docs: https://hermes-agent.nousresearch.com/docs/reference/tools-reference [V-DOC]
### 7. Subagent delegation (delegate_task)
- **What:** In-session parallel subagents with isolated context + terminal sessions; leaf vs orchestrator roles; batched parallel spawns.
- **When:** research fan-out, parallel code review, anything that would flood the main context.
- **Command:** agent-side tool (no CLI). In-session: ask Hermes to "delegate X to subagents." Docs: /docs/user-guide/features (delegation section) [V-DOC]
### 8. Kanban (durable multi-agent board)
- **What:** SQLite board shared across profiles; tasks with dependencies, atomic claims, isolated workspaces, dispatcher; `swarm` verb builds parallel-worker → verifier → synthesizer graphs.
- **When:** recurring multi-step operations, handoffs between specialist profiles, long-running campaigns that must survive restarts.
- **Command:** `hermes kanban create|list|show|swarm|link|complete|watch|stats|dispatch` (40+ verbs) [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/kanban [V-DOC]
### 9. Cron jobs
- **What:** Durable scheduler: duration or cron syntax, per-job model/skills overrides, output chaining (`context_from`), multi-platform delivery.
- **When:** daily reports, monitoring with alerts, weekly reviews.
- **Command:** `hermes cron list|create|edit|pause|resume|run|remove|doctor|status` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/cron [V-DOC]
### 10. Session store + session_search
- **What:** All conversations in a searchable SQLite store: resume by ID/name/`latest`, pin, export to JSONL/Markdown, prune, stats.
- **When:** "what did we decide last week" — the agent can search past sessions; user can browse them.
- **Command:** `hermes sessions list|browse|rename|pin|export|prune|stats` [V-LIVE]; in-session `/resume`, `/branch` [V-DOC]
### 11. Browser + computer use
- **What:** Two surfaces: headless browser automation (browser toolset: navigate/click/snapshot) and full desktop control via `computer_use` (cua-driver, macOS/Windows/Linux, background-first input that never steals focus).
- **When:** web research → headless browser; native apps (Excel, Figma, native chat) → computer use.
- **Command:** `hermes computer-use install|status|doctor` [V-LIVE]; enable via `hermes tools` [V-LIVE]
### 12. Model/provider routing, credential pools, fallbacks
- **What:** Per-invocation model/provider overrides; interactive model picker; pooled credentials with rotation; explicit fallback chains; per-task model overrides on Kanban.
- **Command:** `hermes model` [V-LIVE], `hermes fallback list|add|remove` [V-LIVE], `hermes auth add|list|priority|reset` [V-LIVE]; per-run flags `-m`, `--provider`, `--reasoning` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/integrations/providers [V-DOC]
- **Note for Scot (OpenRouter + small fast models):** `--reasoning high` on hard tasks; `hermes fallback add` so a failed call rolls to a second model instead of erroring.
### 13. Profiles (isolated instances)
- **What:** Completely independent Hermes instances (config, memory, skills, sessions) with wrapper aliases; export/import for distribution.
- **When:** separate work/persona contexts, or one profile per client.
- **Command:** `hermes profile list|create|use|alias|export|import` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/profiles [V-DOC]
### 14. Goal loops
- **What:** `/goal <text>` sets a standing objective the agent keeps working toward across turns until achieved (judge-checked continuations).
- **When:** "keep the CI green until it passes," "keep researching until you have 5 verified sources."
- **Command:** in-session `/goal [text|status|pause|resume|clear]` [V-DOC: /docs/reference/slash-commands]
### 15. Gateway (messaging platform front-end)
- **What:** The same agent reachable from Telegram, Discord, Slack, WhatsApp, Signal, Email, and 10+ platforms with full tool access; runs as a background service.
- **Command:** `hermes gateway run|install|start|status|setup` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ [V-DOC]
### 16. Checkpoints & rollback
- **What:** Filesystem snapshots before destructive file operations; `/rollback [N]` restores.
- **When:** letting the agent loose on important files.
- **Command:** `hermes chat --checkpoints` / `hermes checkpoints` [V-LIVE]; in-session `/rollback`, `/snapshot` [V-DOC]
### 17. Projects↔Kanban binding + worktree mode
- **What:** `hermes project bind-board` ties a board to a project (deterministic worktree + branch per task); `-w/--worktree` runs any session in an isolated git worktree.
- **When:** parallel coding agents that must not collide.
- **Command:** `hermes project bind-board` [V-LIVE]; `hermes -w` [V-LIVE]
### 18. Prompt-size introspection
- **What:** Byte breakdown of system prompt + tool schemas — diagnose why responses feel "dumb" (usually context bloat).
- **Command:** `hermes prompt-size` [V-LIVE]
@@ -1,121 +0,0 @@
# 03 — Command Cheatsheet (every entry verified)
**Verification method:** each VERIFIED-LIVE entry was confirmed against `hermes --help` or `hermes <cmd> --help` on Hermes Agent v0.21.1 (2026.9.7), reference install (Syslog kagentz), 2026-09-11. Raw output: `cli-help-dump.txt`. DOC-ONLY entries come from the official docs (URL given). Nothing is invented.
## (a) CLI — `hermes ...`
### Setup & health
| Command | What it does | Tag |
|---|---|---|
| `hermes setup` | Interactive setup wizard | VERIFIED-LIVE |
| `hermes doctor [--fix] [--live]` | Diagnose config/deps; `--fix` auto-repairs | VERIFIED-LIVE |
| `hermes status [--all] [--deep]` | Component status | VERIFIED-LIVE |
| `hermes config show/edit/get/set/unset/path/env-path/check/migrate` | View/edit config | VERIFIED-LIVE |
| `hermes update` | Update Hermes to latest | VERIFIED-LIVE |
### The Claude Code bridge (highest value for Scot)
| Command | What it does | Tag |
|---|---|---|
| `hermes import-agent claude-code [--dry-run] [--overwrite] [--yes]` | One-command import of a Claude Code setup: maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, memories into Hermes equivalents. Never imports API keys. | VERIFIED-LIVE |
| `hermes import-agent codex` | Same for Codex CLI setups | VERIFIED-LIVE |
| `hermes sessions import` | Import a Claude Code or Codex CLI **session/conversation** into Hermes | VERIFIED-LIVE (subcommand listed in `hermes sessions --help`) |
| `hermes skills trust` | Trust a repo so its project-local skills (`./.hermes/skills`) load — the Hermes analog of `.claude/skills/` | VERIFIED-LIVE |
### Daily driving
| Command | What it does | Tag |
|---|---|---|
| `hermes` / `hermes chat` | Interactive session | VERIFIED-LIVE |
| `hermes -c [NAME]` / `hermes --resume <id\|latest>` | Resume by name or ID | VERIFIED-LIVE |
| `hermes --in DIR --resume latest` | Resume the latest session for a directory | VERIFIED-LIVE |
| `hermes -z "PROMPT"` | One-shot: prints ONLY the final answer (scripting/CI); tools, memory, and AGENTS.md still load | VERIFIED-LIVE |
| `hermes chat -q "PROMPT"` | Single-query mode | VERIFIED-LIVE |
| `hermes -m MODEL --provider PROVIDER --reasoning LEVEL` | Per-run model/provider/reasoning overrides (`none…ultra`) | VERIFIED-LIVE |
| `hermes -s SKILL1,SKILL2` | Preload specific skills for the session | VERIFIED-LIVE |
| `hermes -t TOOLSETS` | Restrict toolsets for this run | VERIFIED-LIVE |
| `hermes -w` | Isolated git worktree session (parallel agents on one repo) | VERIFIED-LIVE |
| `hermes chat --checkpoints` | Enable filesystem checkpoints (`/rollback` to restore) | VERIFIED-LIVE |
| `hermes chat --max-turns N` / `--run-budget SECONDS` | Cap loop iterations / wall-clock budget | VERIFIED-LIVE |
### Context & memory management
| Command | What it does | Tag |
|---|---|---|
| `hermes memory setup/status/off/reset` | External memory provider management (built-in MEMORY.md/USER.md always active) | VERIFIED-LIVE |
| `hermes sessions list/browse/rename/pin/export/prune/stats` | Session store management | VERIFIED-LIVE |
| `hermes skills list/search/install/inspect/browse/config/check/update` | Skill management | VERIFIED-LIVE |
| `hermes skills trust/untrust` | Repo-local skill trust | VERIFIED-LIVE |
| `hermes curator status/run/pause/pin/...` | Background skill maintenance (auto-archive, backups) | VERIFIED-LIVE |
| `hermes prompt-size` | Byte breakdown of system prompt + tool schemas (context-bloat diagnosis) | VERIFIED-LIVE |
| `hermes insights [--days N]` | Usage analytics | VERIFIED-LIVE |
### Tools, MCP, integrations
| Command | What it does | Tag |
|---|---|---|
| `hermes tools` (interactive) / `list/enable/disable` | Per-platform toolset toggles; MCP tools as `server:tool` | VERIFIED-LIVE |
| `hermes mcp add/remove/list/test/configure/picker/catalog/install` | MCP server management (incl. one-click catalog installs) | VERIFIED-LIVE |
| `hermes mcp serve` | Run Hermes AS an MCP server for other agents | VERIFIED-LIVE |
| `hermes computer-use install/status/doctor` | Desktop-control backend (cua-driver) | VERIFIED-LIVE |
| `hermes gateway run/install/start/status/setup` | Messaging gateway (Telegram, Discord, Slack, WhatsApp, …) | VERIFIED-LIVE |
| `hermes send` | Send a message to a configured platform (scripts/cron/CI) | VERIFIED-LIVE |
### Automation & multi-agent
| Command | What it does | Tag |
|---|---|---|
| `hermes cron list/create/edit/pause/resume/run/remove/doctor` | Scheduled jobs (durable, multi-platform delivery) | VERIFIED-LIVE |
| `hermes cron notepad` | Durable per-job key-value notepad across runs | VERIFIED-LIVE |
| `hermes kanban create/list/show/link/complete/swarm/...` | Durable multi-profile task board (40+ verbs) | VERIFIED-LIVE |
| `hermes kanban swarm` | Generate a parallel-workers → verifier → synthesizer task graph | VERIFIED-LIVE |
| `hermes project create/list/add-folder/bind-board` | Named multi-folder workspaces (desktop Projects) | VERIFIED-LIVE |
| `hermes profile list/create/use/alias/export/import` | Isolated Hermes instances | VERIFIED-LIVE |
| `hermes auth add/list/priority/reset` | Pooled credentials per provider (rotation) | VERIFIED-LIVE |
| `hermes fallback list/add/remove` | Fallback model chain (auto-rollover on failure) | VERIFIED-LIVE |
| `hermes model` | Interactive model/provider picker | VERIFIED-LIVE |
| `hermes -yolo` | Bypass command approval prompts (use with care) | VERIFIED-LIVE |
| `hermes pause` / `hermes resume` | Emergency stop / lift (pauses cron, kanban dispatch, gateway turns) | VERIFIED-LIVE |
## (b) In-session slash commands
Source: official slash-commands reference https://hermes-agent.nousresearch.com/docs/reference/slash-commands (DOC-ONLY — slash commands run inside a chat session and were not exercised from this headless research run; the CLI subcommands they map to were verified live). DOC-ONLY.
### Context & session
| Command | What it does |
|---|---|
| `/help` | List all commands (authoritative in your version) |
| `/new` (`/reset`) | Fresh session |
| `/resume [name]` | Resume a named/recent session |
| `/branch` (`/fork`) | Branch the current session |
| `/compress` | Manually compress context (auto-compression also exists) |
| `/undo` | Remove last exchange |
| `/retry` | Resend last message |
| `/title [name]` | Name the session |
| `/save` | Save conversation to file |
| `/history` | Show conversation history |
### Power surfaces
| Command | What it does |
|---|---|
| `/skill <name>` | Load a skill into the session |
| `/skills` | Search/install skills |
| `/reload-skills` | Re-scan skill directory |
| `/tools` / `/toolsets` | Manage tools |
| `/goal [text]` | Set a standing goal the agent works toward across turns (`/goal status/pause/clear` to manage) |
| `/background <prompt>` | Run a prompt in the background |
| `/queue <prompt>` | Queue a prompt for the next turn |
| `/steer <prompt>` | Inject a course-correction after the next tool call without interrupting |
| `/agents` | Show active agents and running tasks |
| `/cron` | Manage cron jobs in-session |
| `/kanban` | Multi-profile collaboration board in-session |
| `/model [name]` | Show/change model mid-session |
| `/reasoning [level]` | Set reasoning effort |
| `/voice [on\|off\|tts]` | Voice mode |
| `/rollback [N]` | Restore filesystem checkpoint (needs `--checkpoints`) |
| `/usage` | Token usage |
| `/insights [days]` | Usage analytics |
| `/platforms` | Gateway platform status |
| `/compact`-equivalent note | Hermes compresses automatically near the context limit; no manual threshold watch needed like Claude Code's `/context` |
### The "work alongside Claude Code" shortlist
1. `hermes import-agent claude-code --dry-run` → migrate the setup (VERIFIED-LIVE)
2. `hermes sessions import` → bring the conversation history over (VERIFIED-LIVE)
3. `hermes skills trust` → load repo-local skills like `.claude/skills/` (VERIFIED-LIVE)
4. `hermes -c` / `hermes --in <repo> --resume latest` → per-directory session continuity (VERIFIED-LIVE)
5. `hermes mcp serve` → expose Hermes to Claude Code as an MCP server (VERIFIED-LIVE) — the reverse direction Claude Code can't do
@@ -1,76 +0,0 @@
# 04 — The Claude Code Bridge: Running Hermes WITH Claude Code
**Prepared:** 2026-09-11. Scot already runs both tools. This file documents the proven integration patterns, citing the installed skills on this host (paths under `/home/hermes/.hermes/skills/`) and official docs.
## Pattern 0 — Import (do this first)
`hermes import-agent claude-code` [VERIFIED-LIVE] maps CLAUDE.md/AGENTS.md instructions,
permission allowlists, MCP servers, skills, and memories into Hermes equivalents. It always
shows a preview, never imports credentials. `hermes sessions import` [VERIFIED-LIVE] pulls
in old Claude Code conversations. After import, Hermes "knows" the projects — the context
gap disappears on day one.
## Pattern 1 — Hermes as orchestrator, Claude Code as worker
Source: installed skill **`autonomous-ai-agents/delegate-coding-agent`** (v1.0.0) + its
reference `references/claude-code.md` (v2.2.1) [V-FILE]. The skill is an official Hermes
skill authored for exactly this.
Two orchestration modes (verbatim from the skill):
- **Print mode (preferred):** `claude -p '<task>' --allowedTools 'Read,Edit' --max-turns 10` —
one-shot, no dialogs, structured JSON output with `session_id`, `num_turns`,
`total_cost_usd`. Ask Hermes: *"delegate this coding task to Claude Code in print mode."*
- **Interactive PTY via tmux:** multi-turn sessions — Hermes starts `tmux new-session`,
sends prompts with `send-keys`, monitors with `capture-pane`. For iterative
refactor → review → fix cycles.
Cross-agent review loop (also from the skill):
```
git diff main...feature | claude -p 'Review this diff for bugs and security issues.' --max-turns 1
```
Hermes runs this, reads the output, and fixes findings itself — Claude Code becomes a
reviewer Hermes coordinates.
Safety rails the skill prescribes: explicit `workdir`, clean git status before launch,
narrow task prompts, `git diff` review, targeted tests before committing.
## Pattern 2 — Parallel workstreams + neutral merge reconciliation
Source: installed skill **`autonomous-ai-agents/merge-reconciler`** [V-FILE].
When Hermes and Claude Code (or two Hermes workers) both edit the same repo and collide:
- Do NOT let either agent resolve the conflict — both are biased toward their own side.
- Spawn a **neutral third agent** with the merge-reconciler skill; it classifies every
conflicted hunk (disjoint-intent / same-question-different-answer / superseded), resolves
under an impartiality contract (touch only conflict markers, surface every design call),
verifies with build/tests, and hands back a summary naming every hunk decision.
- Kanban-native shape: a reconciliation card assigned to a **third profile** with both
workers' cards as parents — parent links carry both sides' completion summaries into the
reconciler's context automatically.
## Pattern 3 — Hermes as MCP server (Claude Code gets Hermes tools)
`hermes mcp serve` [VERIFIED-LIVE] runs Hermes as an MCP server exposing its conversations
and capabilities. Claude Code supports MCP clients (`claude mcp add`), so Claude Code can
consume Hermes as a tool provider — persistent memory, skills, cron — the surfaces Claude
Code lacks. This is the reverse-bridge only Hermes can offer. Docs:
https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp
## Pattern 4 — Import legacy sessions for continuity
`hermes sessions import` [VERIFIED-LIVE] imports a Claude Code session into the Hermes
store; from then on `hermes --resume <id>` / `hermes sessions browse` treat it as native
history. Use when mid-project: the new agent picks up exactly where Claude Code left off.
## Pattern 5 — Desktop GUI automation either side can use
Source: installed skill **`autonomous-ai-agents/computer-use`** (v2.0.0) [V-FILE].
`hermes computer-use install` sets up cua-driver; the `computer_use` toolset drives native
desktop apps background-first (never steals focus/cursor), any-model, cross-platform.
Relevant to the bridge because Claude Code has no desktop automation — if a task needs
Figma/Excel/native apps, that part routes to Hermes while the code routes to Claude Code.
Cmd: `hermes computer-use doctor` for health checks.
## Pattern 6 — The import-agent philosophy in one line
Claude Code holds repo context in `CLAUDE.md`; Hermes holds it in `AGENTS.md` + memory +
skills. `hermes import-agent claude-code` translates the first; the learning loop
("save this as a skill") rebuilds the rest automatically the more Scot uses Hermes.
## Reference paths (for the writer)
- `/home/hermes/.hermes/skills/autonomous-ai-agents/delegate-coding-agent/SKILL.md` and `references/claude-code.md`
- `/home/hermes/.hermes/skills/autonomous-ai-agents/merge-reconciler/SKILL.md`
- `/home/hermes/.hermes/skills/autonomous-ai-agents/computer-use/SKILL.md`
- `hermes import-agent --help` raw output in `cli-help-dump.txt` (lines 554-577)
@@ -1,119 +0,0 @@
# 05 — The Video Watch List (every URL verified 2026-09-11)
**Verification method:** each video was found via YouTube search-results scrape (`videoRenderer` metadata), then confirmed with the YouTube oEmbed endpoint (`curl -s "https://www.youtube.com/oembed?url=<URL>&format=json"`) — every entry below returned HTTP 200 with matching title/author (status PASS). Publish dates, durations, and view counts were read from each watch page's metadata. Raw evidence for all 21 entries: `video-verification.json` in this directory. **21/21 PASS, 0 FAIL.**
## Tier 1 — Hermes-specific, start here
### 1. Learn 95% of Hermes Agent in 31 Minutes
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=Ta2wg6xPaY4 | **Duration:** 31:28 | **Published:** 2026-08-09 | **Views:** ~129k
- **oEmbed:** PASS (title/author match)
- **What it demonstrates:** end-to-end Hermes fundamentals — install, sessions, skills, memory, the learning loop. The most complete single-video orientation found.
- **Watch this when you want** the fastest real overview of the whole harness before touching config.
### 2. Hermes Agent Fundamentals In 29 Minutes
- **Channel:** Tina Huang | **URL:** https://www.youtube.com/watch?v=5_N84t1rUU0 | **Duration:** 29:40 | **Published:** 2026-07-20 | **Views:** ~463k (highest-reach Hermes video found)
- **oEmbed:** PASS
- **What it demonstrates:** conceptual grounding — why Hermes' memory/skills loop differs from one-shot coding agents; practical walkthrough.
- **Watch this when you want** to understand *why* Hermes feels different from Claude Code, not just which buttons to press.
### 3. Every Level of Hermes Agent Explained
- **Channel:** Jack Roberts | **URL:** https://www.youtube.com/watch?v=6GtF_uHbGhw | **Duration:** 25:35 | **Published:** 2026-06-17 | **Views:** ~163k
- **oEmbed:** PASS
- **What it demonstrates:** beginner → advanced ladder of features (memory, skills, automation, multi-agent).
- **Watch this when you want** a map of what to learn next after the basics.
### 4. Hermes Agent Full Tutorial INSTALLATION + USECASES
- **Channel:** CodeHead | **URL:** https://www.youtube.com/watch?v=8GjyOQy19so | **Duration:** 7:47 | **Published:** 2026-05-14 | **Views:** ~64k
- **oEmbed:** PASS
- **What it demonstrates:** install through real use-cases, compact.
- **Watch this when you want** a quick install-to-value demo to share with a colleague.
### 5. Hermes Agent Explained In 5 Minutes
- **Channel:** CodeHead | **URL:** https://www.youtube.com/watch?v=9GpWELm3_XI | **Duration:** 4:53 | **Published:** 2026-05-23 | **Views:** ~251k
- **oEmbed:** PASS
- **What it demonstrates:** 5-minute conceptual pitch of the agent and its learning loop.
- **Watch this when you want** the elevator pitch before committing 30 minutes.
## Tier 2 — Hermes-specific deep dives
### 6. 100 Days With Hermes Agent in 21 Minutes
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=sCa3BtpkziQ | **Duration:** 21:19 | **Published:** 2026-06-17 | **Views:** ~58k
- **oEmbed:** PASS
- **What it demonstrates:** long-horizon usage — what memory/skills accumulation actually looks like after months of daily use.
- **Watch this when you want** to see the payoff of the learning loop over time.
### 7. Hermes Agent - Crash Course for Beginners (AI Agent)
- **Channel:** Adrian Twarog | **URL:** https://www.youtube.com/watch?v=4sAmpcSOVEw | **Duration:** 22:19 | **Published:** 2026-07-21 | **Views:** ~46k
- **oEmbed:** PASS
- **What it demonstrates:** beginner crash course from a well-known dev-YouTube creator.
- **Watch this when you want** a second independent explanation of the basics.
### 8. Hermes Agent: The Ultimate Beginner's Guide
- **Channel:** Metics Media | **URL:** https://www.youtube.com/watch?v=CwPUOVUdApE | **Duration:** 37:08 | **Published:** 2026-04-24 | **Views:** ~119k
- **oEmbed:** PASS
- **What it demonstrates:** long-form beginner guide incl. setup and everyday workflows.
- **Watch this when you want** the most thorough single walkthrough in one sitting.
### 9. Hermes Agent Just Killed OpenClaw (Full Tutorial)
- **Channel:** Leon van Zyl | **URL:** https://www.youtube.com/watch?v=jmtpYUOr7_U | **Duration:** 19:59 | **Published:** 2026-04-28 | **Views:** ~16k
- **oEmbed:** PASS
- **What it demonstrates:** full tutorial framing Hermes against the OpenClaw workflow (MCP config, memory, agents).
- **Watch this when you want** a practitioner's feature-by-feature tutorial.
### 10. Hermes Agent vs OpenClaw
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=zwqhemjHq3E | **Duration:** 15:28 | **Published:** 2026-04-20 | **Views:** ~35k
- **oEmbed:** PASS
- **What it demonstrates:** head-to-head comparison of the two agent harnesses.
- **Watch this when you want** the tradeoffs between Hermes and its main alternative.
### 11. Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model
- **Channel:** Tonbi's AI Garage | **URL:** https://www.youtube.com/watch?v=8tpuky8HpXw | **Duration:** 15:08 | **Published:** 2026-03-11 | **Views:** ~18k
- **oEmbed:** PASS
- **What it demonstrates:** Hermes driven by an OpenRouter-served open model (Qwen 3) — directly relevant to an OpenRouter-connected install.
- **Watch this when you want** to see how small open models behave inside Hermes.
### 12. Use This To Make The Hermes Agent Basically Free
- **Channel:** AI LABS | **URL:** https://www.youtube.com/watch?v=5d02TYoOzfE | **Duration:** 13:08 | **Published:** 2026-07-01 | **Views:** ~55k
- **oEmbed:** PASS
- **What it demonstrates:** running Hermes on cheap/free model backends.
- **Watch this when you want** to cut inference costs on an OpenRouter account.
### 13. Hermes Agent The 24/7 Self-Evolving AI Agent!
- **Channel:** WorldofAI | **URL:** https://www.youtube.com/watch?v=cu2fgknmemA | **Duration:** 9:15 | **Published:** 2026-04-07 | **Views:** ~47k
- **oEmbed:** PASS
- **What it demonstrates:** always-on operation: gateway, cron, background automation.
- **Watch this when you want** to turn Hermes from a chat window into a 24/7 assistant.
## Tier 3 — Adjacent (origin/philosophy; not tutorials)
### 14. Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra
- **Channel:** Peter Yang | **URL:** https://www.youtube.com/watch?v=UWjh5Z4s8jY | **Duration:** 46:45 | **Published:** 2026-08-02 | **Views:** ~37k
- **oEmbed:** PASS
- **What it demonstrates:** interview with Hermes' co-founder on the design philosophy (self-improving agents, skills as memory).
- **Watch this when you want** to understand where the product is going.
### 15. Hermes Agent: Agents that grow with you | Episode #357
- **Channel:** Practical AI | **URL:** https://www.youtube.com/watch?v=UTZhvPXnmwA | **Duration:** 47:34 | **Published:** 2026-05-20 | **Views:** ~1.9k
- **oEmbed:** PASS
- **What it demonstrates:** podcast-depth technical discussion of the agent architecture.
- **Watch this when you want** the engineering story behind the learning loop.
### 16. Did Hermes Agent just kill OpenClaw? (full guide)
- **Channel:** Alex Finn | **URL:** https://www.youtube.com/watch?v=tP6yf22OJdI | **Duration:** 13:55 | **Published:** 2026-03-31 | **Views:** ~132k
- **oEmbed:** PASS
- **What it demonstrates:** guide-style comparison/switch content.
- **Watch this when you want** a switcher's guide perspective.
### 17. Hermes Agent: Why Everyone's Ditching OpenClaw in 2026
- **Channel:** Luke Alexander AI | **URL:** https://www.youtube.com/watch?v=1UgXUjT-QtI | **Duration:** 18:03 | **Published:** 2026-03-26 | **Views:** ~14k
- **oEmbed:** PASS — adjacent, comparison content.
- **Watch this when you want** more comparison context.
## Honesty note (required by task spec)
At least 16 of the 17 entries above are directly Hermes-specific (not merely adjacent); the
"fewer than 5 exist" fallback clause was NOT needed — no padding was necessary. Entries
found in search but excluded as thin/low-signal: `iqN6MVzpJTk` (3.1k views, news-style),
`83nWNRKZTCE` (465 views), `6M2tItdARew` (1.2k views), `P2LIFtrRr2U` (promo-style) — all
also verified PASS and kept in `video-verification.json` as spares. No official Nous
Research YouTube tutorial channel was found in searches; the strongest signal of
Hermes-specific video content is the third-party ecosystem above.
@@ -1,577 +0,0 @@
===== hermes chat --help =====
usage: hermes chat [-h] [-q QUERY | --query-file PATH] [--oneshot]
[--image IMAGE] [-m MODEL] [-t TOOLSETS]
[--reasoning LEVEL] [-s SKILLS] [--provider PROVIDER] [-v]
[-Q] [--resume SESSION_ID] [--no-restore-cwd] [--in DIR]
[--continue [SESSION_NAME]] [--create-if-missing]
[--worktree] [--accept-hooks] [--checkpoints]
[--max-turns N] [--run-budget SECONDS] [--yolo]
[--pass-session-id] [--ignore-user-config] [--ignore-rules]
[--safe-mode] [--source SOURCE] [--tui] [--cli] [--dev]
Start an interactive chat session with Hermes Agent
options:
-h, --help show this help message and exit
-q, --query QUERY Query to run. On a real TTY the prompt seeds an
interactive session (submitted literally as the first
turn); combined with --oneshot or -Q, or on a non-TTY,
it answers and exits.
--query-file PATH Read the single query from a file instead of the
command line ('-' reads stdin). Safe for arbitrary
text: nothing is shell-interpreted, so quotes, $(...),
and backticks are preserved verbatim. Mutually
exclusive with -q.
--oneshot With -q/--query-file: answer the query and exit
(legacy single-query behavior) instead of seeding an
interactive session. Implied on non-TTY stdio and by
-Q/--quiet.
--image IMAGE Optional local image path to attach to a single query
-m, --model MODEL Model to use (e.g., anthropic/claude-sonnet-4)
-t, --toolsets TOOLSETS
Comma-separated toolsets to enable
--reasoning LEVEL Reasoning effort for this session: none, minimal, low,
medium, high, xhigh, max, or ultra. Overrides
agent.reasoning_effort for this run only (same levels
as the /reasoning slash command).
-s, --skills SKILLS Preload one or more skills for the session (repeat
flag or comma-separate)
--provider PROVIDER Inference provider (default: auto). Built-in or a
user-defined name from `providers:` in config.yaml.
-v, --verbose Verbose output
-Q, --quiet Quiet mode for programmatic use: suppress banner,
spinner, and tool previews. Only output the final
response and session info.
--resume, -r SESSION_ID
Resume a previous session by ID (shown on exit), or
'latest' for the most recent session
--no-restore-cwd Don't cd into a resumed session's recorded working
directory.
--in DIR Change into DIR before starting or resuming (scopes '
--resume latest' / -c lookups to DIR's workspace).
--continue, -c [SESSION_NAME]
Resume a session by name, or the most recent if no
name given
--create-if-missing With -c/--continue <name>: if no session matches the
name, create a new session with that title and proceed
(instead of failing with a not-found error).
Programmatic callers that want 'send to this named
thread, making it if needed'.
--worktree, -w Run in an isolated git worktree (for parallel agents
on the same repo)
--accept-hooks Auto-approve any unseen shell hooks declared in
config.yaml without a TTY prompt (see also
HERMES_ACCEPT_HOOKS env var and hooks_auto_accept: in
config.yaml).
--checkpoints Enable filesystem checkpoints before destructive file
operations (use /rollback to restore)
--max-turns N Maximum tool-calling iterations per conversation turn
(default: 500, or agent.max_turns in config)
--run-budget SECONDS Optional wall-clock budget in seconds for each
conversation run. At 80% elapsed the agent gets a one-
time wrap-up notice, and implicit provider stale
timeouts are capped to the remaining budget so one
hung call can't consume the run. Unset = off. Also
configurable as agent.run_budget_seconds in
config.yaml. Intended for one-shot/eval invocations
with a hard ceiling.
--yolo Bypass all dangerous command approval prompts (use at
your own risk)
--pass-session-id Include the session ID in the agent's system prompt
--ignore-user-config Ignore ~/.hermes/config.yaml and fall back to built-in
defaults (credentials in .env are still loaded).
Useful for isolated CI runs, reproduction, and third-
party integrations.
--ignore-rules Skip auto-injection of AGENTS.md, SOUL.md,
.cursorrules, memory, and preloaded skills. Combine
with --ignore-user-config for a fully isolated run.
--safe-mode Troubleshooting mode: disable ALL customizations —
user config, AGENTS.md/memory injection, plugins, and
MCP servers (implies --ignore-user-config and
--ignore-rules). Use to isolate whether a problem
comes from your setup or from Hermes itself.
--source SOURCE Session source tag for filtering (default: cli). Use
'tool' for third-party integrations that should not
appear in user session lists.
--tui Launch the modern TUI instead of the classic REPL
--cli Force the classic prompt_toolkit REPL (overrides
display.interface=tui)
--dev With --tui: run TypeScript sources via tsx (skip dist
build)
===== hermes model --help =====
usage: hermes model [-h] [--refresh] [--portal-url PORTAL_URL]
[--inference-url INFERENCE_URL] [--client-id CLIENT_ID]
[--scope SCOPE] [--no-browser] [--timeout TIMEOUT]
[--ca-bundle CA_BUNDLE] [--insecure]
Interactively select your inference provider and default model
options:
-h, --help show this help message and exit
--refresh Wipe the model picker disk cache and re-fetch every
provider's live /v1/models list.
--portal-url PORTAL_URL
Portal base URL for Nous login (default: production
portal)
--inference-url INFERENCE_URL
Inference API base URL for Nous login (default:
production inference API)
--client-id CLIENT_ID
OAuth client id to use for Nous login (default:
hermes-cli)
--scope SCOPE OAuth scope to request for Nous login
--no-browser Do not attempt to open the browser automatically
during Nous login
--timeout TIMEOUT HTTP request timeout in seconds for Nous login
(default: 15)
--ca-bundle CA_BUNDLE
Path to CA bundle PEM file for Nous TLS verification
--insecure Disable TLS verification for Nous login (testing only)
===== hermes config --help =====
usage: hermes config [-h]
{show,edit,get,set,unset,path,env-path,check,migrate} ...
Manage Hermes Agent configuration
positional arguments:
{show,edit,get,set,unset,path,env-path,check,migrate}
show Show current configuration
edit Open config file in editor
get Print a resolved configuration value
set Set a configuration value
unset Remove a configuration value
path Print config file path
env-path Print .env file path
check Check for missing/outdated config
migrate Update config with new options
options:
-h, --help show this help message and exit
===== hermes cron --help =====
usage: hermes cron [-h] [--accept-hooks]
{list,create,add,edit,pause,resume,run,remove,rm,delete,status,runs,history,incidents,notepad,doctor,tick} ...
Manage scheduled tasks
positional arguments:
{list,create,add,edit,pause,resume,run,remove,rm,delete,status,runs,history,incidents,notepad,doctor,tick}
list List scheduled jobs
create (add) Create a scheduled job
edit Edit an existing scheduled job
pause Pause a scheduled job
resume Resume a paused job
run Run a job on the next scheduler tick
remove (rm, delete)
Remove a scheduled job
status Check if cron scheduler is running
runs (history) Show durable execution attempts
incidents List or acknowledge durable cron failure incidents
notepad Read/write a job's durable notepad (persistent KV
across runs)
doctor Check scheduled jobs for common health issues
tick Run due jobs once and exit
options:
-h, --help show this help message and exit
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
(equivalent to HERMES_ACCEPT_HOOKS=1 /
hooks_auto_accept: true).
===== hermes kanban --help =====
usage: hermes kanban [-h] [--board <slug>]
{init,boards,create,swarm,list,ls,show,assign,set-model,reclaim,reassign,diagnostics,diag,link,unlink,claim,comment,attach,attachments,attach-rm,complete,edit,block,schedule,unblock,request-review,request-changes,reopen-review,promote,archive,tail,dispatch,daemon,watch,stats,notify-subscribe,notify-list,notify-unsubscribe,log,runs,heartbeat,assignees,context,specify,decompose,gc,repair} ...
Durable SQLite-backed task board shared across Hermes profiles. Tasks are
claimed atomically, can depend on other tasks, and are executed by a named
profile in an isolated workspace. See https://hermes-
agent.nousresearch.com/docs/user-guide/features/kanban or docs/hermes-
kanban-v1-spec.pdf for the full design.
positional arguments:
{init,boards,create,swarm,list,ls,show,assign,set-model,reclaim,reassign,diagnostics,diag,link,unlink,claim,comment,attach,attachments,attach-rm,complete,edit,block,schedule,unblock,request-review,request-changes,reopen-review,promote,archive,tail,dispatch,daemon,watch,stats,notify-subscribe,notify-list,notify-unsubscribe,log,runs,heartbeat,assignees,context,specify,decompose,gc,repair}
init Create kanban.db if missing (idempotent)
boards Manage kanban boards (one board per project /
workstream)
create Create a new task
swarm Create a Kanban Swarm v1 graph (parallel workers →
verifier → synthesizer)
list (ls) List tasks
show Show a task with comments + events
assign Assign or reassign a task
set-model Set or clear a task's model/provider override (takes
effect on the next dispatch)
reclaim Release an active worker claim on a running task
reassign Reassign a task to a different profile, optionally
reclaiming first
diagnostics (diag) List active diagnostics on the current board
link Add a parent->child dependency
unlink Remove a parent->child dependency
claim Atomically claim a ready task (prints resolved
workspace path)
comment Append a comment
attach Attach a local file to a task
attachments List a task's attachments
attach-rm Delete an attachment by id
complete Mark one or more tasks done
edit Edit recovery fields on an already-completed task
block Mark one or more tasks blocked
schedule Park one or more tasks in Scheduled (waiting on time,
not human input)
unblock Return blocked/scheduled tasks to ready, or todo while
parents remain open
request-review Move a task to 'review' (implementation done, awaiting
review) — NOT a block
request-changes Reviewer verdict: return the active review run to its
implementer
reopen-review Send one or more review tasks back for changes (review
-> ready/todo)
promote Manually move one or more todo/blocked tasks to ready
(recovery path)
archive Archive one or more tasks
tail Follow a task's event stream
dispatch One dispatcher pass: reclaim stale, promote ready,
spawn workers
daemon DEPRECATED — dispatcher now runs in the gateway. Use
`hermes gateway start`.
watch Live-stream task_events to the terminal (Ctrl+C to
exit)
stats Per-status + per-assignee counts + oldest-ready age
notify-subscribe Subscribe a gateway source to a task's terminal events
(used by /kanban subscribe in the gateway adapter)
notify-list List notification subscriptions (optionally for a
single task)
notify-unsubscribe Remove a gateway subscription from a task
log Print the worker log for a task (from <kanban-
root>/kanban/logs/)
runs Show attempt history for a task (one row per run:
profile, outcome, elapsed, summary)
heartbeat Emit a heartbeat event for a running task (worker
liveness signal)
assignees List known profiles + per-profile task counts (union
of ~/.hermes/profiles/ and current assignees on the
board)
context Print the full context a worker sees for a task (title
+ body + parent results + comments).
specify Flesh out a triage-column task into a concrete spec
(title + body) and promote it to todo. Uses the
auxiliary LLM configured under
auxiliary.triage_specifier.
decompose Decompose a triage-column task into a graph of child
tasks routed to specialist profiles by description.
Falls back to specify-style single-task promotion when
the task doesn't benefit from fan-out. Uses
auxiliary.kanban_decomposer.
gc Garbage-collect archived-task workspaces, old events,
and old logs
repair Check kanban.db integrity and auto-repair index-only
corruption
options:
-h, --help show this help message and exit
--board <slug> Board slug to operate on. Defaults to the current
board (set via `hermes kanban boards switch <slug>` or
the HERMES_KANBAN_BOARD env var). Use `hermes kanban
boards list` to see all boards.
===== hermes skills --help =====
usage: hermes skills [-h]
{trust,untrust,browse,search,install,inspect,list,check,update,audit,uninstall,reset,list-modified,diff,opt-out,opt-in,repair-official,publish,snapshot,tap,config} ...
Search, install, inspect, audit, configure, and manage skills from skills.sh,
well-known agent skill endpoints, GitHub, ClawHub, and other registries.
positional arguments:
{trust,untrust,browse,search,install,inspect,list,check,update,audit,uninstall,reset,list-modified,diff,opt-out,opt-in,repair-official,publish,snapshot,tap,config}
trust Trust a project so its repo-local skills
(./.hermes/skills, ./.agents/skills) load
untrust Revoke project-skill trust for a repo
browse Browse all available skills (paginated)
search Search skill registries
install Install a skill
inspect Preview a skill without installing
list List installed skills
check Check installed hub skills for updates
update Update installed hub skills
audit Re-scan installed hub skills
uninstall Remove a hub-installed skill
reset Reset a bundled skill — clears 'user-modified'
tracking so updates work again
list-modified List bundled skills you've edited (which `hermes
update` keeps)
diff Show how your copy of a bundled skill differs from the
stock version
opt-out Stop bundled skills from being seeded into this
profile
opt-in Re-enable bundled-skill seeding (undo opt-out)
repair-official Backfill or restore official optional skills from repo
source
publish Publish a skill to a registry
snapshot Export/import skill configurations
tap Manage skill sources
config Interactive skill configuration — enable/disable
individual skills
options:
-h, --help show this help message and exit
===== hermes sessions --help =====
usage: hermes sessions [-h]
{list,export,delete,prune,archive,optimize,clean-markers,optimize-storage,repair,repair-routing,recover,stats,rename,pin,unpin,pinned,retitle-skills,browse,import} ...
View and manage the SQLite session store
positional arguments:
{list,export,delete,prune,archive,optimize,clean-markers,optimize-storage,repair,repair-routing,recover,stats,rename,pin,unpin,pinned,retitle-skills,browse,import}
list List recent sessions
export Export sessions to JSONL, Markdown, or QMD
delete Delete a specific session
prune Delete old sessions (filterable by time window,
source, title, ...)
archive Bulk-archive (soft-hide) sessions matching filters —
no deletion
optimize Reclaim disk space: merge FTS5 segments + VACUUM (no
data change)
clean-markers Permanently clear stale tool-call marker content left
by sessions from before #78148
optimize-storage Migrate the search index to the compact v23 layout
(reclaims disk on large DBs)
repair Repair a malformed state.db schema so hidden sessions
reappear
repair-routing Re-stamp gateway sessions that lost their routing
identity
recover Rebuild canonical session data into a separate clean
database
stats Show session store statistics
rename Set or change a session's title
pin Pin session(s) — durable keep flag, exempt from auto-
archive
unpin Remove the pin (durable keep flag) from session(s)
pinned List pinned sessions
retitle-skills Re-title sessions whose auto-title came from a
/skill's own text
browse Interactive session picker — browse, search, and
resume sessions
import Import a Claude Code or Codex CLI session into Hermes
options:
-h, --help show this help message and exit
===== hermes mcp --help =====
usage: hermes mcp [-h] [--accept-hooks]
{serve,add,remove,rm,list,ls,test,configure,config,login,reauth,picker,catalog,install} ...
Manage MCP server connections and run Hermes as an MCP server. MCP servers
provide additional tools via the Model Context Protocol. Use 'hermes mcp add'
to connect to a new server, or 'hermes mcp serve' to expose Hermes
conversations over MCP.
positional arguments:
{serve,add,remove,rm,list,ls,test,configure,config,login,reauth,picker,catalog,install}
serve Run Hermes as an MCP server (expose conversations to
other agents)
add Add an MCP server (discovery-first install)
remove (rm) Remove an MCP server
list (ls) List configured MCP servers
test Test MCP server connection
configure (config) Toggle tool selection
login Force re-authentication for an OAuth-based MCP server
reauth Re-authenticate one OAuth MCP server, or all of them
(--all)
picker Interactive catalog picker (also the default for
`hermes mcp`)
catalog List Nous-approved MCPs available for one-click
install
install Install a catalog MCP by name (e.g. `hermes mcp
install n8n`)
options:
-h, --help show this help message and exit
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
(equivalent to HERMES_ACCEPT_HOOKS=1 /
hooks_auto_accept: true).
===== hermes profile --help =====
usage: hermes profile [-h]
{list,use,create,delete,describe,show,alias,rename,export,import,install,update,info} ...
positional arguments:
{list,use,create,delete,describe,show,alias,rename,export,import,install,update,info}
list List all profiles
use Set sticky default profile
create Create a new profile
delete Delete a profile
describe Read or set a profile's description (used by the
kanban orchestrator)
show Show profile details
alias Manage wrapper scripts
rename Rename a profile ('default': sets a display name; id
unchanged)
export Export a profile to archive
import Import a profile from archive
install Install a profile distribution from a git URL or local
directory
update Re-pull a distribution and apply updates (user data
preserved)
info Show a profile's distribution manifest (version,
requirements, source)
options:
-h, --help show this help message and exit
===== hermes memory --help =====
usage: hermes memory [-h] {setup,status,off,reset} ...
Set up and manage external memory provider plugins. Available providers:
honcho, openviking, mem0, hindsight, holographic, retaindb, byterover. Only
one external provider can be active at a time. Built-in memory
(MEMORY.md/USER.md) is always active.
positional arguments:
{setup,status,off,reset}
setup Interactive provider selection and configuration
status Show current memory provider config
off Disable external provider (built-in only)
reset Erase all built-in memory (MEMORY.md and USER.md)
options:
-h, --help show this help message and exit
===== hermes tools --help =====
usage: hermes tools [-h] [--summary] {list,disable,enable,post-setup} ...
Enable, disable, or list tools for CLI, Telegram, Discord, etc. Built-in
toolsets use plain names (e.g. web, memory). MCP tools use server:tool
notation (e.g. github:create_issue). Run 'hermes tools' with no subcommand for
the interactive configuration UI.
positional arguments:
{list,disable,enable,post-setup}
list Show all tools and their enabled/disabled status
disable Disable toolsets or MCP tools
enable Enable toolsets or MCP tools
post-setup Run a provider's post-setup install hook
(npm/pip/binary)
options:
-h, --help show this help message and exit
--summary Print a summary of enabled tools per platform and exit
===== hermes project --help =====
usage: hermes project [-h]
{create,list,ls,show,add-folder,remove-folder,rename,set-primary,use,archive,restore,bind-board} ...
Projects are human-named workspaces that can span multiple folders / repos.
They anchor desktop session grouping and, when bound to a kanban board, give
tasks a deterministic worktree + branch convention. State is per-profile.
positional arguments:
{create,list,ls,show,add-folder,remove-folder,rename,set-primary,use,archive,restore,bind-board}
create Create a new project
list (ls) List projects
show Show a project's details
add-folder Add a folder to a project
remove-folder Remove a folder from a project
rename Rename a project
set-primary Set the primary folder
use Set the active project
archive Archive a project
restore Restore an archived project
bind-board Bind a kanban board to a project
options:
-h, --help show this help message and exit
===== hermes gateway --help =====
usage: hermes gateway [-h] [--accept-hooks]
{run,start,stop,restart,status,install,uninstall,list,setup,migrate-legacy,enroll} ...
Manage the messaging gateway (Telegram, Discord, WhatsApp, Weixin, and more)
positional arguments:
{run,start,stop,restart,status,install,uninstall,list,setup,migrate-legacy,enroll}
run Run gateway in foreground (recommended for WSL,
Docker, Termux)
start Start the installed systemd/launchd background service
stop Stop gateway service
restart Restart gateway service
status Show gateway status
install Install gateway as a systemd/launchd background
service
uninstall Uninstall gateway service
list List all profiles and their gateway status
setup Configure messaging platforms
migrate-legacy Remove legacy hermes.service units from pre-rename
installs
enroll Enroll this gateway with a relay connector (writes
relay auth creds to .env)
options:
-h, --help show this help message and exit
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
(equivalent to HERMES_ACCEPT_HOOKS=1 /
hooks_auto_accept: true).
===== hermes computer-use --help =====
usage: hermes computer-use [-h] {install,status,doctor,permissions} ...
Install or check the cua-driver binary used by the `computer_use` toolset.
Supported on macOS, Windows, and Linux. Use `hermes computer-use install` to
fetch and run the upstream cua-driver installer. This is equivalent to the
post-setup hook that `hermes tools` runs when you first enable the Computer
Use toolset, and is a stable target for re-running the install if it didn't
fire (e.g. when toggling the toolset on a returning-user setup). Use `hermes
computer-use doctor` to run cua-driver's `health_report` MCP tool and surface
its check matrix (TCC, bundle identity, version, platform support, ...) in
human-readable form.
positional arguments:
{install,status,doctor,permissions}
install Install or repair the cua-driver binary
(macOS/Windows/Linux)
status Print whether cua-driver is installed and on PATH
doctor Run cua-driver `health_report` and surface the check
matrix
permissions Check or grant macOS Accessibility + Screen Recording
(macOS)
options:
-h, --help show this help message and exit
===== hermes doctor --help =====
usage: hermes doctor [-h] [--fix] [--live] [--ack ADVISORY_ID]
Diagnose issues with Hermes Agent setup
options:
-h, --help show this help message and exit
--fix Attempt to fix issues automatically
--live Opt-in: run one bounded, read-only real-call health probe
per configured tool backend
(Firecrawl/FAL/browser/MCP/TTS/STT) after the static
checks. Makes real network calls.
--ack ADVISORY_ID Acknowledge a security advisory by ID and exit. After
ack, the advisory will no longer trigger startup banners.
Run `hermes doctor` first to see active advisories and
their IDs.
===== hermes status --help =====
usage: hermes status [-h] [--all] [--deep]
Display status of Hermes Agent components
options:
-h, --help show this help message and exit
--all Show all details (redacted for sharing)
--deep Run deep checks (may take longer)
===== hermes import-agent --help =====
usage: hermes import-agent [-h] [--source SOURCE] [--dry-run] [--overwrite]
[--yes]
[{claude-code,codex}]
One-command import of another coding agent's setup into Hermes. Maps
CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills,
and memories into their Hermes equivalents. Always shows a preview before
making changes. API keys and credentials are never imported — run 'hermes
setup' for those.
positional arguments:
{claude-code,codex} Which agent to import from (default: auto-detect
~/.claude or ~/.codex)
options:
-h, --help show this help message and exit
--source SOURCE Path to the agent's config directory (default:
~/.claude or ~/.codex)
--dry-run Preview only — stop after showing what would be
imported
--overwrite Overwrite existing Hermes items on name conflicts
(default: skip)
--yes, -y Skip confirmation prompts
@@ -1,33 +0,0 @@
import json, subprocess, re
data = json.load(open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json"))
have = {r["videoId"] for r in data}
extra = []
for v in ["8GjyOQy19so", "zwqhemjHq3E"]:
if v in have:
continue
url = f"https://www.youtube.com/watch?v={v}"
oe = subprocess.run(["curl", "-s", f"https://www.youtube.com/oembed?url={url}&format=json"],
capture_output=True, text=True, timeout=30)
try:
oej = json.loads(oe.stdout)
o = {"status": "PASS", "title": oej.get("title"), "author": oej.get("author_name")}
except Exception:
o = {"status": "FAIL", "raw": oe.stdout[:200]}
wp = subprocess.run(["curl", "-s", "-L", url,
"-H", "User-Agent: Mozilla/5.0 (Windows NT 10.0) Chrome/124.0",
"-H", "Accept-Language: en-US"], capture_output=True, text=True, timeout=30)
html = wp.stdout
pub = re.search(r'"publishDate":"([\d\-T:Z]+)"', html)
dur = re.search(r'"lengthSeconds":"(\d+)"', html)
views = re.search(r'"viewCount":"(\d+)"', html)
rec = {"videoId": v, "oembed": o,
"publishDate": pub.group(1) if pub else None,
"lengthSeconds": int(dur.group(1)) if dur else None,
"views": int(views.group(1)) if views else None}
extra.append(rec)
print(json.dumps(rec))
data.extend(extra)
json.dump(data, open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json", "w"), indent=2)
print("total videos in ledger:", len(data))
@@ -1,38 +0,0 @@
import json, subprocess, re, sys
vids = ["9GpWELm3_XI","iqN6MVzpJTk","Ta2wg6xPaY4","8tpuky8HpXw","tP6yf22OJdI",
"UWjh5Z4s8jY","5d02TYoOzfE","83nWNRKZTCE","6M2tItdARew","5_N84t1rUU0",
"UTZhvPXnmwA","1UgXUjT-QtI","6GtF_uHbGhw","CwPUOVUdApE","sCa3BtpkziQ",
"jmtpYUOr7_U","cu2fgknmemA","4sAmpcSOVEw","P2LIFtrRr2U"]
results = []
for v in vids:
url = f"https://www.youtube.com/watch?v={v}"
oe = subprocess.run(["curl","-s",f"https://www.youtube.com/oembed?url={url}&format=json"],
capture_output=True, text=True, timeout=30)
try:
oej = json.loads(oe.stdout)
oembed = {"status":"PASS","title":oej.get("title"),"author":oej.get("author_name")}
except Exception:
oembed = {"status":"FAIL","raw":oe.stdout[:200]}
# watch page for date + duration
wp = subprocess.run(["curl","-s","-L",url,"-H","User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/124.0",
"-H","Accept-Language: en-US,en;q=0.9"], capture_output=True, text=True, timeout=30)
html = wp.stdout
pub = re.search(r'"publishDate":"([\d\-T:Z]+)"', html)
upd = re.search(r'"uploadDate":"([\d\-T:Z]+)"', html)
dur = re.search(r'"lengthSeconds":"(\d+)"', html)
views = re.search(r'"viewCount":"(\d+)"', html)
results.append({
"videoId": v, "oembed": oembed,
"publishDate": pub.group(1) if pub else None,
"uploadDate": upd.group(1) if upd else None,
"lengthSeconds": int(dur.group(1)) if dur else None,
"views": int(views.group(1)) if views else None,
"watchpage_bytes": len(html),
})
print(json.dumps(results[-1]))
with open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json","w") as f:
json.dump(results, f, indent=2)
print("saved video-verification.json")
@@ -1,271 +0,0 @@
[
{
"videoId": "9GpWELm3_XI",
"oembed": {
"status": "PASS",
"title": "Hermes Agent Explained In 5 Minutes",
"author": "CodeHead"
},
"publishDate": "2026-05-23T08:00:27-07:00",
"uploadDate": "2026-05-23T08:00:27-07:00",
"lengthSeconds": 293,
"views": 251275,
"watchpage_bytes": 1419048
},
{
"videoId": "iqN6MVzpJTk",
"oembed": {
"status": "PASS",
"title": "Hermes Agent by Nous Research: The Open-Source Agent Model Everyone Is Switching To",
"author": "Praveen Govindaraj"
},
"publishDate": "2026-02-26T21:56:54-08:00",
"uploadDate": "2026-02-26T21:56:54-08:00",
"lengthSeconds": 193,
"views": 3140,
"watchpage_bytes": 1305120
},
{
"videoId": "Ta2wg6xPaY4",
"oembed": {
"status": "PASS",
"title": "Learn 95% of Hermes Agent in 31 Minutes",
"author": "Sharbel A."
},
"publishDate": "2026-08-09T07:00:19-07:00",
"uploadDate": "2026-08-09T07:00:19-07:00",
"lengthSeconds": 1888,
"views": 129313,
"watchpage_bytes": 1468130
},
{
"videoId": "8tpuky8HpXw",
"oembed": {
"status": "PASS",
"title": "Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model",
"author": "Tonbi's AI Garage"
},
"publishDate": "2026-03-11T07:00:14-07:00",
"uploadDate": "2026-03-11T07:00:14-07:00",
"lengthSeconds": 907,
"views": 18167,
"watchpage_bytes": 1326925
},
{
"videoId": "tP6yf22OJdI",
"oembed": {
"status": "PASS",
"title": "Did Hermes Agent just kill OpenClaw? (full guide)",
"author": "Alex Finn"
},
"publishDate": "2026-03-31T06:15:10-07:00",
"uploadDate": "2026-03-31T06:15:10-07:00",
"lengthSeconds": 835,
"views": 132128,
"watchpage_bytes": 1390199
},
{
"videoId": "UWjh5Z4s8jY",
"oembed": {
"status": "PASS",
"title": "Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra",
"author": "Peter Yang"
},
"publishDate": "2026-08-02T06:00:12-07:00",
"uploadDate": "2026-08-02T06:00:12-07:00",
"lengthSeconds": 2804,
"views": 36613,
"watchpage_bytes": 1386664
},
{
"videoId": "5d02TYoOzfE",
"oembed": {
"status": "PASS",
"title": "Use This To Make The Hermes Agent Basically Free",
"author": "AI LABS"
},
"publishDate": "2026-07-01T07:00:26-07:00",
"uploadDate": "2026-07-01T07:00:26-07:00",
"lengthSeconds": 788,
"views": 54858,
"watchpage_bytes": 1412833
},
{
"videoId": "83nWNRKZTCE",
"oembed": {
"status": "PASS",
"title": "Meet the AI Agent That Grows With You Hermes Agent by Nous Research",
"author": "Eddy Says Hi #EddySaysHi"
},
"publishDate": "2026-03-21T13:00:09-07:00",
"uploadDate": "2026-03-21T13:00:09-07:00",
"lengthSeconds": 366,
"views": 465,
"watchpage_bytes": 1254068
},
{
"videoId": "6M2tItdARew",
"oembed": {
"status": "PASS",
"title": "The AI Agent That Never Forgets: Meet Hermes Agent by Nous Research",
"author": "Siggi"
},
"publishDate": "2026-03-11T13:53:17-07:00",
"uploadDate": "2026-03-11T13:53:17-07:00",
"lengthSeconds": 371,
"views": 1199,
"watchpage_bytes": 1267812
},
{
"videoId": "5_N84t1rUU0",
"oembed": {
"status": "PASS",
"title": "Hermes Agent Fundamentals In 29 Minutes",
"author": "Tina Huang"
},
"publishDate": "2026-07-20T09:37:02-07:00",
"uploadDate": "2026-07-20T09:37:02-07:00",
"lengthSeconds": 1780,
"views": 462764,
"watchpage_bytes": 1561142
},
{
"videoId": "UTZhvPXnmwA",
"oembed": {
"status": "PASS",
"title": "Hermes Agent: Agents that grow with you |Episode #357|",
"author": "Practical AI"
},
"publishDate": "2026-05-20T15:00:17-07:00",
"uploadDate": "2026-05-20T15:00:17-07:00",
"lengthSeconds": 2853,
"views": 1888,
"watchpage_bytes": 1329765
},
{
"videoId": "1UgXUjT-QtI",
"oembed": {
"status": "PASS",
"title": "Hermes Agent: Why Everyone's Ditching OpenClaw in 2026",
"author": "Luke Alexander AI"
},
"publishDate": "2026-03-26T07:10:29-07:00",
"uploadDate": "2026-03-26T07:10:29-07:00",
"lengthSeconds": 1083,
"views": 13625,
"watchpage_bytes": 1321848
},
{
"videoId": "6GtF_uHbGhw",
"oembed": {
"status": "PASS",
"title": "Every Level of Hermes Agent Explained",
"author": "Jack Roberts"
},
"publishDate": "2026-06-17T12:27:42-07:00",
"uploadDate": "2026-06-17T12:27:42-07:00",
"lengthSeconds": 1535,
"views": 162977,
"watchpage_bytes": 1577121
},
{
"videoId": "CwPUOVUdApE",
"oembed": {
"status": "PASS",
"title": "Hermes Agent: The Ultimate Beginner\u2019s Guide",
"author": "Metics Media"
},
"publishDate": "2026-04-24T06:59:04-07:00",
"uploadDate": "2026-04-24T06:59:04-07:00",
"lengthSeconds": 2228,
"views": 118612,
"watchpage_bytes": 1637578
},
{
"videoId": "sCa3BtpkziQ",
"oembed": {
"status": "PASS",
"title": "100 Days With Hermes Agent in 21 Minutes",
"author": "Sharbel A."
},
"publishDate": "2026-06-17T07:47:26-07:00",
"uploadDate": "2026-06-17T07:47:26-07:00",
"lengthSeconds": 1279,
"views": 57581,
"watchpage_bytes": 1454230
},
{
"videoId": "jmtpYUOr7_U",
"oembed": {
"status": "PASS",
"title": "Hermes Agent Just Killed OpenClaw (Full Tutorial)",
"author": "Leon van Zyl"
},
"publishDate": "2026-04-28T04:19:38-07:00",
"uploadDate": "2026-04-28T04:19:38-07:00",
"lengthSeconds": 1199,
"views": 15988,
"watchpage_bytes": 1510488
},
{
"videoId": "cu2fgknmemA",
"oembed": {
"status": "PASS",
"title": "Hermes Agent The 24/7 Self-Evolving AI Agent!",
"author": "WorldofAI"
},
"publishDate": "2026-04-07T00:01:34-07:00",
"uploadDate": "2026-04-07T00:01:34-07:00",
"lengthSeconds": 555,
"views": 46889,
"watchpage_bytes": 1545788
},
{
"videoId": "4sAmpcSOVEw",
"oembed": {
"status": "PASS",
"title": "Hermes Agent - Crash Course for Beginners (AI Agent)",
"author": "Adrian Twarog"
},
"publishDate": "2026-07-21T01:20:41-07:00",
"uploadDate": "2026-07-21T01:20:41-07:00",
"lengthSeconds": 1338,
"views": 46484,
"watchpage_bytes": 1585548
},
{
"videoId": "P2LIFtrRr2U",
"oembed": {
"status": "PASS",
"title": "Hermes Agent: New FREE OpenClaw Alternative!",
"author": "Julian Goldie SEO"
},
"publishDate": "2026-03-09T14:00:32-07:00",
"uploadDate": "2026-03-09T14:00:32-07:00",
"lengthSeconds": 735,
"views": 9975,
"watchpage_bytes": 1353286
},
{
"videoId": "8GjyOQy19so",
"oembed": {
"status": "PASS",
"title": "Hermes Agent Full Tutorial INSTALLATION + USECASES",
"author": "CodeHead"
},
"publishDate": "2026-05-14T08:00:23-07:00",
"lengthSeconds": 467,
"views": 64285
},
{
"videoId": "zwqhemjHq3E",
"oembed": {
"status": "PASS",
"title": "Hermes Agent vs OpenClaw",
"author": "Sharbel A."
},
"publishDate": "2026-04-20T07:14:00-07:00",
"lengthSeconds": 928,
"views": 35360
}
]
@@ -1,10 +0,0 @@
import re, sys
html = open(sys.argv[1], encoding='utf-8', errors='ignore').read()
ids = re.findall(r'"videoRenderer":\{"videoId":"([\w-]{11})"', html)
print("videoRenderer hits:", len(set(ids)))
for vid in dict.fromkeys(ids):
m = re.search(r'"videoId":"%s".{0,3000}?"title":\{"runs":\[\{"text":"(.*?)"\}' % vid, html, re.S)
ch = re.search(r'"videoId":"%s".{0,6000}?"ownerText":\{"runs":\[\{"text":"(.*?)"' % vid, html, re.S)
dur = re.search(r'"videoId":"%s".{0,4000}?"lengthText":\{"accessibility".{0,400}?"simpleText":"(.*?)"' % vid, html, re.S)
print((vid, m.group(1) if m else "?", ch.group(1) if ch else "?", dur.group(1) if dur else "?"))
@@ -1,42 +0,0 @@
# 06 — Sources
Access date for ALL entries: **2026-09-11** (via citation ledger `sources.py`; doc URLs additionally confirmed HTTP 200 by curl -L).
## Official docs (hermes-agent.nousresearch.com)
| # | URL | Supported |
|---|-----|-----------|
| 1 | https://hermes-agent.nousresearch.com/docs | Docs index; overall feature map |
| 3 | https://hermes-agent.nousresearch.com/docs/user-guide/configuration | Config sections, SOUL.md, checkpoints |
| 4 | https://hermes-agent.nousresearch.com/docs/reference/slash-commands | Slash command registry (03) |
| 5 | https://hermes-agent.nousresearch.com/docs/reference/tools-reference | Toolset inventory (02 §6) |
| 6 | https://hermes-agent.nousresearch.com/docs/user-guide/features/cron | Cron surface (02 §9) |
| 7 | https://hermes-agent.nousresearch.com/docs/user-guide/features/kanban | Kanban surface (02 §8) |
| 8 | https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp | MCP surface + `hermes mcp serve` (02 §5, 04 P3) |
| 9 | https://hermes-agent.nousresearch.com/docs/user-guide/features/memory | Memory surface (02 §2) |
| 10 | https://hermes-agent.nousresearch.com/docs/user-guide/profiles | Profiles (02 §13) |
| 11 | https://hermes-agent.nousresearch.com/docs/integrations/providers | Model/provider routing (02 §12) |
| 12 | https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ | Gateway platforms (02 §15) |
| 13 | https://hermes-agent.nousresearch.com/docs/user-guide/features/curator | Skill maintenance (02 §3) |
| 14 | https://hermes-agent.nousresearch.com/docs/reference/cli-commands | CLI command index cross-check (03) |
| 15 | https://hermes-agent.nousresearch.com/docs/reference/skills-catalog | Skills catalog (02 §3) |
## GitHub
| # | URL | Supported |
|---|-----|-----------|
| 2 | https://github.com/nousresearch/hermes-agent | Repo identity, learning-loop description (01, 02) |
## Local primary sources (not web URLs; verified on this host)
- Live CLI help output, Hermes Agent v0.21.1 (2026.9.7), reference install (Syslog kagentz): `hermes --help` + `hermes {chat,model,config,cron,kanban,skills,sessions,mcp,profile,memory,tools,project,gateway,computer-use,doctor,status,import-agent} --help` → raw dump `cli-help-dump.txt` (577 lines). Basis for all VERIFIED-LIVE tags in 01/02/03/04.
- `/home/hermes/.hermes/skills/autonomous-ai-agents/delegate-coding-agent/SKILL.md` (v1.0.0) + `references/claude-code.md` (v2.2.1) → 04 Patterns 1-2, 01 Claude Code column.
- `/home/hermes/.hermes/skills/autonomous-ai-agents/merge-reconciler/SKILL.md` → 04 Pattern 2.
- `/home/hermes/.hermes/skills/autonomous-ai-agents/computer-use/SKILL.md` (v2.0.0) → 04 Pattern 5, 02 §11.
## YouTube verification
- YouTube search results pages (scraped 2026-09-11): `https://www.youtube.com/results?search_query=hermes+agent+nous+research` and `...nous+research+hermes+agent+official`
- oEmbed endpoint per video: `https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=<ID>&format=json` — all 21 checked IDs returned PASS (HTTP 200, title/author match). Raw evidence incl. publishDate/lengthSeconds/viewCount per watch page: `video-verification.json`.
- 17 listed in 05-videos.md + 4 spares; 21/21 pass, 0 fail.
## Explicit gaps (could not close)
1. No official Nous Research-produced tutorial video was found — video list is third-party ecosystem content (disclosed in 05).
2. Slash commands were verified DOC-ONLY (https://hermes-agent.nousresearch.com/docs/reference/slash-commands); they require an interactive session to exercise, which this headless run does not have. CLI equivalents were verified live.
3. `hermes-agent.nousresearch.com/docs/developer-guide/` returned 404 — developer docs live in-repo (`AGENTS.md` in the GitHub repo), not as a docs site section.
+170 -48
View File
@@ -1,11 +1,14 @@
---
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
# ⛔ The guest/host-keyed report-only gate the GC executor MUST honour lives in the body
# "Hard gate" YAML block below — that block is authoritative and is the only copy.
kind: responsibility
name: disk-gc-threat-response
description: >
Recurring disk health scan, garbage collection, and threat response
across 15 Proxmox CTs + 3 GPU bare-metal hosts. Triggered by incident
across 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts
(fleet verified against `pvesh get /cluster/resources` 2026-09-12). Triggered by incident
2026-07-04 where CT 105 (kagentz) hit 87% disk (49G/59G) from
Docker image bloat — 5 dangling images, 15 build cache layers.
Recovered 35.67GB. Second incident 2026-07-09: amdpve (.15) Docker
@@ -25,7 +28,7 @@ logged within 5 minutes of discovery.
## Scope
All 15 CTs via `pct-run` + 3 GPU bare-metal hosts via direct SSH.
All 20 Proxmox guests (17 LXC via `pct-run` + 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts via direct SSH.
Docker hosts get special attention:
| Host | CT | Disk Risk | GC Strategy |
@@ -40,7 +43,7 @@ Docker hosts get special attention:
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
## Threat Levels
## Threat Levels (GUEST filesystems)
| Level | Threshold | Response | Escalation |
|-------|-----------|----------|------------|
@@ -50,6 +53,39 @@ Docker hosts get special attention:
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
## Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)
Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are **report-only** — no automatic deletion of media or datastore content ever.
| Level | Threshold | Response | Escalation |
|-------|-----------|----------|------------|
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert (state-change only) |
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
**State lives in a small JSON state file:** `state/host-disk-bands.json`, keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
```
storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN
```
**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.
**Action classes by volume type:**
- **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
- **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
- **pbs-datastore** (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.
**Justification (measured 2026-09-15, firstmate):** storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.
## Requires
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
@@ -81,6 +117,32 @@ and escalation trail.
- May also be invoked manually: `prose run disk-gc-threat-response`
- Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
## Scanner: scripts/disk-gc-scan.py
The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability
verdicts deterministic:
1. **Retry on failure:** Each probe retries once before declaring a guest unreachable.
2. **Named probe target:** Every rendered line names the guest, CT id, node, and
access method actually used.
3. **Failure kind printed:** An unreachable guest is reported with its failure kind
(timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare
"unreachable" verdict.
4. **Per-guest access method:** The correct access path is selected from a per-guest
map so the wrong path cannot be picked by an executor improvising:
- CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees
loop0/59G instead of the real 99G filesystem)
- CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM)
- All other CTs = `pct-run <ct_id>` (which uses `pct exec` via SSH to the node)
5. **Every figure traces to a named probe:** The scan output prints the exact command
that produced each disk figure, so two different guests can never render
identical numbers without the probe commands proving it.
Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output).
The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate
from the `report_only_guests` YAML block above.
## Shape
- `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
@@ -99,35 +161,71 @@ and escalation trail.
## Execution
### Host filesystems: report-only, NEVER auto-delete
Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
**CT 111 / hostname `tdunna` / 192.168.68.129 is DETECT-AND-REPORT-ONLY.** It belongs to Theo.
The captain ruled 2026-08-17 and re-confirmed 2026-09-10 that Theo handles CT 111 himself.
At **every** threat level — AMBER, RED, or CRITICAL — the executor must:
- push the threat row and alert the owner, and
- **never** call `gc-executor`, and **never** run any GC command against that guest: no
`apt-get clean/autoremove`, no `journalctl --vacuum-*`, no `find /var/log -delete`, no
`/tmp`/`/var/tmp` deletion, no snap removal, no `docker system prune`.
This gate is keyed on **guest id / hostname / IP**, not on an agent name. The frontmatter
`report_only_agents` marker (e.g. `koby`) names an AGENT while the scan unit is a GUEST, so an
agent-name marker can silently miss the guest it lives on — it must never be the only gate.
**The authoritative machine-readable exclusion list is the YAML block below.** The executor
reads it at run time; `scripts/disk-gc-plan.py` turns a fleet scan into the action plan using it.
Extend the list here, never by hand-maintaining a second copy. The Execution loop below MUST
call that planner and MUST NOT reimplement the gate.
```yaml
# disk-gc report-only guests — authoritative. Keyed on guest/host, not agent.
report_only_guests:
- guest: 111
hostname: tdunna
ip: 192.168.68.129
node: storepve
reason: "Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10"
```
### Loop
```prose
let fleet = call disk-scanner
scope: all
let threats = []
for ct in fleet:
if ct.usage_pct >= 95:
push threats { ct: ct.id, level: "CRITICAL", pct: ct.usage_pct }
else if ct.usage_pct >= 85:
push threats { ct: ct.id, level: "RED", pct: ct.usage_pct }
else if ct.usage_pct >= 75:
push threats { ct: ct.id, level: "AMBER", pct: ct.usage_pct }
-- The report-only gate is IMPLEMENTED IN scripts/disk-gc-plan.py and MUST NOT be
-- reimplemented here. That planner reads the contract's `report_only_guests` YAML block
-- and matches on guest id OR hostname OR IP, so the tested gate is the executed gate.
let plan = call disk-gc-plan
fleet: fleet
-- sort by severity descending
sort threats by pct desc
for row in plan:
if row.action == "report-only":
-- Excluded guest: alert only. No gc-executor call is constructed for it, at any level.
call alerter
threat: row
result: { action: "report-only", reason: row.reason }
else:
let result = call gc-executor
ct: row.target
level: row.level
strategy: lookup-gc-strategy(row.target)
for threat in threats:
let result = call gc-executor
ct: threat.ct
level: threat.level
strategy: lookup-gc-strategy(threat.ct)
call alerter
threat: threat
result: result
call alerter
threat: row
result: result
call summary-reporter
fleet: fleet
threats: threats
plan: plan
```
## GC Strategies by Host Type
@@ -182,6 +280,12 @@ done
## Alert Templates
### HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)
```
⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
Action: {volume_type-specific action}
```
### AMBER (75-84%)
```
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
@@ -239,7 +343,9 @@ dangling images and orphaned build cache. No automated GC was in place.
## Incident Log: 2026-07-09 — amdpve docker bloat
### Discovery
Scheduled fleet disk scan via `pct-run` across all 15 CTs + 3 GPU bare-metal hosts.
Scheduled fleet disk scan across all 20 Proxmox guests (17 LXC via `pct-run`, 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts.
> **Report-only gate applies to this scan:** CT 111 (`tdunna`, 192.168.68.129) is alerted but never
> garbage-collected at any level.
amdpve (.15) flagged at 78% (AMBER threshold: 75%).
### Diagnosis
@@ -272,25 +378,35 @@ one-off GPU builds. No automated post-migration cleanup was in place.
- Contract now scans GPU bare-metal hosts alongside CTs
- Access via `pct-run` script for all CTs (no hardcoded IPs)
## Access Matrix (documented 2026-07-09)
## Access Matrix (verified against `pvesh get /cluster/resources` 2026-09-12)
### CT Access (via pct-run)
| CT | Name | Node | Status |
|----|------|------|--------|
| 100 | abiba | minipve | local |
| 102 | adguard | minipve | ✅ reachable |
| 104 | authentik | minipve | ✅ reachable |
| 105 | kagentz | minipve | ✅ reachable |
| 106 | ra-h-os | storepve | ✅ reachable |
| 107 | pbs | storepve | ✅ reachable |
| 108 | media | storepve | ✅ reachable |
| 110 | gitea | minipve | ✅ reachable |
| 111 | tdunna | amdpve | ✅ reachable |
| 112 | tanko | amdpve | ✅ reachable |
| 113 | baggy | amdpve | ✅ reachable |
| 115 | scottdenya | amdpve | ✅ reachable |
| 116 | syslog-api | minipve | ✅ reachable |
| 117 | zulip | storepve | ✅ reachable |
### Guest Access (via `pct-run` — CT id only, node resolved by `scripts/pct-run.sh`)
| Guest | Name | Node | Type | Status |
|------|------|------|------|--------|
| 100 | abiba | minipve | lxc | ✅ reachable (probed via pct-run like any other guest; no local shortcut) |
| 102 | adguard | minipve | lxc | ✅ reachable |
| 104 | authentik | minipve | lxc | ✅ reachable |
| 105 | kagentz | **amdpve** | lxc | ✅ reachable (was documented as minipve — corrected) |
| 106 | ra-h-os | storepve | lxc | ✅ reachable |
| 107 | pbs | storepve | lxc | ✅ reachable |
| 108 | media | storepve | lxc | ✅ reachable |
| 110 | gitea | minipve | lxc | ✅ reachable |
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
| 112 | tanko | amdpve | lxc | ✅ reachable |
| 113 | baggy | amdpve | lxc | ✅ reachable |
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
| 116 | syslog-api | minipve | lxc | ✅ reachable |
| 117 | zulip | storepve | lxc | ✅ reachable |
| 118 | jdownloader | storepve | lxc | ✅ reachable |
| 119 | infisical-vault | minipve | lxc | ✅ reachable |
| 120 | adguard2 | amdpve | lxc | ✅ reachable |
### QEMU VMs (via direct SSH)
| VM | Name | Node | IP | Status |
|----|------|------|-----|--------|
| 101 | llm-gpu (workload now bare metal .8) | acerpve | — | ✅ reachable |
| 103 | ocu-llm (workload now bare metal .110) | ocupve | — | ✅ reachable |
| 109 | docker-vm | storepve | 192.168.68.7 | ✅ reachable |
### GPU Bare Metal (via direct SSH)
| Host | IP | GPU | Status |
@@ -299,9 +415,15 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| ocu-llm | 192.168.68.110 | RTX 5070 | ✅ reachable |
| amdpve | 192.168.68.15 | Strix Halo | ✅ reachable |
### KVM VM (via direct SSH)
| Host | IP | Role | Status |
|------|-----|------|--------|
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
> **Fleet count:** 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts. Corrected
> 2026-09-12: CT 105 → amdpve, CT 111 → storepve, and guests 118/119/120 were missing.
>
> **CT 100 probe gap (folded in):** CT 100 previously reported "unreachable (not reported)" every
> run. Root cause is the same stale access layer: `pct-run` resolves the guest's node from its map,
> and the map/contract must reflect `pvesh /cluster/resources`. Verified working from inside CT 100:
> `scripts/pct-run.sh 100 "df -P / | tail -1"` → `23% /`. Probe CT 100 through `pct-run` like any
> other guest — never through a local-only path, since the scanner itself runs inside CT 100 and a
> container has no `pct` binary.
>
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.
+4
View File
@@ -261,6 +261,10 @@ repaired.
not exist` on .15. `agent-health-check.py` now carries the live-verified
`storepve` mapping (the script is not the topology source of truth); the
CRITICAL contract itself needs an authorized correction.
**✅ Resolved 2026-09-12:** the topology was corrected in its owner,
`infrastructure-control.prose.md` (CT 111 → storepve, CT 105 → amdpve), and
`scripts/pct-run.sh` now matches. This snapshot is left as observed; treat
those owner documents as authoritative.
2. **Strix Halo `:8080` firewall claim is stale.** `prose-ai-review.sh`
ground-truth rule #4 and `gpu-monitor.prose.md` say `:8080` is firewalled to
`.116` only and `.24` cannot probe it. Live on .15:
+23 -23
View File
@@ -8,12 +8,12 @@ description: >
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx).
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability.
Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context).
Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K.
For >128K on NVIDIA hosts → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba
triggers:
@@ -91,8 +91,8 @@ Single source of truth for models, aliases, rpm caps, weights and fallback chain
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there.
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
**No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
is retired and returns 400 `Invalid model name`.
## Routing Configuration (LiteLLM — July 2026)
@@ -190,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
## Prometheus & Grafana
@@ -207,7 +207,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf`, alias `strix-moe`, 256K context (n_ctx 262144), --parallel 2 --kv-unified, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
@@ -216,17 +216,17 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled.
- **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
- **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
- **Alias-retirement follow-up (2026-09-12)**: the agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, and the executable `audit-hermes-config.py` still reference the retired `gpu-light`/`gemma-4-12b` names, and `audit-hermes-config.py` Rule 8 currently fails a config whose vision/web_extract model is `gpu-vision`. These are tracked separately and are intentionally NOT updated in this change.
- **Alias-retirement sweep (2026-09-12)**: `gemma-4-12b`, `gpu-light` and `crew-auto` are retired and replaced by `gpu-vision` / no cap respectively. The agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, `gpu-self-heal.prose.md`, `hermes-key-enforcement.prose.md`, `inference-optimization.prose.md`, `litellm-client-timeouts.prose.md` and the executable `audit-hermes-config.py` were all updated to the live canonical alias in the same change. **koby's config on .129 still names `gpu-light` (and `gemma-4-E4B`); .129 is report-only, so that is recorded for its owner and NOT edited here.**
## GPU Inference Benchmarks (Current)
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
| Strix Halo (.15) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe) | **65** | 140 | — | **256K** |
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks from 2026-07-17. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe), n_ctx 262144, --parallel 2 --kv-unified. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
GPU contexts: RTX 3090 (.8) and RTX 5070 (.110) at 128K; Strix Halo (.15) at 256K (2026-09-12).
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
@@ -239,7 +239,7 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
### Stable Aliases — CRITICAL
All agent configs MUST use stable role-based aliases, never model-specific names:
- `compression.model: strix-moe`
- `compression.model: syslog-auto`
- `auxiliary.vision.model: gpu-vision`
- `delegation.model: gpu-dense`
- `auxiliary.web_extract.model: gpu-vision`
@@ -247,27 +247,27 @@ All agent configs MUST use stable role-based aliases, never model-specific names
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12)
- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
- Mumuni compression model alias: `strix-moe`
- Mumuni compression model alias: `syslog-auto`
### Mumuni Agent Profile
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs:
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and `audit-hermes-config.py`; whether Mumuni's LIVE config currently complies is a separate operational question.
| Setting | Value | Notes |
|---------|-------|-------|
| `model.default` | `syslog-auto` | Balanced default (pool router) |
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
| `compression.model` | `syslog-auto` | Rule 7: auto-routing, prevents Strix Halo overload |
| `aux.compression.model` | `syslog-auto` | Must match `compression.model` (Rule 7) |
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) |
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
| `memory.memory_char_limit` | 800 | Brief memory entries |
+2 -2
View File
@@ -27,8 +27,8 @@ agent: abiba
┌──────┐ ┌──────┐ ┌────────┐
│.8:8080│ │.110 │ │.116:80 │
│RTX3090│ │:8080 │ │nginx │
│gemma │ │RTX5070│ │router │
└──────┘ │qwen27B│ │LiteLLM │
│qwen │ │RTX5070│ │router │
└──────┘ │vision │ │LiteLLM │
└──────┘ │dashboard│
└────────┘
```
+13 -12
View File
@@ -12,7 +12,7 @@ description: >
Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing.
Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
Stable role-based aliases (strix-moe, gpu-dense, gpu-vision) from gpu-fleet.
agent: abiba
depends_on:
- gpu-monitor.prose.md (live data source on .24:9100)
@@ -48,12 +48,12 @@ depends_on:
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf | ~10/64GB (16%) | 256K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
@@ -64,7 +64,7 @@ Key notes:
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
- **Fix**:
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gpu-vision → gpu-dense, gpu-dense → gpu-vision)
3. If all GPUs hot, alert about cooling infrastructure
- **Verify**: Temp drops below 80°C within 5 minutes
- **Escalate after**: 3 verification failures → Zulip alert
@@ -142,10 +142,10 @@ Key notes:
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- **Detect**: Benchmark tok/s vs baseline for each GPU at its current context
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
- Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
@@ -157,13 +157,14 @@ Key notes:
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
- **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**:
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (gpu-light, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks. Weight: 0.15 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity).
- RTX 5070 (gpu-vision, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks.
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model).
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Weights are not restated here** — the live `syslog-auto` pool weights and rpm caps live in CT 116 `/opt/inference-harness/litellm_config.yaml`, the single source of truth.
- **Fix**:
- Alert if any GPU is handling workload outside its designated role
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-vision, strix-moe)
- Track per-GPU request distribution via LiteLLM spend logs
- **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
@@ -311,6 +312,6 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
### L6: Stable Aliases Replace Model Names
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15; `gpu-light` was superseded by `gpu-vision` on 2026-09-12.
- Self-heal must use aliases for reporting and alerting, not model-specific names.
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
+47 -8
View File
@@ -5,12 +5,30 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
author: Abiba (pi agent)
---
# Hermes Agent Baseline — Canonical Good State
## Reachability Detection
Before checking agent baseline, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no api_key in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Quick Restore
```bash
@@ -24,11 +42,12 @@ done
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|-------|-----|------|-----|---------------|------------|----------|
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
| Koby | 111 | storepve | .129 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
> **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
> CT 111 (tdunna, 192.168.68.129, storepve) is report-only — Theo's box; alert only, never garbage-collect.
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
Keys are stored in Infisical vault (project=agents, env=production) and injected at
@@ -80,7 +99,7 @@ custom_providers:
auxiliary:
vision:
provider: harness
model: gemma-4-12b # or syslog-auto
model: gpu-vision # RTX 5070 stable alias (Rule 8; do not use syslog-auto for aux)
base_url: http://192.168.68.116/litellm/v1
api_key_env: LITELLM_API_KEY
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
@@ -95,13 +114,33 @@ auxiliary:
threshold: 0.65
target_ratio: 0.3
provider: harness
model: syslog-auto # or gemma-4-12b
model: syslog-auto # Rule 7: compression must be syslog-auto
base_url: http://192.168.68.116/litellm/v1
api_key_env: LITELLM_API_KEY
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
timeout: 120
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Known Bug: `api_key_env` Ignored by Auxiliary Client
**Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478)
@@ -185,7 +224,9 @@ Abiba (CT100) runs pi via PM2 with the Zulip extension.
Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`.
**models.json** — Must only list models authorized for the agent's LiteLLM key.
Key is injected via `infisical run --` wrapper at PM2 startup:
`/v1/models` is key-scoped and the live registry is CT 116
`/opt/inference-harness/litellm_config.yaml`; treat the list below as a snapshot and re-read
the registry before applying. Key is injected via `infisical run --` wrapper at PM2 startup:
```json
{
"providers": {
@@ -197,9 +238,7 @@ Key is injected via `infisical run --` wrapper at PM2 startup:
{ "id": "syslog-auto" },
{ "id": "strix-moe" },
{ "id": "gpu-dense" },
{ "id": "gpu-light" },
{ "id": "qwen3.6-27B-code" },
{ "id": "gemma-4-12b" }
{ "id": "gpu-vision" }
]
}
}
+69 -29
View File
@@ -19,6 +19,24 @@ description: >
- agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array
## Reachability Detection
Before auditing the config template, verify the host is reachable and can be checked. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "base_url:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "base_url:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no base_url in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Agent Keys (LiteLLM — Current 2026-07-11)
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
@@ -92,17 +110,17 @@ work immediately after restart.
```yaml
# ─── Model Selection ───
model:
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
default: <agent_model> # e.g., strix-moe, gpu-dense, syslog-auto
provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K).
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
# and falls back to 256K when /v1/models lacks a context
# field (llama-server does). Without this override, agents
# silently run syslog-auto at 256K (verified 2026-08-09).
# Set 65536 if using gemma-4-12b directly (tight VRAM).
# Set 65536 if pinning a single model directly (tight VRAM).
fallback_providers:
provider: deepseek
@@ -133,7 +151,7 @@ compression:
enabled: true
model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
provider: harness
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30
protect_last_n: 40
@@ -143,25 +161,25 @@ compression:
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# model: gpu-vision # stable alias (NOT a raw model name)
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
# api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents.
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
vision:
provider: harness
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
model: gpu-vision # stable alias for RTX 5070
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY
timeout: 60
download_timeout: 30
web_extract:
provider: harness
model: gpu-light # stable alias for RTX 5070
model: gpu-vision # stable alias for RTX 5070
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY
timeout: 30
@@ -173,10 +191,10 @@ auxiliary:
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
# delegation.model and x_search.model use gpu-dense (NOT retired raw name).
delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
model: gpu-dense # stable alias for RTX 3090
provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY
@@ -199,6 +217,26 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
3. **After update**: Restart Hermes on the agent host
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Configuration Rules
### Rule 1: Shared Infra Is Locked
@@ -252,10 +290,11 @@ The following MUST be identical across ALL profiles:
- For agents needing longer outputs: raise to 8192, but never omit
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- **OPERATIONAL DECISION (2026-07-23): Use `syslog-auto` for compression across all agents.**
The `syslog-auto` alias routes to the Strix Halo, but uses the weighted pool instead of pinning
to `strix-moe` directly. This prevents sustained Strix Halo thermal load because the pool can
@@ -264,41 +303,41 @@ The following MUST be identical across ALL profiles:
- All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK)
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 128K context) — the designated compression GPU. This frees the
(64GB UMA, 256K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
- **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
- `auxiliary.compression.model: syslog-auto` (Strix Halo)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
- See `devops-hermes-compression` skill for full reference
### Rule 9: Compression Threshold for 128K Models
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
- See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
- **Koby Exception**: Per captain ruling 2026-08-11, Koby is a DeepSeek-primary external agent; its primary model remains `deepseek-v4-flash` (via api.deepseek.com to preserve DeepSeek-specific reasoning, while other sections follow Rule 10.
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
- `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe
and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
- `syslog-auto` is the LiteLLM routing model — it load-balances across the live pool
(see CT 116 `/opt/inference-harness/litellm_config.yaml` for the current members and weights). Using it protects against:
- Model name typos that cause 403 errors and silent worker failures
- Single GPU downtime (routing falls back automatically)
- Key/model authorization mismatches
@@ -321,7 +360,8 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
verify ALL FOUR of these against the live config. They are the only root causes found in production:
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K).
A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
+50 -6
View File
@@ -117,6 +117,44 @@ model:
api_key_env: LITELLM_API_KEY
```
## Reachability Detection
Before checking for hardcoded keys, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key: sk-" "/root/.hermes/"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no hardcoded keys in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Detection Query
Run on any Hermes host to detect violations:
@@ -169,16 +207,22 @@ The agent picks up the new key via `infisical run --` at gateway startup.
**Keys are permanent and use bare agent name aliases.**
- **Duration**: `null` — keys never expire. This is enforced by `default_key_generate_params` in `litellm_config.yaml`.
- **Duration**: `null` — keys never expire by default. **Expiry must be set EXPLICITLY at creation** with the `duration` parameter (e.g., `90d` for 90 days). The 90-day default is the standard; however, the config default is **NOT honoured** by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns `expires=null`). This has been recorded in `/opt/inference-harness/litellm_config.yaml` to prevent re-filing as a bug.
- **Daily Audit**: A daily audit job runs at 00:00 UTC (`/usr/local/bin/litellm-key-renewal-ct116.sh`, cron 00:00). It is **AUDIT-ONLY** and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES `abiba-pi` and `koby` (report-only, and .129 must never be touched), and logs `RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED` when renewal is skipped. **Renewal is NOT implemented** — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end.
- **Exclusions**: `abiba-pi` and every firstmate/secondmate/crewmate key stay **WITHOUT an expiry** until a proven renewal path exists. `koby` is **report-only** (never touched). These exclusions are enforced by the audit job.
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use.
- **Max budget**: $100 per key (config default).
```yaml
# In litellm_config.yaml — ensures all future keys inherit these defaults:
# NOT currently set in the authority; recommended value. CT 116 litellm_config.yaml has no
# default_key_generate_params block today, and a key generated with no explicit models comes back
# with an EMPTY models list. `models` is a literal key-generation parameter, so this is a value to
# ADD — re-read the live registry at CT 116 /opt/inference-harness/litellm_config.yaml and
# re-verify before applying.
litellm_settings:
default_key_generate_params:
models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b"]
models: ["syslog-auto", "gpu-dense", "gpu-vision", "strix-moe"]
duration: null # ← permanent
max_budget: 100
metadata:
@@ -286,13 +330,13 @@ auxiliary:
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
model: gemma-4-12b
model: gpu-vision
provider: harness
compression:
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
model: gemma-4-12b
model: syslog-auto
provider: harness
```
+3 -3
View File
@@ -47,7 +47,7 @@ connectivity recovery including end-to-end DM validation.
## Requires
- SSH access to target host (direct or via amdpve for CTs)
- SSH access to target host (direct, or via the guest's Proxmox node for CTs)
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
- Python 3 with `httpx` installed on target
@@ -56,7 +56,7 @@ connectivity recovery including end-to-end DM validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------|
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
| Field | Value | Trust |
@@ -72,7 +72,7 @@ connectivity recovery including end-to-end DM validation.
### Step 1: Resolve Target
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
### Step 2: Pull Latest Plugin Source
+3 -3
View File
@@ -43,7 +43,7 @@ gateway restart, and connection validation.
## Requires
- SSH access to target host (direct or via amdpve for CTs)
- SSH access to target host (direct, or via the guest's Proxmox node for CTs)
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
- Python 3 with `httpx` installed on target
- Zulip server accessible at `https://chat.sysloggh.net`
@@ -52,7 +52,7 @@ gateway restart, and connection validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------|
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
| Field | Value | Trust |
@@ -67,7 +67,7 @@ gateway restart, and connection validation.
### Step 1: Locate Target
Map `target` to connectivity parameters from the live-state table above.
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
### Step 2: Deploy Zulip Adapter
+8 -6
View File
@@ -4,7 +4,7 @@ kind: responsibility
description: >
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
assignments, agent context management, and prompt caching — to reduce response
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12).
id: 067NC6KP02RG60S50M40E30928
---
@@ -23,7 +23,7 @@ management, and prompt caching — without sacrificing agent capability.
.123, any others on .129/.122) including compression, model, context_window,
prompt_caching, memory settings
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
qwen .8:8080, gemma .110:8080)
gpu-dense .8:8080, gpu-vision .110:8080)
### Maintains
@@ -56,14 +56,14 @@ duration.
**Context is the root cause.** Every ~46K prompt token costs ~87s of
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: qwen for code/standard queries; gemma for
compression/auxiliary; strix-moe for compression tasks.
- **Route by task**: gpu-dense for code/standard queries; gpu-vision for
vision/web-auxiliary; syslog-auto for compression.
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
compact at 51K, not 85K. Target 15% tail (not 30%).
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
### Shape
@@ -96,8 +96,10 @@ call enable-prompt-caching
hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110]
-- Phase 5: Verify end-to-end latency
-- `models` is a literal verification parameter (a snapshot only): the authoritative registry is
-- CT 116 /opt/inference-harness/litellm_config.yaml; re-read it before use.
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
models: [syslog-auto, gpu-dense, gpu-vision, strix-moe]
```
+92 -15
View File
@@ -16,8 +16,8 @@ description: >
**Last verified:** 2026-08-15 — hwepve removed from Tabiri cluster
(now 5 nodes: minipve, amdpve, storepve, acerpve, ocupve). hwepve
(192.168.68.4) is a standalone PVE node + NetBird routing peer;
London relocation pending. CTs 100 (abiba) and 105 (kagentz) moved
to minipve.
London relocation pending. CT 100 (abiba) is on minipve; CT 105
(kagentz) is on amdpve.
---
# Infrastructure Control Pattern
@@ -105,16 +105,16 @@ description: >
| Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | abiba, kagentz, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
| minipve | .12 | 16C | 30GB | abiba, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | kagentz, tanko, baggy, scottdenya, adguard2 | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
> minipve at .10, not acerpve. Abiba (CT 100) and kagentz (CT 105) are on
> minipve (moved from hwepve 2026-08-15). Mumuni runs inside Abiba CT100
> (.24); CT 114 (mumuni) no longer exists in the cluster.
> **Note:** CTs on storepve include jdownloader (CT 118) and tdunna (CT 111).
> AdGuard (CT 102) is on minipve at .10, not acerpve. Abiba (CT 100) is on
> minipve (moved from hwepve 2026-08-15); kagentz (CT 105) is on amdpve.
> Mumuni runs inside Abiba CT100 (.24); CT 114 (mumuni) no longer exists in the cluster.
>
> **hwepve (192.168.68.4) — STANDALONE (removed from Tabiri 2026-08-15):**
> Huawei MateBook 16 (KLVL-WXX9), pve-manager/9.2.10, kernel 7.0.14-8-pve.
@@ -222,8 +222,8 @@ description: >
**Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gemma)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
- 192.168.68.15:9400 (Strix Halo — strix-moe)
- harness-litellm:4000 (LiteLLM health)
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
@@ -364,6 +364,79 @@ For docker-vm specifically:
- No PBS backup in 48h → fail
```
### Backup Safety Preconditions (2026-09-15)
#### Background & Rationale
Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first:
1. **acerpve thin-pool VM 101** (acerpve, 192.168.68.9, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root.
2. **amdpve 0700 tmpdir** (amdpve, 192.168.68.15, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup.
3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood.
> ⚠️ **Hostname Resolution Warning (2026-09-15)**: The PVE node hostnames (acerpve, amdpve, minipve, storepve, ocupve) all resolve to the VPS (72.61.0.17, the Netbird VPS at srv1079750.hstgr.cloud) via the wildcard `*.dns.sysloggh.net` record, NOT to the actual nodes. So `ssh acerpve` lands on the VPS. **Nodes must be addressed by IP**: acerpve 192.168.68.9, amdpve 192.168.68.15, storepve 192.168.68.6, minipve 192.168.68.12, ocupve 192.168.68.5. Guest CTs are reached through their node (`pct exec`). Guest hostnames that resolve on the LAN (e.g. kagentz = 192.168.68.14) are fine. (The DNS address records are a separate decision — row: dag-daemon-node-hostnames-resolve-to-the-vps-20260915.)
#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts)
Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass:
```bash
# Check 1: Pool headroom (PRIMARY - yields percentages directly)
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname, see warning above
lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data
# Example output (acerpve, 192.168.68.9):
# LV Data% Meta% LSize
# data 29.95 1.22 <816.21g
# Required thresholds (documented minimum):
# data_percent < 90% (80% recommended for safety margin)
# metadata_percent < 70% (metadata fills faster than data)
# Check 2: Verify pool is not in error state (dmsetup shows the raw DM device)
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname
dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1
# dmsetup status pve-data-tpool field order (verified on 192.168.68.9):
# $1=start $2=length $3="thin-pool" $4=transaction-id
# $5=metadata_used/metadata_total (blocks) $6=data_used/data_total (sectors)
# remaining fields are flags ("-", "rw", "discard_passdown", "queue_if_no_space", ...)
# This is only used for the ERROR-STATE check; use the lvs command above for percentages.
# metadata_percent = $5 / ($5 split by /) [second number in pair]
```
**Minimum thresholds**: If either `data_percent >= 90%` or `metadata_percent >= 70%`, the backup MUST NOT start. State explicitly that these are hard stops, not warnings.
**Why this is a precondition**: A full or errored thin pool fails EVERY volume on the VG at once, including the host root. This is not a soft failure - it takes down the entire Proxmox host.
#### Staging Directory Requirement (2026-09-14 incident)
Any custom vzdump `tmpdir` MUST be world-traversable and writable exactly like `/var/tmp` (mode 1777). The archive step of vzdump runs in an unprivileged user namespace and cannot traverse a root-owned 0700 directory.
**Symptom to recognize**: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES` on every container in the backup run.
**Fix**: `chmod 1777 <custom-tmpdir>` before starting vzdump.
#### Task Start Rule for Truncating Shells
When starting a backup task from a shell that may truncate output (e.g., pipes, `head`), always use:
```bash
pvesh create /storage/backup --output-format json -- ... | head -2
# ❌ Can kill the backup task ("broken pipe" status)
```
Instead, capture JSON output without piping to truncating commands:
```bash
# Use --output-format json and capture to variable
result=$(pvesh create /storage/backup --output-format json -- ...)
# Then parse result if needed
```
#### GPU Host Backup Status (acerpve VM 101)
VM 101 (llm-gpu) and VM 103 (ocu-llm) have NO scheduled backup job. Coverage is manual one-off runs only. This is intentional for VM 101 until the thin-pool failure mechanism is understood and documented.
## Section 5: Network Services — Monitoring
### 5.1 Service Inventory
@@ -602,13 +675,13 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 102 | adguard | **minipve** | **.10** | DNS | ❌ |
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
| 104 | authentik | minipve | .11 | OIDC | ❌ |
| 105 | kagentz | minipve | — | Agent Zero | ✅ |
| 105 | kagentz | amdpve | — | Agent Zero | ✅ |
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
| 107 | pbs | storepve | — | Backups | ❌ |
| 108 | media | storepve | — | Media | ❌ |
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | **.17** | Git | ❌ |
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
@@ -616,6 +689,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 117 | zulip | storepve | .19 | Chat | ❌ |
| 118 | jdownloader | storepve | .20 | JDownloader LXC (dedicated, migrated from docker-vm 2026-08-01) | ✅ |
| 119 | infisical-vault | minipve | — | Vault | ❌ |
| 120 | adguard2 | amdpve | — | DNS (secondary AdGuard) | ❌ |
## Appendix C: Docker Compose Files Location
@@ -636,8 +710,8 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| CT | Name | Node | pct-run |
|-----|------|------|---------|
| 100 | abiba | minipve | `pct-run 100` |
| 105 | kagentz | minipve | `pct-run 105` |
| 111 | tdunna | amdpve | `pct-run 111` |
| 105 | kagentz | amdpve | `pct-run 105` |
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
| 112 | tanko | amdpve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` |
@@ -648,6 +722,9 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| 107 | proxmox-backup | storepve | `pct-run 107` |
| 108 | media | storepve | `pct-run 108` |
| 117 | zulip | storepve | `pct-run 117` |
| 118 | jdownloader | storepve | `pct-run 118` |
| 119 | infisical-vault | minipve | `pct-run 119` |
| 120 | adguard2 | amdpve | `pct-run 120` |
| 102 | adguard | **minipve** | `pct-run 102` |
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
+127 -98
View File
@@ -17,7 +17,7 @@ description: >
⚠️ This contract is target-state aspirational — but GPU export + alerting
are now as-built (verified 2026-08-09).
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
version: 1.0.0
version: 1.0.1
---
## Architecture
@@ -120,127 +120,156 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
service answered — report the code, never "down". A redirect is not a failure.
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe,
and that is a statement about YOUR PROBE, not about the service.
2. **A failed probe is never a service verdict.** Print
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
then report. Apply the same shape as scripts/disk-gc-scan.py.
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
(retried at 25s: also timeout)" is actionable.
### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
tool calls; never repeat a prior report unless a live probe fails.**
**PROBE SHAPE (per standing rules above):**
- Every probe prints the target name + URL + HTTP code (or failure kind)
- Retry once on connection failure at longer timeout
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
- Report the actual probe command and its result, not a summary verdict
```bash
# Provenance — run first; paste the absolute path into the report
pwd -P
# Zulip API health (POST ping)
source /etc/litellm-monitor.env
ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
# Expected: 200 (HTTP 000 = unreachable/cache)
# ============================================================
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
# ============================================================
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
if [ -z "$zulip_key" ]; then
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
else
ZULIP_USER="abiba-bot@chat.sysloggh.net"
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
# Expected: 200 (bare-200 probe; any other status is an ALERT)
fi
# PM2 process health
# ============================================================
# 2. PM2 PROCESS HEALTH
# ============================================================
pm2 jlist
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
# spoton-service removed 2026-09-14 (not in live set)
# GPU exporters (may be down per DEPLOYMENT STATUS)
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
# ============================================================
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
# ============================================================
# Probe /metrics (the Prometheus scrape target), not bare /
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
if [ "$code" == "000" ]; then
# Retry with longer timeout
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "GPU exporter http://$host:9400/metrics -> $code"
fi
done
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
# Router health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (Router is up and responding)
# ============================================================
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Router http://192.168.68.116/health -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
# ============================================================
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
fi
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 —
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host,
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
# DOWN = connection refused (000) or timeout only.
# ============================================================
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
# ============================================================
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
printf '%s:8006 -> %s\n' "$node" \
"$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
if [ "$code" == "000" ]; then
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "PVE API https://$node:8006/api2/json/version -> $code"
fi
done
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault.
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)
# ============================================================
# 7. PROMETHEUS TARGETS — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# Grafana health
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# ============================================================
# 8. GRAFANA HEALTH — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
# LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4000/metrics | head -20
# Expected: Prometheus-formatted metrics output
# ============================================================
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
```
**Report format**: Begin every report with the **absolute path the probe executed
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
distinguishable from a real fault at read time. Summarize actual results from
each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE
API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is
DOWN; empty output is a warning. For probes whose expected result is a bare `200`
(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected
status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200`
expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect)
is a stale expectation, not a fault.
from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault
at read time. For each probe, print the target name, the full URL, and the HTTP
code (or failure kind with retry details). Apply the standing probe rules: any
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
a failure.
### Phase 1: GPU Exporters
**NVIDIA (.8 and .110)**:
1. Download `nvidia_gpu_exporter` binary
2. Create systemd service `nvidia-gpu-exporter.service`
3. Start and enable
**AMD (.15)**:
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
2. Parses `amdgpu_top --json -d 1000` output
3. Exposes key metrics at `:9400/metrics` via Python http.server
4. Create systemd service
5. Start and enable
### Phase 2: Prometheus
1. Create `/opt/monitoring/` directory on CT 116
2. Write `prometheus.yml` with scrape configs for all targets
3. Add to docker-compose (or separate compose file)
4. Start container
### Phase 3: Grafana
1. Create `/opt/monitoring/grafana/` directories
2. Provision Prometheus datasource
3. Provision GPU fleet dashboard JSON
4. Provision LiteLLM dashboard JSON
5. Add to docker-compose
6. Start container
### Phase 4: Verification
1. Verify all 3 GPU exporters return 200 at :9400/metrics
2. Verify Prometheus targets all UP at :9090/targets
3. Verify Grafana accessible at :3001 with dashboards
4. Verify LiteLLM metrics flowing to Prometheus
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
## Verification Commands
```bash
# GPU exporters
curl -s http://192.168.68.8:9400/metrics | grep nvidia
curl -s http://192.168.68.110:9400/metrics | grep nvidia
curl -s http://192.168.68.15:9400/metrics | grep amdgpu
# Prometheus
curl -s http://192.168.68.116:9090/api/v1/targets
# Grafana
curl -s http://192.168.68.116:3001/api/health
# LiteLLM metrics (already live)
curl -s http://192.168.68.116:4000/metrics | head -20
```
+13 -3
View File
@@ -67,8 +67,10 @@ description: >
4. **If action == "create"**:
- Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date)
- Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" }
- Duration is null (permanent) — inherited from litellm default_key_generate_params
- Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"]
- Duration is whatever the caller passes; NO default enforcement exists today (CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key with no explicit models returns an empty models list). Agent keys are permanent by policy, not by that block. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
- Return the new key
5. **If action == "rotate"**:
@@ -318,7 +320,15 @@ directly call OpenRouter via Python's requests library. Converting would require
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
```bash
# PRIMARY (proven, runs on CT 116 with no extra tooling):
docker exec harness-litellm printenv LITELLM_MASTER_KEY
# Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching)
# The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it)
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
```
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
+4 -7
View File
@@ -26,9 +26,8 @@ description: >
| Model | avg latency | avg TTFT | p-profile (24h) |
|---|---|---|---|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
| strix-moe | 7.5s | — | Strix Halo, healthy |
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
| gpu-vision (retired gemma-4-12b, RTX 5070) | 2.6s | — | RTX 5070, healthy |
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
backend), full recovery 07:00-08:00 with ZERO client failures once requests
@@ -53,13 +52,11 @@ proxy queuing.
### 2. Auxiliary tasks — keep template timeouts, one correction
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
these are fine.
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on `gemma-4-12b` (retired 2026-09-12); the live RTX 5070 alias is `gpu-vision`.
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
syslog-auto; delegation defaults that assume fast responses will 408 the
same way.
(gpu-dense backend) is the same speed class as syslog-auto; delegation
defaults that assume fast responses will 408 the same way.
### 3. Retry policy — backoff, not repetition
+42 -7
View File
@@ -19,7 +19,7 @@ description: >
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
- Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).
---
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
@@ -44,7 +44,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
## Parameters
@@ -69,7 +69,16 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks — the master key must never be used for inference
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
`strix-moe`, `syslog-auto`). Retrieve from the executor's host via:
```
monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
```
If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys").
The master key must never be used for inference.
## GPU Fleet Topology
@@ -144,12 +153,19 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
check runs on the **backend edge**, not the public edge, so these paths carry the
`/litellm/` prefix:
- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
- Auth uses the dedicated `monitor` agent key. Retrieve via:
`ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"`
Do NOT use the master key for inference — the master key is for admin endpoints only
(`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via:
`ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"`
- KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`,
`gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe
returns 403 and the host is not covered.
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
alias) and re-run — never drop the host from the probe to make the check pass.
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
@@ -161,8 +177,27 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
8. **Check agent keys**:
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
- **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use:
`ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"`
- If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys")
9. **Check Grafana**:
- GET {{grafana_url}}/api/health → expect 200
10. **Compile and report** — Determine overall_status from individual check results
## Executor Script (2026-09-13)
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
checks defined above and reports results in a standardized format. Paste its output in
the status line.
- Hand-rolled probes are **not** an acceptable substitute for the script.
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
because the `harness-docker-stats` container binds to localhost on CT 116.
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
in the remote curl command with proper quoting.
Expected output on a healthy fleet: 11/11 passing checks.
+2 -2
View File
@@ -90,8 +90,8 @@ raw data never provided.
| Worker | Model | Toolsets | Role | Use When |
|--------|-------|----------|------|----------|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
+7 -1
View File
@@ -2,6 +2,10 @@
kind: responsibility
name: pm2-self-heal
description: >
PM2 process health check for abiba-telegram, abiba-zulip, gitea-runner, and zulip-watchdog.
gpu-monitor is systemd-managed (gpu-monitor.service), NOT PM2.
gpu-watchdog is decommissioned and folded into gpu-monitor.service.
gitea-runner is KEPT. abiba-zulip is KEPT (online for days).
---
## Maintains
@@ -9,7 +13,6 @@ description: >
- abiba-telegram: { status: "online", uptime: string, restarts: number }
- abiba-zulip: { status: "online", uptime: string, restarts: number }
- gitea-runner: { status: "online", uptime: string, restarts: number }
- spoton-service: { status: "online", uptime: string, restarts: number }
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp
@@ -36,6 +39,9 @@ description: >
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
spoton incident). Alerts include the restart count.
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
four processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog). The spoton
reference above is historical context for the crash-loop guard, not a live process.
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
+1 -1
View File
@@ -87,7 +87,7 @@ agent: abiba
| storepve | 192.168.68.6 | PVE |
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
## Operations
@@ -0,0 +1,44 @@
disk-gc report-only verification — CT 111 / tdunna / 192.168.68.129
date: 2026-09-12T19:11:36Z
host: abiba (this scanner runs INSIDE CT 100 / abiba)
branch head: 6e612ce37b1b9a9688b04e7f816848e7a4185fca
command: scripts/disk-gc-plan.py --scan <live fleet scan>
PURPOSE: prove that on a REAL fleet scan, CT 111 is alerted and NO gc-executor action
is emitted for it at any level. No GC command was executed against .129.
--- live fleet scan (df -P / via scripts/pct-run.sh for LXC, direct SSH for hosts) ---
tdunna 84%
acerpve 192.168.68.9 77%
amdpve 192.168.68.15 76%
ocu-llm 192.168.68.110 69%
storepve 192.168.68.6 65%
kagentz 61%
tanko 56%
minipve 192.168.68.12 49%
authentik 45%
infisical-vault 39%
adguard 38%
ocupve 192.168.68.5 38%
scottdenya 35%
syslog-api 34%
llm-gpu 192.168.68.8 25%
abiba 23%
baggy 22%
jdownloader 21%
gitea 16%
ra-h-os 13%
zulip 12%
docker-vm 192.168.68.7 11%
adguard2 10%
media 9%
proxmox-backup-server 4%
--- planner output (action plan) ---
111 AMBER 84.0% -> REPORT-ONLY (no GC) — Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10
acerpve AMBER 77.0% -> gc-executor
amdpve AMBER 76.0% -> gc-executor
--- verdict ---
CT 111 (tdunna) 84% AMBER -> report-only; no gc-executor row emitted; no GC run on .129.
Owned hosts acerpve .9 (77%) and amdpve .15 (76%) -> gc-executor (ours).
+64 -26
View File
@@ -340,6 +340,36 @@ def check_gpu_ports():
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
# ═══════════════════════════════════════════════════════════════════
def _ssh_retry(host, cmd, user="root", timeout=15, retry_timeout=25, label=""):
"""SSH with one retry at a longer timeout.
Returns (stdout_or_None, probe_failed_bool, fail_kind).
When probe_failed is True, fail_kind is one of: timeout, ssh-failed.
"""
import subprocess as _sp
def _attempt(tmo, conn_tmo):
try:
r = _sp.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", f"ConnectTimeout={conn_tmo}",
f"{user}@{host}", cmd],
capture_output=True, text=True, timeout=tmo)
return r.stdout.strip() if r.returncode == 0 else None
except _sp.TimeoutExpired:
return "__timeout__"
except:
return None
result = _attempt(timeout, 8)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
prefix = f"{label} " if label else ""
print(f" probe-failed: {prefix}ssh {user}@{host} — {kind} (retrying at {retry_timeout}s…)")
result = _attempt(retry_timeout, 15)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
return None, True, kind
return result, False, None
def check_agents():
for name, agent in AGENTS.items():
host = agent.get("host")
@@ -355,42 +385,49 @@ def check_agents():
is_dsh = agent.get("runtime") == "dsh"
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
since = "since 2026-08-27" if is_dsh else "since the harness purge"
live = ssh(host, "true", user=user)
print(f" {'✅' if live is not None else '❌'} {name}: {label} — "
f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
if live is None:
_fail(f"unreachable:{name}", name)
live, probe_failed, fail_kind = _ssh_retry(host, "true", user=user)
if probe_failed:
print(f" ❌ {name}: {label} — probe-failed: ssh {user}@{host} {fail_kind} "
f"(retried at 25s: also {fail_kind}) [CT {ct}]")
_fail(f"probe-failed:{name}:{fail_kind}", name)
else:
print(f" ✅ {name}: {label} — no Hermes gateway {since} "
f"(ssh {user}@{host} OK, CT {ct})")
continue
if not host or not user:
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
continue
# Resolve the Hermes gateway PID once, before the report-only branch:
# the summary line below renders `pid`, and it used to be bound only in
# the report-only path — leaving it unbound on the abiba/koonimo path
# raised UnboundLocalError and crashed the whole check. Agents without
# a gateway get pid=?.
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid:
# Resolve the Hermes gateway PID with retry. The probe target is
# explicit: ssh {user}@{host} pgrep -f hermes gateway.
pid, probe_failed, fail_kind = _ssh_retry(
host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid and not probe_failed:
pid, probe_failed, fail_kind = _ssh_retry(
host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid and not probe_failed:
pid = "?"
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
if probe_failed:
print(f" ❌ {name}: probe-failed: ssh {user}@{host} {fail_kind} "
f"(retried at 25s: also {fail_kind}) [CT {ct}] — gateway status UNDETERMINED")
_fail(f"probe-failed:{name}:{fail_kind}", name)
continue
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only (captain's 2026-08-17 ruling)
if report_only:
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
# Still check gateway status for reporting purposes
if pid == "?":
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> no process found (reported only, NOT counted) [CT {ct}]")
_fail(f"gateway-down:{name}", name)
continue
else:
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
continue # Skip the rest of the check for Koby
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> pid={pid} (running, reported only, NOT repaired) [CT {ct}]")
continue # Skip the rest of the check for Koby
# Gateway state file
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
state, _, _ = _ssh_retry(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
if state:
try:
st = json.loads(state)
@@ -408,21 +445,22 @@ def check_agents():
]
streaming = "no"
for p in adapter_paths:
has_edit = ssh(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
has_edit, _, _ = _ssh_retry(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
if has_edit and has_edit != "0":
streaming = "yes"
break
# Recent errors
recent_errors = ssh(host,
recent_errors, _, _ = _ssh_retry(
host,
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
user=user)
recent_errors = (recent_errors or "0").strip().split("\n")[-1]
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
f"{name}: probe: ssh {user}@{host} — gw={gw_state} zulip={zulip} "
f"streaming={streaming} errors_10m={recent_errors.strip() or '0'} pid={pid} [CT {ct}]")
# ═══════════════════════════════════════════════════════════════════
+206
View File
@@ -0,0 +1,206 @@
#!/usr/bin/env python3
"""disk-gc-plan — turn a fleet disk scan into the GC action plan.
This is the executable side of `disk-gc-threat-response.prose.md`. It exists so the
report-only gate is enforced by code that can be tested, rather than by prose the
executor might misread.
THE HARD GATE: guests listed in the contract's `report_only_guests` block are
DETECT-AND-REPORT-ONLY at EVERY level (AMBER, RED, CRITICAL). This tool will never
emit a `gc-executor` action for one, so no GC command can be constructed for it.
The gate is keyed on GUEST identity — guest id, hostname, or IP — never on an agent
name. An agent-name marker can silently miss the guest it lives on; a guest marker
cannot.
The authoritative exclusion list lives in the contract itself (the fenced ```yaml
block containing `report_only_guests:`). This tool reads it from there so there is
only ever one copy.
Usage:
disk-gc-plan.py --scan scan.json # [{"id":111,"usage_pct":84}, ...]
cat scan.json | disk-gc-plan.py # same, via stdin
disk-gc-plan.py --scan scan.json --json # machine-readable plan
Scan entries may carry any of: id / guest / vmid / ct / ctid, hostname / name, ip.
A threshold-crossing entry with no recognizable identity is reported, never GC'd.
Exit codes: 0 ok, 1 usage/parse error.
"""
from __future__ import annotations
import argparse
import json
import pathlib
import re
import sys
import yaml
REPO = pathlib.Path(__file__).resolve().parent.parent
DEFAULT_CONTRACT = REPO / "disk-gc-threat-response.prose.md"
AMBER, RED, CRITICAL = 75, 85, 95
IDENTITY_FIELDS = ("id", "guest", "vmid", "ct", "ctid", "hostname", "name", "ip")
IDENTITY_TYPE_PREFIX = re.compile(r"^(?:lxc|qemu)/")
UNIDENTIFIED_REASON = "unidentified target - refusing to schedule GC"
def load_report_only_guests(contract_path: pathlib.Path) -> list[dict]:
"""Read the authoritative report_only_guests block out of the contract.
The contract carries it as a fenced ```yaml block. Parsing the declared,
machine-readable block is the intended interface — the contract owns the list.
"""
text = contract_path.read_text(encoding="utf-8")
for block in re.findall(r"```yaml\n(.*?)```", text, re.S):
if "report_only_guests:" in block:
data = yaml.safe_load(block)
guests = data.get("report_only_guests") or []
if not isinstance(guests, list):
raise SystemExit("report_only_guests must be a list")
if not guests:
raise SystemExit(
"report_only_guests is empty or missing - refusing to plan GC "
"without the report-only gate"
)
for guest in guests:
if not isinstance(guest, dict) or not _keys(guest):
raise SystemExit(
"report_only_guests entry has no recognizable identity key "
f"(expected one of: {', '.join(IDENTITY_FIELDS)}): {guest!r}"
)
return guests
raise SystemExit(
f"no authoritative report_only_guests block found in {contract_path}"
)
def _canonical_number(number: float) -> str:
if float(number).is_integer():
return str(int(number))
return str(number).strip().lower()
def _normalize_identity(value: object) -> str:
"""Canonicalise a guest identity so differently-encoded ids compare equal:
numeric and numeric-string ids collapse to an integer string, Proxmox
type prefixes and leading zeros are stripped, and hostnames/IPs are only
trimmed and lowercased."""
if isinstance(value, bool):
return str(value).strip().lower()
if isinstance(value, (int, float)):
return _canonical_number(float(value))
text = str(value).strip().lower()
text = IDENTITY_TYPE_PREFIX.sub("", text)
try:
return _canonical_number(float(text))
except ValueError:
return text
def _keys(entry: dict) -> set[str]:
"""Guest/host identity keys, shared by exclusions and scan entries so the two
sides of the gate can never key on different fields."""
out: set[str] = set()
for field in IDENTITY_FIELDS:
value = entry.get(field)
if value is None:
continue
key = _normalize_identity(value)
if key:
out.add(key)
return out
def level_for(pct: float) -> str | None:
if pct >= CRITICAL:
return "CRITICAL"
if pct >= RED:
return "RED"
if pct >= AMBER:
return "AMBER"
return None
def build_plan(scan: list[dict], report_only: list[dict]) -> list[dict]:
excluded = [(e, _keys(e)) for e in report_only]
plan: list[dict] = []
for entry in scan:
pct = entry.get("usage_pct")
if pct is None:
continue
level = level_for(float(pct))
if level is None:
continue # GREEN: log only, no action
scan_keys = _keys(entry)
target = next(
(entry.get(k) for k in IDENTITY_FIELDS if entry.get(k) not in (None, "")),
"?",
)
if not scan_keys:
plan.append({
"target": target,
"level": level,
"pct": float(pct),
"action": "report-only",
"reason": UNIDENTIFIED_REASON,
})
continue
match = next((e for e, keys in excluded if keys & scan_keys), None)
if match is not None:
plan.append({
"target": target,
"level": level,
"pct": float(pct),
"action": "report-only",
"reason": match.get("reason", "").strip(),
})
else:
plan.append({
"target": target,
"level": level,
"pct": float(pct),
"action": "gc-executor",
})
plan.sort(key=lambda row: row["pct"], reverse=True)
return plan
def main() -> int:
ap = argparse.ArgumentParser(description="Plan disk GC actions with the report-only gate.")
ap.add_argument("--scan", help="JSON file: list of {id|ct|hostname|ip, usage_pct}")
ap.add_argument("--contract", default=str(DEFAULT_CONTRACT))
ap.add_argument("--json", action="store_true", help="emit the plan as JSON")
args = ap.parse_args()
raw = pathlib.Path(args.scan).read_text() if args.scan else sys.stdin.read()
try:
scan = json.loads(raw)
except json.JSONDecodeError as exc:
print(f"invalid scan JSON: {exc}", file=sys.stderr)
return 1
if not isinstance(scan, list):
print("scan must be a JSON list", file=sys.stderr)
return 1
report_only = load_report_only_guests(pathlib.Path(args.contract))
plan = build_plan(scan, report_only)
if args.json:
print(json.dumps(plan, indent=2))
return 0
if not plan:
print("no threats (nothing at or above 75%)")
return 0
for row in plan:
if row["action"] == "report-only":
print(f" {row['target']} {row['level']} {row['pct']}% -> REPORT-ONLY (no GC) — {row['reason']}")
else:
print(f" {row['target']} {row['level']} {row['pct']}% -> gc-executor")
return 0
if __name__ == "__main__":
sys.exit(main())
+339
View File
@@ -0,0 +1,339 @@
#!/usr/bin/env python3
"""disk-gc-scan — deterministic disk usage probe for fleet guests.
This is the executable scanner side of `disk-gc-threat-response.prose.md`. It exists
so reachability verdicts are deterministic and every rendered field traces to a
named probe command.
DESIGN PRINCIPLES (per task disk-gc-probe-false-unreachable-20260913):
1. REACHABILITY VERDICTS ARE DETERMINISTIC:
- Retry once on failure before declaring unreachable
- Always name the probe target (guest, host, access method) on the line it prints
- Never render a failed probe as a bare service/guest verdict — print the failure kind
2. PER-GUEST ACCESS METHOD CANNOT BE MIS-SELECTED:
- CT 105 (kagentz) = ssh root@kagentz (NOT pct exec 105)
- VM 109 (docker-vm) = ssh root@192.168.68.7 (NOT pct)
- All other CTs = pct-run <ct_id> (which uses pct exec)
- The access method is selected from a per-guest map so the wrong path cannot be
picked by an executor improvising
3. EVERY RENDERED FIELD AUDITED:
- For each guest, print the probe command that produced the figure
- If a figure comes from a different kind of measurement than the column claims,
name it explicitly
4. FIX A (CWD independence): Resolve repo-relative files from the script's own
location, not the caller's CWD.
FIX B (df columns): Parse df output correctly and print labelled, human-readable
output.
Usage:
disk-gc-scan.py # scan all guests
disk-gc-scan.py --json # machine-readable output
Exit codes: 0 ok (all guests probed), 1 probe error
"""
from __future__ import annotations
import json
import pathlib
import subprocess
import sys
import time
import os
from dataclasses import dataclass
from typing import Optional
# Resolve repo-relative files from the script's own location, not the caller's CWD
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
HELPER_PCT_RUN = SCRIPT_DIR / "pct-run.sh"
# Per-guest access method map. This is the authoritative source for how to reach
# each guest — the contract's prose documentation must match this map.
#
# Access methods:
# - "pct-run": use pct-run.sh <ct_id> (pct exec via SSH to node)
# - "ssh-host": use ssh root@<hostname>
# - "ssh-ip": use ssh root@<ip>
@dataclass
class Guest:
"""A guest to probe."""
ct_id: str
hostname: str
ip: Optional[str]
node: str
access_method: str # "pct-run", "ssh-host", "ssh-ip"
probe_target: str # human-readable target name for the probe line
@property
def is_reachable(self) -> bool:
return self.probe_result is not None and self.probe_result.exit_code == 0
@property
def usage_pct(self) -> Optional[float]:
return self.probe_result.usage_pct if self.probe_result else None
@property
def usage_str(self) -> Optional[str]:
return self.probe_result.usage_str if self.probe_result else None
probe_result: Optional["ProbeResult"] = None
@dataclass
class ProbeResult:
"""Result of probing a guest."""
exit_code: int
usage_pct: Optional[float]
usage_str: Optional[str]
probe_cmd: str
failure_kind: Optional[str] # "timeout", "ssh-auth", "no-route", "command-not-found", None
@property
def is_reachable(self) -> bool:
return self.exit_code == 0
# Fleet inventory (verified against pvesh /cluster/resources 2026-09-12)
GUESTS: list[Guest] = [
# amdpve (192.168.68.15)
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
access_method="pct-run", probe_target="scottdenya (CT 115, amdpve)"),
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
# minipve (192.168.68.12)
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
access_method="pct-run", probe_target="adguard (CT 102, minipve)"),
Guest(ct_id="104", hostname="authentik", ip="192.168.68.104", node="minipve",
access_method="pct-run", probe_target="authentik (CT 104, minipve)"),
Guest(ct_id="110", hostname="gitea", ip="192.168.68.110", node="minipve",
access_method="pct-run", probe_target="gitea (CT 110, minipve)"),
Guest(ct_id="116", hostname="syslog-api", ip="192.168.68.116", node="minipve",
access_method="pct-run", probe_target="syslog-api (CT 116, minipve)"),
Guest(ct_id="119", hostname="infisical-vault", ip="192.168.68.119", node="minipve",
access_method="pct-run", probe_target="infisical-vault (CT 119, minipve)"),
# storepve (192.168.68.6)
Guest(ct_id="106", hostname="ra-h-os", ip="192.168.68.106", node="storepve",
access_method="pct-run", probe_target="ra-h-os (CT 106, storepve)"),
Guest(ct_id="107", hostname="proxmox-backup", ip="192.168.68.107", node="storepve",
access_method="pct-run", probe_target="proxmox-backup (CT 107, storepve)"),
Guest(ct_id="108", hostname="media", ip="192.168.68.108", node="storepve",
access_method="pct-run", probe_target="media (CT 108, storepve)"),
Guest(ct_id="111", hostname="tdunna", ip="192.168.68.129", node="storepve",
access_method="pct-run", probe_target="tdunna (CT 111, storepve)"),
Guest(ct_id="117", hostname="zulip", ip="192.168.68.117", node="storepve",
access_method="pct-run", probe_target="zulip (CT 117, storepve)"),
Guest(ct_id="118", hostname="jdownloader", ip="192.168.68.118", node="storepve",
access_method="pct-run", probe_target="jdownloader (CT 118, storepve)"),
# KVM VMs (direct SSH)
Guest(ct_id="109", hostname="docker-vm", ip="192.168.68.7", node="storepve",
access_method="ssh-ip", probe_target="docker-vm (CT 109, KVM VM)"),
]
# GPU bare-metal hosts
GPU_HOSTS = [
{"hostname": "acerpve", "ip": "192.168.68.9", "gpu": "RTX 3090",
"probe_target": "RTX 3090 (bare metal .9)"},
{"hostname": "ocupve", "ip": "192.168.68.110", "gpu": "RTX 5070",
"probe_target": "RTX 5070 (bare metal .110)"},
{"hostname": "amdpve", "ip": "192.168.68.15", "gpu": "Strix Halo",
"probe_target": "Strix Halo (bare metal .15)"},
]
CONNECT_TIMEOUT = 5
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
"""Run a command and return (exit_code, stdout, stderr)."""
try:
result = subprocess.run(
cmd, shell=True, capture_output=True, text=True, timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 124, "", "timeout"
except Exception as e:
return 1, "", str(e)
def probe_guest(guest: Guest) -> ProbeResult:
"""Probe a single guest and return the result.
Access method is selected from guest.access_method:
- "pct-run": pct-run.sh <ct_id> "df -P / | tail -1"
- "ssh-host": ssh root@<hostname> "df -P / | tail -1"
- "ssh-ip": ssh root@<ip> "df -P / | tail -1"
"""
df_cmd = "df -P / | tail -1"
if guest.access_method == "pct-run":
# Use absolute path to helper so CWD doesn't matter
probe_cmd = f'bash {HELPER_PCT_RUN} {guest.ct_id} "{df_cmd}"'
elif guest.access_method == "ssh-host":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.hostname} "{df_cmd}"'
elif guest.access_method == "ssh-ip":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.ip} "{df_cmd}"'
else:
raise ValueError(f"unknown access_method: {guest.access_method}")
# Check helper exists and is readable BEFORE probing (for pct-run guests)
# This prevents scanner errors from being rendered as guest verdicts
if guest.access_method == "pct-run":
if not HELPER_PCT_RUN.exists():
print(f"SCANNER ERROR: helper not found: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
if not os.access(str(HELPER_PCT_RUN), os.R_OK):
print(f"SCANNER ERROR: helper not readable: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
# Retry once on failure before declaring unreachable
for attempt in range(2):
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
if exit_code == 0:
# Parse df output: Filesystem 1024-blocks Used Available Capacity Mounted on
# parts[0]=Filesystem, parts[1]=Total (1K blocks), parts[2]=Used, parts[3]=Available, parts[4]=Capacity
parts = stdout.split()
if len(parts) >= 5:
capacity_str = parts[4] # e.g., "34%"
usage_pct = float(capacity_str.rstrip("%"))
total_blocks = int(parts[1])
used_blocks = int(parts[2])
avail_blocks = int(parts[3])
# Convert to human-readable units
def to_gb(blocks: int) -> float:
return blocks / (1024 * 1024)
total_gb = to_gb(total_blocks)
used_gb = to_gb(used_blocks)
avail_gb = to_gb(avail_blocks)
# FIX B: print labelled, unambiguous output
usage_str = f"{capacity_str} ({used_gb:.1f}G used of {total_gb:.1f}G total, {avail_gb:.1f}G free)"
return ProbeResult(
exit_code=0,
usage_pct=usage_pct,
usage_str=usage_str,
probe_cmd=probe_cmd,
failure_kind=None,
)
else:
# Unexpected output format
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="parse-error",
)
else:
# Classify failure kind
if exit_code == 124:
failure_kind = "timeout"
elif "Connection timed out" in stderr or "timed out" in stderr:
failure_kind = "timeout"
elif "Permission denied" in stderr or "password" in stderr.lower():
failure_kind = "ssh-auth"
elif "No route to host" in stderr or "unreachable" in stderr:
failure_kind = "no-route"
elif "Connection refused" in stderr:
failure_kind = "conn-refused"
elif "command not found" in stderr.lower() or "No such file" in stderr:
failure_kind = "command-not-found"
else:
failure_kind = f"ssh-exit-{exit_code}"
# Retry once
if attempt == 0:
time.sleep(1)
continue
return ProbeResult(
exit_code=exit_code,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind=failure_kind,
)
# Should not reach here, but just in case
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="unknown",
)
def scan_fleet() -> list[dict]:
"""Scan all guests and return the results."""
results = []
for guest in GUESTS:
probe_result = probe_guest(guest)
guest.probe_result = probe_result
row = {
"target": guest.probe_target,
"ct_id": guest.ct_id,
"hostname": guest.hostname,
"node": guest.node,
"access_method": guest.access_method,
"reachable": probe_result.is_reachable,
"usage_pct": probe_result.usage_pct,
"usage_str": probe_result.usage_str,
"probe_cmd": probe_result.probe_cmd,
"failure_kind": probe_result.failure_kind,
}
results.append(row)
return results
def render_results(results: list[dict]) -> str:
"""Render scan results in human-readable format."""
lines = []
lines.append("=== Disk GC Scan ===")
lines.append("")
for row in results:
if row["reachable"]:
lines.append(f" ✅ {row['target']}: {row['usage_str']}")
lines.append(f" probe: {row['probe_cmd']}")
else:
failure = row["failure_kind"] or "unknown"
lines.append(f" ❌ {row['target']}: UNREACHABLE ({failure})")
lines.append(f" probe: {row['probe_cmd']}")
return "\n".join(lines)
def main() -> int:
import argparse
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
ap.add_argument("--json", action="store_true", help="machine-readable output")
args = ap.parse_args()
results = scan_fleet()
if args.json:
print(json.dumps(results, indent=2))
else:
print(render_results(results))
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
# (a probe error means the probe itself failed, not just that the guest was unreachable)
return 0
if __name__ == "__main__":
sys.exit(main())
+32
View File
@@ -0,0 +1,32 @@
#!/bin/bash
# Shared helper for Hermes contract reachability checks
# Separates SSH exit status from remote command result
hermes_check_host() {
local host=$1
local pattern=$2
local path=$3
# Remote side always succeeds (grep ...; true), so ssh exit code = connection status only
local out
out=$(ssh -o BatchMode=yes -o ConnectTimeout=3 root@"$host" "grep -RIn '$pattern' '$path' 2>/dev/null; true" 2>/dev/null)
local status=$?
if [ $status -ne 0 ]; then
echo "$host: UNREACHABLE (ssh exit $status)"
elif [ -n "$out" ]; then
echo "$host: VIOLATION: $out"
else
echo "$host: COMPLIANT (no matches found)"
fi
}
# Standalone mode: scripts/hermes-reachability-check.sh <host> <pattern> <path>
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
if [ $# -ne 3 ]; then
echo "Usage: $0 <host> <pattern> <path>" >&2
exit 2
fi
hermes_check_host "$1" "$2" "$3"
exit 0
fi
+304
View File
@@ -0,0 +1,304 @@
#!/usr/bin/env python3
"""
LiteLLM Health Check - Contract executor
Runs all checks defined in litellm-health.prose.md and reports results.
"""
import subprocess
import sys
import json
import time
import random
# Configuration
BACKEND_HOST = "192.168.68.116"
GPU_HOSTS = {
"gpu-dense": "192.168.68.8",
"gpu-vision": "192.168.68.110",
"strix-moe": "192.168.68.15"
}
def run_command(cmd, timeout=15):
"""Run a command and return (exit_code, stdout, stderr)"""
try:
result = subprocess.run(
cmd,
shell=True,
capture_output=True,
text=True,
timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 1, "", "TIMEOUT"
except Exception as e:
return 1, "", str(e)
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
"""Probe HTTP endpoint and return (status_code, failure_kind)
Returns:
(code, None) if successful or HTTP response received
(000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc.
"""
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
if follow_redirects:
cmd += " -L"
cmd += " '" + url + "'"
try:
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0:
# Determine failure kind from curl exit code
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
if rc == 28:
return (000, "timeout after " + str(timeout) + "s")
elif rc == 7:
return (000, "connection refused")
elif rc == 6:
return (000, "dns failure")
elif rc == 35:
return (000, "ssl error")
elif rc == 52:
return (000, "empty response")
else:
return (000, "curl exit " + str(rc))
return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response")
except subprocess.TimeoutExpired:
return (000, "timeout after " + str(timeout) + "s")
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
cmd = "curl -s -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
cmd += " '" + url + "'"
rc, stdout, stderr = run_command(cmd, timeout)
# Return first 200 chars, single line
body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else ""
return body
def check_liveliness():
"""Step 1: Liveliness probe"""
code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
def check_containers():
"""Step 2: Container health via SSH"""
cmd = "ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@192.168.68.116 'docker ps --format \"{{.Names}} {{.Status}}\"'"
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Containers", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
lines = stdout.split('\n') if stdout else []
container_count = len([l for l in lines if l.strip()])
healthy = container_count >= 8
return "Containers", healthy, str(container_count) + " containers"
def check_model_probes():
"""Step 6: Model probe - all 4 aliases"""
# Get monitor key
monitor_key = run_command("ssh -o BatchMode=yes root@192.168.68.116 \"grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2\"")[1]
results = []
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
# Single-host aliases: 30s timeout each
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
if code == 000 and failure_kind:
# Report probe failure with kind, do not assert a service verdict
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
elif code == 200:
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
elif code in (401, 403):
# Credential fault - capture body and key alias
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
# Resolve key alias
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
# Pool alias (syslog-auto): 60s timeout, retry once on 000
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
# Retry once with same timeout
time.sleep(1)
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)"))
elif code in (401, 403):
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
alias = "monitor-20260813"
results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
return results
def check_admin_key_list():
"""Step 8: Admin API key list - use two-step approach"""
# Step 1: Get master key
mk_cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"docker exec harness-litellm printenv LITELLM_MASTER_KEY\""
mk_rc, mk_stdout, mk_stderr = run_command(mk_cmd)
if mk_rc != 0:
return "Admin Key List", False, "credential-missing (ssh failed: " + mk_stderr + ")"
mk = mk_stdout
if not mk or "NO-CURL" in mk:
return "Admin Key List", False, "credential-missing (empty or NO-CURL)"
# Step 2: Call using the key - use double quotes inside SSH command
cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"curl -s -H \\\"Authorization: Bearer " + mk + "\\\" http://127.0.0.1:4000/key/list\""
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Admin Key List", False, "admin-call-failed (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Try to parse the response
try:
data = json.loads(stdout)
# Response is a dict with "keys" field
if isinstance(data, dict) and "keys" in data:
key_count = len(data["keys"])
elif isinstance(data, list):
key_count = len(data)
else:
key_count = 0
if key_count == 0:
return "Admin Key List", False, "admin-call-failed (empty response)"
return "Admin Key List", True, str(key_count) + " keys"
except Exception as e:
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
def check_github_status():
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
code, _ = probe_http("https://status.github.com/api/status.json", timeout=15)
# GitHub status API returns 301 redirect, which is expected behavior
return "GitHub Status", code == 301, str(code)
def check_prometheus():
"""Step 4: Prometheus health"""
code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
def check_grafana():
"""Step 9: Grafana health"""
code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
def check_docker_stats():
"""Step 10: Docker Stats health - fetch from CT 116 host"""
# Docker stats is on localhost from CT 116
cmd = "ssh -o BatchMode=yes root@192.168.68.116 'curl -s http://127.0.0.1:9324/metrics | head -20'"
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Docker Stats", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Check response is non-empty
if not stdout or len(stdout) < 100:
return "Docker Stats", False, "empty response"
return "Docker Stats", True, "200 (target: 127.0.0.1:9324/metrics from CT 116)"
def main():
print("🏥 LiteLLM Health Check v1.0.0")
print("📍 Backend edge: http://" + BACKEND_HOST)
print("")
all_pass = True
# Run all checks
checks = [
check_liveliness(),
check_containers(),
check_prometheus(),
check_grafana(),
]
for result in checks:
name, passed, detail = result
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
if not passed:
all_pass = False
# Model probes
model_results = check_model_probes()
for name, passed, detail in model_results:
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
if not passed:
all_pass = False
# Admin key list
admin_result = check_admin_key_list()
status = "✅" if admin_result[1] else "❌"
print(" " + status + " Admin Key List: " + admin_result[2])
if not admin_result[1]:
all_pass = False
# GitHub status
github_result = check_github_status()
status = "✅" if github_result[1] else "❌"
print(" " + status + " " + github_result[0] + ": " + github_result[2])
if not github_result[1]:
all_pass = False
# Docker stats
docker_stats_result = check_docker_stats()
status = "✅" if docker_stats_result[1] else "❌"
print(" " + status + " " + docker_stats_result[0] + ": " + docker_stats_result[2])
if not docker_stats_result[1]:
all_pass = False
print("")
if all_pass:
print("✅ All checks passed")
return 0
else:
print("❌ Some checks failed")
return 1
if __name__ == "__main__":
sys.exit(main())
+7 -6
View File
@@ -11,11 +11,13 @@ set -euo pipefail
# ── CT ID → PVE Node mapping (maintained HERE, not in prose contracts) ──
declare -A CT_NODES=(
# amdpve (192.168.68.15)
[111]=amdpve # tdunna
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
[112]=amdpve # tanko
[113]=amdpve # baggy
[115]=amdpve # scottdenya
[120]=amdpve # adguard2 (added 2026-09-12)
# minipve (192.168.68.12)
[100]=minipve # abiba (was hwepve)
[102]=minipve # adguard (was acerpve)
[104]=minipve # authentik
[110]=minipve # gitea
@@ -25,17 +27,16 @@ declare -A CT_NODES=(
[106]=storepve # ra-h-os
[107]=storepve # proxmox-backup
[108]=storepve # media
[111]=storepve # tdunna (was amdpve — corrected 2026-09-12; live per pvesh)
[117]=storepve # zulip
[118]=storepve # jdownloader
# acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
[100]=minipve # abiba (was hwepve)
[105]=minipve # kagentz (was hwepve)
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
#
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090)
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070)
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH)
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090) [QEMU VM on acerpve]
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070) [QEMU VM on ocupve]
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH) [QEMU VM on storepve]
)
# Each node must be root-accessible via SSH hostname
+7 -5
View File
@@ -47,17 +47,19 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (5 nodes):**
- amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
- minipve (192.168.68.12): abiba, kagentz, adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
- amdpve (192.168.68.15): kagentz, tanko, baggy, scottdenya, adguard2
- minipve (192.168.68.12): abiba, adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
- acerpve (192.168.68.9): llm-gpu
- ocupve (192.168.68.5): ocu-llm
**CT IDs (verified 2026-07-24 against PVE API):**
**CT IDs (verified 2026-09-12 against PVE API):**
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
107:pbs 108:media 110:gitea 111:tdunna 112:tanko
113:baggy 115:scottdenya 116:syslog-api 117:zulip
118:jdownloader 119:infisical-vault
118:jdownloader 119:infisical-vault 120:adguard2
**CT 111 (tdunna, 192.168.68.129) is REPORT-ONLY — Theo's box; alert only, never garbage-collect.**
**NO CT 122, CT 123, or .19 exist in the cluster.**
+191
View File
@@ -0,0 +1,191 @@
"""Regression tests for the 2026-09-12 retired-alias sweep in audit-hermes-config.py.
WHY THIS FILE EXISTS: the executable audit pushed agent configs toward a DEAD alias.
Rule 8 required `auxiliary.vision.model == "gpu-light"` and
`auxiliary.web_extract.model == "gpu-light"`, but `gpu-light` (and its raw predecessor
`gemma-4-12b`) were retired on 2026-09-12 and now return 400 `Invalid model name`; the
live RTX 5070 alias is `gpu-vision`. A config that adopted the correct canonical alias
therefore FAILED our own audit, so the audit was actively enforcing a broken config.
These tests execute the real CLI (`python3 audit-hermes-config.py <config>`) and assert
observable behaviour — exit code and the emitted rule message — for the live alias and
for both retired names. No network, vault, or SSH access is required.
"""
from __future__ import annotations
import pathlib
import subprocess
import sys
ROOT = pathlib.Path(__file__).resolve().parent.parent
AUDIT = ROOT / "audit-hermes-config.py"
BASE = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: {alias}
provider: harness
web_extract:
model: {alias}
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/v1
"""
def _run_config(tmp_path, name, text):
cfg = tmp_path / name
cfg.write_text(text)
proc = subprocess.run(
[sys.executable, str(AUDIT), str(cfg)],
capture_output=True, text=True,
)
return proc.returncode, proc.stdout
def _run(tmp_path, alias):
return _run_config(tmp_path, f"{alias}.yaml", BASE.format(alias=alias))
def test_live_canonical_alias_passes(tmp_path):
"""The RTX 5070 alias that actually resolves must satisfy Rule 8."""
code, out = _run(tmp_path, "gpu-vision")
assert code == 0, out
assert "RESULT: PASS" in out
def test_retired_gpu_light_is_rejected(tmp_path):
"""A config pinned to the retired alias must fail, not pass."""
code, out = _run(tmp_path, "gpu-light")
assert code == 1, out
assert "auxiliary.vision.model must be gpu-vision" in out
assert "RESULT: FAIL" in out
def test_retired_gemma_is_rejected(tmp_path):
"""The retired raw model name must fail Rule 8 as well."""
code, out = _run(tmp_path, "gemma-4-12b")
assert code == 1, out
assert "auxiliary.vision.model must be gpu-vision" in out
assert "RESULT: FAIL" in out
def test_corrected_compression_example_passes(tmp_path):
"""The corrected workaround (vision=gpu-vision, compression=syslog-auto) must PASS."""
code, out = _run(tmp_path, "gpu-vision")
assert code == 0, out
assert "[Rule 7] compression.model must be syslog-auto (got 'syslog-auto')" in out
assert "[Rule 7] auxiliary.compression.model must be syslog-auto (got 'syslog-auto')" in out
assert "RESULT: PASS" in out
def test_retired_alias_in_delegation_is_rejected(tmp_path):
"""delegation.model has no dedicated value rule, so a retired name there used to PASS."""
code, out = _run_config(
tmp_path,
"delegation-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(
"delegation:\n provider: harness",
"delegation:\n provider: harness\n model: gpu-light",
),
)
assert code == 1, out
assert "delegation.model = 'gpu-light' is retired" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_custom_providers_is_rejected(tmp_path):
"""custom_providers[*].model is model-bearing; a retired name there must fail."""
code, out = _run_config(
tmp_path,
"custom-provider-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(
" - name: harness\n key_env: LITELLM_API_KEY",
" - name: harness\n model: gpu-light\n key_env: LITELLM_API_KEY",
),
)
assert code == 1, out
assert "custom_providers[0].model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
def test_retired_raw_name_fails(tmp_path):
"""Retired raw names no longer resolve (400), so they fail; the 2026-09-12 registry change moved qwen3.6-27B-code from raw-but-live to non-resolving."""
code, out = _run_config(
tmp_path,
"retired-qwen.yaml",
BASE.format(alias="gpu-vision").replace(
"delegation:\n provider: harness",
"delegation:\n provider: harness\n model: qwen3.6-27B-code",
),
)
assert code != 0, out
assert "delegation.model = 'qwen3.6-27B-code' is retired and no longer resolves" in out
assert "use gpu-dense" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_fallback_providers_is_rejected(tmp_path):
"""fallback_providers.model is model-bearing; a retired name there must fail."""
code, out = _run_config(
tmp_path,
"fallback-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(" model: deepseek-v4-flash", " model: gpu-light"),
)
assert code == 1, out
assert "fallback_providers.model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_x_search_is_rejected(tmp_path):
"""x_search.model was previously not enumerated; the derivation must catch it."""
code, out = _run_config(
tmp_path,
"x-search-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(
"delegation:\n provider: harness",
"delegation:\n provider: harness\nx_search:\n model: gpu-light",
),
)
assert code == 1, out
assert "x_search.model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
"""A nested auxiliary sub-block outside the named three must still be derived."""
code, out = _run_config(
tmp_path,
"nested-aux-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(
" compression:\n model: syslog-auto\n provider: harness\ndelegation:",
" compression:\n model: syslog-auto\n provider: harness\n"
" tasks:\n summarize:\n model: gpu-light\ndelegation:",
),
)
assert code == 1, out
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
+185
View File
@@ -0,0 +1,185 @@
"""Regression tests for the disk-gc report-only gate (CT 111 / tdunna / .129).
WHY THIS FILE EXISTS: `disk-gc-threat-response.prose.md` defined AMBER as "GC scheduled
for next run" and its Execution loop called `gc-executor` for EVERY threat, with no
guest-level exclusion. CT 111 (tdunna, 192.168.68.129) belongs to Theo and is
report-only per the captain (2026-08-17, re-confirmed 2026-09-10) — so a single AMBER
reading on that guest would have scheduled GC commands (apt clean, journal vacuum,
log/tmp deletion, snap removal) against someone else's box. The only marker was
frontmatter `report_only_agents`, which names an AGENT while the scan unit is a GUEST.
These tests execute the real planner (`scripts/disk-gc-plan.py`) and assert observable
behaviour: an excluded guest never produces a `gc-executor` action at any level, while
our own guests still do.
"""
from __future__ import annotations
import json
import pathlib
import subprocess
import sys
ROOT = pathlib.Path(__file__).resolve().parent.parent
PLAN = ROOT / "scripts" / "disk-gc-plan.py"
def _plan(scan, tmp_path):
scan_file = tmp_path / "scan.json"
scan_file.write_text(json.dumps(scan))
proc = subprocess.run(
[sys.executable, str(PLAN), "--scan", str(scan_file), "--json"],
capture_output=True, text=True,
)
assert proc.returncode == 0, proc.stderr
return json.loads(proc.stdout)
def _actions_for(plan, target):
return [row for row in plan if str(row["target"]) == str(target)]
def test_excluded_guest_never_gets_gc_at_any_level(tmp_path):
"""CT 111 at AMBER, RED and CRITICAL — always report-only, never gc-executor."""
for pct, level in ((84, "AMBER"), (90, "RED"), (97, "CRITICAL")):
plan = _plan([{"id": 111, "hostname": "tdunna", "ip": "192.168.68.129",
"usage_pct": pct}], tmp_path)
rows = _actions_for(plan, 111)
assert rows, f"CT 111 must still be reported at {level}"
assert rows[0]["level"] == level
assert rows[0]["action"] == "report-only", rows
assert not any(r["action"] == "gc-executor" for r in rows)
def test_exclusion_matches_on_any_identity_key(tmp_path):
"""The gate is keyed on guest/host, so id, ct, ctid, hostname or IP all match."""
for entry in ({"id": 111, "usage_pct": 95},
{"ct": 111, "usage_pct": 95},
{"ctid": 111, "usage_pct": 95},
{"hostname": "tdunna", "usage_pct": 95},
{"ip": "192.168.68.129", "usage_pct": 95}):
plan = _plan([entry], tmp_path)
assert all(r["action"] == "report-only" for r in plan), (entry, plan)
def test_exclusion_matches_encoded_identities(tmp_path):
"""A differently-encoded CT 111 id must not slip past the gate to gc-executor."""
encodings = ({"id": 111.0},
{"id": "111.0"},
{"id": "lxc/111"},
{"id": "qemu/111"},
{"id": "0111"},
{"id": " 111 "})
for alias in encodings:
plan = _plan([{**alias, "usage_pct": 97}], tmp_path)
assert plan, alias
assert plan[0]["action"] == "report-only", (alias, plan)
def test_our_own_guests_still_get_gc(tmp_path):
"""acerpve .9 and amdpve .15 are ours — they must still be acted on."""
plan = _plan([{"hostname": "acerpve", "ip": "192.168.68.9", "usage_pct": 77},
{"hostname": "amdpve", "ip": "192.168.68.15", "usage_pct": 76}], tmp_path)
assert len(plan) == 2
assert all(r["action"] == "gc-executor" for r in plan), plan
def test_below_threshold_emits_nothing(tmp_path):
"""GREEN guests produce no action at all."""
assert _plan([{"id": 111, "usage_pct": 40}], tmp_path) == []
def test_agent_name_alone_does_not_gate_a_guest(tmp_path):
"""An agent-name marker must not be the gate: an unrelated guest still gets GC."""
plan = _plan([{"id": 999, "hostname": "koby", "usage_pct": 95}], tmp_path)
assert plan and plan[0]["action"] == "gc-executor"
def test_unidentified_threat_fails_closed(tmp_path):
"""A threshold-crossing entry with no recognized identity must not schedule GC."""
plan = _plan([{"usage_pct": 97}], tmp_path)
assert plan, "an unidentified threat must still be reported"
assert plan[0]["action"] == "report-only", plan
assert plan[0]["reason"], plan
def test_exclusion_entry_without_identity_fails_closed(tmp_path):
"""A mis-typed exclusion entry must break the run, never silently disable the gate."""
contract = tmp_path / "broken.prose.md"
contract.write_text(
"```yaml\n"
"report_only_guests:\n"
" - node: storepve\n"
" reason: \"typo - no guest identity\"\n"
"```\n"
)
scan_file = tmp_path / "scan.json"
scan_file.write_text(json.dumps([{"ct": 111, "usage_pct": 97}]))
proc = subprocess.run(
[sys.executable, str(PLAN), "--scan", str(scan_file),
"--contract", str(contract), "--json"],
capture_output=True, text=True,
)
assert proc.returncode != 0, proc.stdout
assert "gc-executor" not in proc.stdout
assert "identity" in proc.stderr.lower(), proc.stderr
def test_empty_exclusion_block_fails_closed(tmp_path):
"""An emptied report_only_guests list must break the run, not disable the gate."""
contract = tmp_path / "empty.prose.md"
contract.write_text("```yaml\nreport_only_guests: []\n```\n")
scan_file = tmp_path / "scan.json"
scan_file.write_text(json.dumps([{"ct": 111, "usage_pct": 97}]))
proc = subprocess.run(
[sys.executable, str(PLAN), "--scan", str(scan_file),
"--contract", str(contract), "--json"],
capture_output=True, text=True,
)
assert proc.returncode != 0, proc.stdout
assert "gc-executor" not in proc.stdout
assert "report-only gate" in proc.stderr.lower(), proc.stderr
def _plan_with_contract(scan, contract_text, tmp_path, name):
contract = tmp_path / name
contract.write_text(contract_text)
scan_file = tmp_path / f"scan-{name}.json"
scan_file.write_text(json.dumps(scan))
proc = subprocess.run(
[sys.executable, str(PLAN), "--scan", str(scan_file),
"--contract", str(contract), "--json"],
capture_output=True, text=True,
)
assert proc.returncode == 0, proc.stderr
return json.loads(proc.stdout)
def test_gate_is_read_from_the_contract_block(tmp_path):
"""The gate is data-driven by the contract block: the planner excludes the guest
when the block names it and acts on it when the block does not. Executes the real
planner interface against both fixtures so the behaviour change is observable."""
scan = [{"id": 111, "hostname": "tdunna", "ip": "192.168.68.129", "usage_pct": 95}]
with_gate = (
"```yaml\n"
"report_only_guests:\n"
" - guest: 111\n"
" hostname: tdunna\n"
" ip: 192.168.68.129\n"
" reason: \"fixture reason\"\n"
"```\n"
)
without_gate = (
"```yaml\n"
"report_only_guests:\n"
" - guest: 999\n"
" reason: \"fixture excludes a different guest\"\n"
"```\n"
)
gated = _plan_with_contract(scan, with_gate, tmp_path, "gated.prose.md")
assert gated[0]["action"] == "report-only", gated
assert gated[0]["reason"] == "fixture reason", gated
ungated = _plan_with_contract(scan, without_gate, tmp_path, "ungated.prose.md")
assert ungated[0]["action"] == "gc-executor", ungated
assert ungated[0]["action"] != gated[0]["action"]
+32 -4
View File
@@ -127,20 +127,48 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
## Execution
### Liveness rule (scoped)
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
### Step 1: Zulip Server Liveness
```bash
curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY'
# Probe the Zulip API (authenticated, any HTTP status = ALIVE)
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code"
fi
```
Expected: `200`. If not → mark `zulip_server_status: "down"`, skip per-platform checks, alert.
Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault.
### Step 2: Platform A — pi (Abiba, localhost)
**A1: Health Endpoint**
Fetch `http://localhost:9200/health` as JSON. Check:
```bash
# Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE)
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health)
echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Abiba extension http://127.0.0.1:9200/health -> $code"
fi
```
Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.**
Check the JSON payload:
| Field | Healthy | Critical |
|-------|---------|----------|