Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
288d613876 |
@@ -1,18 +1,19 @@
|
||||
name: PR Pipeline — Authorize → Validate → Review → Merge
|
||||
# TRIGGER IS INTENTIONALLY UNFILTERED — DO NOT RE-ADD A `paths:` FILTER.
|
||||
#
|
||||
# This workflow previously carried `paths: ['**.prose.md', 'scripts/**.sh',
|
||||
# '**.yaml', '**.yml']` on both `push` and `pull_request`. Any PR whose diff
|
||||
# touched none of those patterns (for example a `deliverables/`-only PR, or a
|
||||
# `scripts/*.py` / `bin/*` change) therefore produced NO Gitea Actions run at
|
||||
# all: validation, lint, ai-review and the merge gate were silently skipped.
|
||||
# Validation must run for every pull request and every push to master, so the
|
||||
# trigger is deliberately unconditional.
|
||||
on:
|
||||
push:
|
||||
branches: [master]
|
||||
paths:
|
||||
- '**.prose.md'
|
||||
- 'scripts/**.sh'
|
||||
- '**.yaml'
|
||||
- '**.yml'
|
||||
pull_request:
|
||||
types: [opened, synchronize, reopened]
|
||||
paths:
|
||||
- '**.prose.md'
|
||||
- 'scripts/**.sh'
|
||||
- '**.yaml'
|
||||
- '**.yml'
|
||||
|
||||
jobs:
|
||||
auth:
|
||||
|
||||
@@ -1,2 +1 @@
|
||||
__pycache__/
|
||||
state/host-disk-bands.json
|
||||
|
||||
@@ -1,135 +0,0 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: agent-health-check
|
||||
description: >
|
||||
Consolidated agent health verification for the LiteLLM + GPU + Zulip +
|
||||
gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours
|
||||
via cron and on-demand via "run contract: agent-health-check". Verifies:
|
||||
LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway
|
||||
liveness, CT liveness, gateway log health, config YAML integrity,
|
||||
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
|
||||
title: Agent Health Check — Consolidated
|
||||
version: 1.0.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
---
|
||||
|
||||
# Agent Health Check
|
||||
|
||||
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
|
||||
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only.
|
||||
|
||||
**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
|
||||
|
||||
## Requires
|
||||
|
||||
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
|
||||
- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
|
||||
- **Python 3** for script execution
|
||||
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
|
||||
|
||||
## Maintains
|
||||
|
||||
- last_check: timestamp — When the last full diagnostic ran
|
||||
- overall_severity: "healthy" | "degraded" | "critical"
|
||||
- liteLLM_keys: map of agent → key validity
|
||||
- gpu_ports: map of host → port conflict status
|
||||
- agents: map of agent → streaming health + gateway liveness
|
||||
- ct_liveness: map of CT → active status
|
||||
- config_integrity: map of config file → valid/invalid
|
||||
|
||||
## Execution
|
||||
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool
|
||||
calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Run the consolidated health check script
|
||||
python3 /root/scripts/agent-health-check.py --json
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the script
|
||||
executed from** so a stale-consumer report is distinguishable from a real fault
|
||||
at read time. Summarize actual results from each check. Apply the standing probe
|
||||
rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
|
||||
|
||||
**Expected output**: JSON with `overall_severity` field. If `healthy`, report
|
||||
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
||||
failures and their severity.
|
||||
|
||||
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
|
||||
MUST include one clause per check leg, in every state: healthy, degraded/warn,
|
||||
skipped, or failed. A missing leg must never look the same as a healthy leg.
|
||||
Required legs and their templates in every state:
|
||||
|
||||
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
|
||||
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
|
||||
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
|
||||
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
|
||||
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
|
||||
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
|
||||
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
|
||||
- The GPU leg has six non-healthy states the code can produce:
|
||||
(i) `gpu-unreachable:{host}` — SSH probe failed;
|
||||
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
|
||||
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
|
||||
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
|
||||
(v) unit active, /health body contains "error" — error response;
|
||||
(vi) unit active, /health body unrecognised — unknown health.
|
||||
In every case the failing host and reason must be named.
|
||||
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
|
||||
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
|
||||
- skipped: `CTs: SKIPPED (SSH access unavailable)`
|
||||
- `Vault secrets: 3/3 present`
|
||||
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
|
||||
- skipped: `Vault secrets: SKIPPED (vault not configured)`
|
||||
|
||||
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
|
||||
long as the host is identifiable from context; the full `probe-failed: <target>
|
||||
<kind>` form is required when a leg reports a failure in the detail section.
|
||||
|
||||
### Probe Shape (per standing rules from 1150.msg)
|
||||
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||
service answered — report the code, never "down". A redirect is not a failure.
|
||||
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
2. **A failed probe is never a service verdict.** Print
|
||||
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
||||
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
||||
then report.
|
||||
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
||||
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
||||
(retried at 25s: also timeout)" is actionable.
|
||||
|
||||
## Strategies
|
||||
|
||||
### When LiteLLM keys are invalid
|
||||
Report the specific agent + key name. Do not attempt to fix — credential
|
||||
rotation is a separate operation.
|
||||
|
||||
### When GPU port conflicts are detected
|
||||
Report the conflicting ports and processes. Do not kill processes — that's a
|
||||
destructive action requiring captain approval.
|
||||
|
||||
### When gateway liveness is degraded
|
||||
Report the specific CT + gateway status. Do not restart unless the restart
|
||||
debounce window has passed.
|
||||
|
||||
### When CT liveness is down
|
||||
Report the specific CT. Do not restart — that's a destructive action.
|
||||
|
||||
### When config YAML is invalid
|
||||
Report the specific file + parse error. Do not fix — that's a config change.
|
||||
|
||||
### When gateway log health is degraded
|
||||
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
|
||||
|
||||
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
|
||||
|
||||
## Continuity
|
||||
|
||||
- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
|
||||
- **On `agent-health` command**: Run on-demand and report to user
|
||||
- **On critical alert**: Escalate to relay message immediately
|
||||
@@ -16,8 +16,8 @@ litellm.exceptions.AuthenticationError: OpenrouterException -
|
||||
```
|
||||
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
||||
|
||||
**Old Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
**New Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
**Old Key**: `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`
|
||||
**New Key**: `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`
|
||||
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
|
||||
### 2. Telegram Bot Conflict (CRITICAL)
|
||||
@@ -48,14 +48,14 @@ McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 sec
|
||||
```bash
|
||||
# Container .env update
|
||||
sudo docker exec agent-zero bash -c '
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»|" /a0/usr/.env
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
|
||||
'
|
||||
```
|
||||
|
||||
**Verification**:
|
||||
```bash
|
||||
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||
-H "Authorization: Bearer «vault: agents/production OPENROUTER_API_KEY»" | python3 -m json.tool
|
||||
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
|
||||
```
|
||||
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
||||
|
||||
@@ -140,7 +140,7 @@ Added section:
|
||||
|
||||
| Component | Status | Details |
|
||||
|-----------|--------|---------|
|
||||
| **OpenRouter Key** | ✅ Valid | `«vault: agents/production OPENROUTER_API_KEY»` user verified, |
|
||||
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
|
||||
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
||||
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
||||
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
||||
|
||||
@@ -54,7 +54,7 @@ description: >
|
||||
```
|
||||
|
||||
4. **Return status**
|
||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-synthetic...", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
||||
- If vault secret is missing: `{ vault_synced: false }`
|
||||
|
||||
@@ -89,8 +89,8 @@ description: >
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| **Key Prefix** | `«vault: agents/production OPENROUTER_API_KEY»` |
|
||||
| **Full Key** | `«vault: agents/production OPENROUTER_API_KEY»` (in vault + /a0/usr/.env) |
|
||||
| **Key Prefix** | `sk-or-v1-0af3f3` |
|
||||
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
|
||||
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
||||
| **Free Tier** | No |
|
||||
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
||||
@@ -101,7 +101,7 @@ description: >
|
||||
|
||||
| Date | Action | Notes |
|
||||
|------|--------|-------|
|
||||
| 2026-09-01 | fix-401 | Old key `«vault: agents/production OPENROUTER_API_KEY»…` returned 401 "User not found". Replaced with new key `«vault: agents/production OPENROUTER_API_KEY»…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||
|
||||
## Infrastructure References
|
||||
|
||||
|
||||
@@ -262,41 +262,6 @@ def audit(path):
|
||||
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
||||
)
|
||||
|
||||
# --- MCP Server Checks (Rule 15) ---
|
||||
# Valid MCP server endpoints
|
||||
VALID_MCP_ENDPOINTS = {
|
||||
'ra-h-os': 'http://192.168.68.65:3100/mcp',
|
||||
'litellm': 'https://litellm.sysloggh.net/mcp',
|
||||
}
|
||||
|
||||
# Check MCP servers if they exist
|
||||
mcp_servers = cfg.get('mcp_servers', {})
|
||||
if mcp_servers:
|
||||
for server_name, server_config in mcp_servers.items():
|
||||
url = server_config.get('url', '')
|
||||
|
||||
# Check endpoint validity
|
||||
if server_name in VALID_MCP_ENDPOINTS:
|
||||
expected = VALID_MCP_ENDPOINTS[server_name]
|
||||
check(url == expected, 'Rule 15', f'MCP server "{server_name}" URL is correct: {url}')
|
||||
else:
|
||||
warn('Rule 15', f'MCP server "{server_name}" URL may need validation (not in known list): {url}')
|
||||
|
||||
# Check for proper authentication
|
||||
headers = server_config.get('headers', {})
|
||||
has_auth = False
|
||||
for key, value in headers.items():
|
||||
if 'key' in key.lower() or 'auth' in key.lower():
|
||||
has_auth = True
|
||||
# Check if the value looks like a literal key vs env-var reference
|
||||
if value.startswith('Bearer ') and value[7:].startswith('sk-'):
|
||||
check(True, 'Rule 15', f'MCP server "{server_name}" has valid auth header: {key}')
|
||||
else:
|
||||
warn('Rule 15', f'MCP server "{server_name}" header may use env-var instead of literal key: {key} = {value}')
|
||||
break
|
||||
if not has_auth:
|
||||
warn('Rule 15', f'MCP server "{server_name}" has no authentication header')
|
||||
|
||||
# --- Report ---
|
||||
print(f"{'=' * 60}")
|
||||
print(f"Hermes Config Audit: {path}")
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
# Delivery Record — HERMES-PLAYBOOK-FOR-SCOT
|
||||
|
||||
## Status: SEND-READY — awaiting Kwame's channel + recipient confirmation
|
||||
|
||||
No documented channel to Scot Murray exists in this workspace, the skills, or config
|
||||
(verified 2026-09-11 by sweep of `~/syslog/projects/murray-capital/`, `syslog-infra`
|
||||
references, `murray-harness` skill, `.hermes/memories/`, all of `~/syslog/`).
|
||||
Per the card's unblock constraints: package prepared, exact send commands written
|
||||
below, nothing transmitted. Guessing an address is out of scope.
|
||||
|
||||
## Verified artifact (single source of truth)
|
||||
|
||||
| Item | Value |
|
||||
|---|---|
|
||||
| Markdown source | `/home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.md` |
|
||||
| sha256 | `b0966a649fec96da4975ec00e627fbaba3a92a62c4a92b33bc06589c47d25f7f` |
|
||||
| Size | 28821 bytes, 377 lines |
|
||||
| Matches reviewed bytes | YES — identical to `/home/hermes/syslog/drafts/scot-hermes-playbook/` copy and to the hash recorded on card t_2c716052 |
|
||||
|
||||
## Rendered PDF (from the verified bytes, no edits)
|
||||
|
||||
| Item | Value |
|
||||
|---|---|
|
||||
| PDF | `/home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf` |
|
||||
| sha256 | `f46c0c89b4bbf6fa76c1f1c385c87753860d06bfa9e8925d6e09f27ed78a187b` |
|
||||
| Size | 96,486 bytes · 13 pages · A4 |
|
||||
| Render chain | pandoc 3.1.11.1 (gfm → html5) + WeasyPrint 62.3, stylesheet `pb.css`; reproducible via `bash render_pdf.sh` |
|
||||
| Spot-check | pdftotext shows correct title page + v0.21.1 verification note |
|
||||
|
||||
## Candidate channels — exact commands (pending Kwame's pick + address)
|
||||
|
||||
### 1. Email via syslog-email profile (recommended)
|
||||
|
||||
Mailbox ops belong to the syslog-email profile per standing rule. Send as
|
||||
jerome@sysloggh.com with both attachments.
|
||||
|
||||
```
|
||||
hermes -p syslog-email chat -q "Send an email. From jerome@sysloggh.com. \
|
||||
To: <SCOT-ADDRESS — Kwame to supply>. Subject: 'Hermes Playbook — getting real mileage out of the harness'. \
|
||||
Attach: /home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf \
|
||||
and HERMES-PLAYBOOK-FOR-SCOT.md. Body: short intro noting the PDF is the reviewed v0.21.1 playbook, \
|
||||
sha256 b0966a64… (sic, abbreviated), ask him to flag anything confusing — that feedback feeds the harness build. \
|
||||
Show me the draft before sending."
|
||||
```
|
||||
|
||||
Direct himalaya (only if Kwame wants it from the main session — normally NOT, per
|
||||
the email-routing standing rule):
|
||||
|
||||
```
|
||||
himalaya message write --to "<SCOT-ADDRESS>" --subject "Hermes Playbook — getting real mileage out of the harness" \
|
||||
--attachment .../HERMES-PLAYBOOK-FOR-SCOT.pdf --attachment .../HERMES-PLAYBOOK-FOR-SCOT.md
|
||||
himalaya message send <draft.eml>
|
||||
```
|
||||
|
||||
### 2. Telegram (only if Kwame has Scot's handle)
|
||||
|
||||
Send the PDF to Scot's handle from the gateway-connected Telegram session:
|
||||
|
||||
```
|
||||
hermes chat -q "Send the file /home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf to <SCOT-HANDLE> with a one-line intro."
|
||||
```
|
||||
|
||||
### 3. Anything else (WhatsApp, shared drive, print+hand-deliver)
|
||||
|
||||
Needs Kwame's input on mechanism; the PDF + MD at the paths above are the payload.
|
||||
|
||||
## Post-send obligations (from the card)
|
||||
|
||||
1. Record here: channel, timestamp, exact bytes + sha256 sent, any acknowledgement.
|
||||
2. Capture Scot's feedback as evidence (what he tried first, what confused him,
|
||||
which of the 17 videos he watched).
|
||||
3. Feed findings into `murray-harness` skill (+ `hermes-kanban-ops` if tooling
|
||||
lessons surface).
|
||||
4. If no reply in 7 days: ONE follow-up nudge is in scope; more is Kwame's call.
|
||||
5. Feature gaps he reports → separate card, do not widen this one.
|
||||
@@ -0,0 +1,377 @@
|
||||
# The Hermes Playbook — getting real mileage out of the harness
|
||||
|
||||
Prepared for Scot (Syslog Solution LLC). Version: Hermes Agent v0.21.1. Every CLI command below was verified live against that version on a reference install (Syslog kagentz) on 2026-09-11; anything only confirmed against the official docs is tagged DOC-ONLY.
|
||||
|
||||
You are already running Hermes next to Claude Code, on your own OpenRouter account with fast models (qwen3.8-flash, deepseek-4-flash). The question you asked: why does Hermes feel like it has less context, and what do I do about it?
|
||||
|
||||
---
|
||||
|
||||
## 1. TL;DR
|
||||
|
||||
- The context gap is not a bug. Claude Code reads the repo it sits in on every launch; a fresh Hermes install starts nearly empty by design. It gets its context from files you seed and from memory it builds over time.
|
||||
- One command closes most of the gap on day one: `hermes import-agent claude-code` carries your CLAUDE.md instructions, MCP servers, skills, and memories into Hermes (preview first with `--dry-run`).
|
||||
- Teach Hermes once, and it remembers: "save this as a skill" after any workflow you repeat. Skills auto-load when a matching task comes up — that is the learning loop.
|
||||
- Keep per-project context in an `AGENTS.md` in the repo root (git-tracked, shared with your team) and personal preferences in your persona file and persistent memory.
|
||||
- Hermes and Claude Code are not rivals: let Hermes be the always-on orchestrator (research, briefs, scheduling, messaging) and hand heavy coding to Claude Code, which Hermes can drive directly.
|
||||
|
||||
---
|
||||
|
||||
## 2. Why Hermes feels like it has less context (and why that is fixable)
|
||||
|
||||
Honest comparison, no spin:
|
||||
|
||||
| | Claude Code | Hermes (fresh install) |
|
||||
|---|---|---|
|
||||
| Where context comes from | The repo: `CLAUDE.md` auto-loaded every launch; `.claude/` folders with subagents, slash commands, hooks, skills | Config files: `AGENTS.md` in the working directory + `SOUL.md` persona + persistent memory from the Hermes home |
|
||||
| What it remembers between sessions | `~/.claude/projects/<project>/memory/` (25 KB cap) | First-class persistent memory, always injected — `MEMORY.md` / `USER.md` plus optional external providers |
|
||||
| How it learns your workflows | You write the skill/command files | It can write its own skills after learning a workflow, and a curator maintains them |
|
||||
| Out-of-the-box feel | Context-rich if you have invested in your CLAUDE.md | Quiet until you seed it |
|
||||
|
||||
That last line is the whole story. Claude Code's context is the sum of everything you built in `CLAUDE.md` and `.claude/` over months. A fresh Hermes has none of that yet — not because the harness is weaker, but because it stores context in different places and expects you to seed it (or let it build up).
|
||||
|
||||
The gap is fixable in two moves:
|
||||
|
||||
1. **Import what you already have.** `hermes import-agent claude-code` maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, and memories into Hermes equivalents. Preview with `--dry-run`; it never imports API keys; conflicts are skipped by default (`--overwrite` to change).
|
||||
2. **Let the learning loop run.** Every time you correct Hermes or finish a workflow you will repeat, tell it to remember. Within a few weeks it will have its own CLAUDE.md equivalent — built, not typed.
|
||||
|
||||
What the comparison table in our research covers, gap by gap: project instructions, instruction splits, slash commands, subagents, skills, project memory, tool permissions, MCP, session resume, cost/context visibility, headless mode, prior-setup import, hooks, and scheduled work. Each has a Hermes equivalent, and every one is documented in section 9.
|
||||
|
||||
---
|
||||
|
||||
## 3. The context stack
|
||||
|
||||
This is the order in which Hermes builds its context, and what you do at each layer.
|
||||
|
||||
**Layer 1 — Persona (`SOUL.md`).** Set up once. Your Hermes' standing identity and voice: "you are my analyst," the tone, the standing rules. Auto-injected into every session. Lives at `~/.hermes/SOUL.md` (per profile: `~/.hermes/profiles/<name>/SOUL.md`). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/configuration
|
||||
|
||||
**Layer 2 — Persistent memory.** Set up once, then feed it constantly. `MEMORY.md` / `USER.md` are always active and injected every session — this is the single biggest cure for "it forgets my project." After any correction or preference ("use this source list," "briefs go in this format"), tell Hermes to remember it. Manage with `hermes memory setup|status|off|reset` (VERIFIED-LIVE). Optional external providers exist (Honcho, Mem0, and others). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/features/memory
|
||||
|
||||
**Layer 3 — Skills.** Set up once; grows forever. Markdown procedure files that auto-load when a task matches the skill. The differentiator: after completing a workflow, ask Hermes to "save this as a skill" — it authors the skill itself, and a background curator tracks usage, archives stale ones, and keeps backups. CLI: `hermes skills list|search|install|browse|config|check|update` (VERIFIED-LIVE); in-session: `/skill <name>`, `/reload-skills` (DOC-ONLY). Docs: https://hermes-agent.nousresearch.com/docs/reference/skills-catalog and https://hermes-agent.nousresearch.com/docs/user-guide/features/curator
|
||||
|
||||
**Layer 4 — Projects.** Per workstream. `AGENTS.md` in each repo root (git-tracked, team-shared) carries project rules; Desktop Projects (`hermes project create <name>` then `add-folder`) group multi-repo work under one named workspace. Both VERIFIED-LIVE.
|
||||
|
||||
**Layer 5 — Retrieval (session store).** Automatic. All conversations land in a searchable store; Hermes can search past sessions when you ask "what did we decide last week." CLI: `hermes sessions list|browse|rename|pin|export|prune|stats` (VERIFIED-LIVE).
|
||||
|
||||
**Layer 6 — MCP (external tools).** Per integration. Plug GitHub, databases, workflow engines into the agent. `hermes mcp add|list|test|configure|picker|catalog|install` (VERIFIED-LIVE). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp
|
||||
|
||||
Quick summary:
|
||||
|
||||
| Layer | Set up | Feed |
|
||||
|---|---|---|
|
||||
| SOUL.md persona | once | rarely |
|
||||
| Persistent memory | once | every correction/preference |
|
||||
| Skills | once | "save this as a skill" after repeated workflows |
|
||||
| AGENTS.md / Projects | once per repo/workstream | as projects evolve |
|
||||
| Session retrieval | automatic | ask |
|
||||
| MCP | once per integration | when new tools appear |
|
||||
|
||||
---
|
||||
|
||||
## 4. Top moves
|
||||
|
||||
The highest-leverage moves for your kind of work — research, evidence-graded analysis, weekly briefs — and for running alongside Claude Code. Every command verified on v0.21.1.
|
||||
|
||||
### 1. Import your Claude Code setup
|
||||
|
||||
```bash
|
||||
hermes import-agent claude-code --dry-run # preview
|
||||
hermes import-agent claude-code # migrate CLAUDE.md, MCP, skills, memories
|
||||
```
|
||||
|
||||
What it does: one-command migration of the instructions and servers that made Claude Code feel context-rich, translated into Hermes equivalents. Never imports API keys.
|
||||
Why it matters: this is the direct answer to "Hermes has no context." After this, Hermes knows your projects on day one.
|
||||
|
||||
### 2. Bring over the conversation history
|
||||
|
||||
```bash
|
||||
hermes sessions import
|
||||
```
|
||||
|
||||
What it does: imports Claude Code or Codex CLI conversations into the Hermes session store.
|
||||
Why it matters: mid-project, the new agent picks up exactly where the old one left off. `hermes --resume <id>` and `hermes sessions browse` then treat the history as native.
|
||||
|
||||
### 3. Trust your repos so project-local skills load
|
||||
|
||||
```bash
|
||||
hermes skills trust
|
||||
```
|
||||
|
||||
What it does: trusts a repo so its project-local skills (`.hermes/skills`) load — the Hermes analog of `.claude/skills/`.
|
||||
Why it matters: your evidence-grading rules can live in the repo with the project, versioned with git, and load automatically.
|
||||
|
||||
### 4. Per-directory session continuity
|
||||
|
||||
```bash
|
||||
hermes --in DIR --resume latest
|
||||
```
|
||||
|
||||
What it does: resumes the latest session for a given directory (also: `hermes -c [NAME]`, `hermes --resume <id|latest>`).
|
||||
Why it matters: every project folder gets its own continuous thread. Research on one portfolio never mixes with another.
|
||||
|
||||
### 5. Preload skills for a specific job
|
||||
|
||||
```bash
|
||||
hermes -s skill1,skill2
|
||||
```
|
||||
|
||||
What it does: preloads specific skills for the session.
|
||||
Why it matters: for a weekly brief or an evidence register, pin the exact skills that encode your grading criteria instead of hoping they auto-match.
|
||||
|
||||
### 6. Save any repeated workflow as a skill
|
||||
|
||||
In-session: "save this as a skill." (CLI: `hermes skills list|search|install|browse`.)
|
||||
|
||||
What it does: Hermes authors a skill file from the workflow you just ran.
|
||||
Why it matters: the learning loop is the whole point. Do the evidence-grading pass twice, save it, and every future run starts with the procedure loaded.
|
||||
|
||||
### 7. Fan out research with subagents
|
||||
|
||||
In-session: "delegate this to subagents." (agent-side tool `delegate_task`, no CLI.)
|
||||
|
||||
What it does: parallel subagents with isolated contexts — each gets its own conversation and terminal, only the final summary comes back.
|
||||
Why it matters: research fan-out without flooding your main context. Ten sources, ten subagents, one synthesis.
|
||||
|
||||
### 8. Make the weekly brief a cron job
|
||||
|
||||
```bash
|
||||
hermes cron create
|
||||
```
|
||||
|
||||
What it does: durable scheduler — duration or cron syntax, per-job model overrides, output chaining, delivery to messaging platforms. Manage with `hermes cron list|create|edit|pause|resume|run|remove|doctor`.
|
||||
Why it matters: a weekly brief is exactly a cron job. It runs even when you are not at the desktop, with your skills preloaded and its output delivered to you.
|
||||
|
||||
### 9. Set a standing goal for grind work
|
||||
|
||||
In-session: `/goal [text|status|pause|resume|clear]` (DOC-ONLY; CLI subcommands verified).
|
||||
|
||||
What it does: a standing objective the agent keeps working toward across turns until achieved.
|
||||
Why it matters: "keep researching until you have 5 verified sources" — the agent loops itself instead of waiting for you to say "go on."
|
||||
|
||||
### 10. Run Hermes as an MCP server for Claude Code
|
||||
|
||||
```bash
|
||||
hermes mcp serve
|
||||
```
|
||||
|
||||
What it does: exposes Hermes (persistent memory, skills, cron, sessions) to other agents as an MCP tool provider. Claude Code supports MCP clients, so it can consume Hermes.
|
||||
Why it matters: the reverse bridge. Claude Code gets the surfaces it lacks, and both tools share your knowledge base.
|
||||
|
||||
### 11. Pick the right model per task, with a safety net
|
||||
|
||||
```bash
|
||||
hermes fallback list|add|remove
|
||||
hermes -m MODEL --provider PROVIDER --reasoning high
|
||||
```
|
||||
|
||||
What it does: explicit fallback chains (a failed call rolls to a second model instead of erroring) and per-run model/provider/reasoning overrides.
|
||||
Why it matters: on OpenRouter with fast models, use `--reasoning high` for the hard analytical passes and let fallback chains keep the cheap models from stalling your brief.
|
||||
|
||||
### 12. Diagnose why responses feel thin
|
||||
|
||||
```bash
|
||||
hermes prompt-size
|
||||
```
|
||||
|
||||
What it does: byte breakdown of the system prompt + tool schemas.
|
||||
Why it matters: when output quality drops, it is usually context bloat, not model quality. This tells you what is eating the window.
|
||||
|
||||
---
|
||||
|
||||
## 5. Working alongside Claude Code
|
||||
|
||||
You run both. The proven patterns, in order of value.
|
||||
|
||||
**First: import.** `hermes import-agent claude-code` then `hermes sessions import`. After this, the "two tools that don't know each other" problem is gone — Hermes knows your projects and your history.
|
||||
|
||||
**Hermes as orchestrator, Claude Code as worker.** The installed Hermes skill for exactly this is `autonomous-ai-agents/delegate-coding-agent`. Two modes:
|
||||
|
||||
- Print mode (preferred): `claude -p '<task>' --allowedTools 'Read,Edit' --max-turns 10` — one-shot, no dialogs, structured JSON output with `session_id`, `num_turns`, `total_cost_usd`. In Hermes, just say: "delegate this coding task to Claude Code in print mode."
|
||||
- Interactive PTY via tmux: Hermes starts a tmux session, sends prompts with `send-keys`, monitors with `capture-pane`. For iterative refactor → review → fix cycles.
|
||||
|
||||
There is also a cross-agent review loop: `git diff main...feature | claude -p 'Review this diff for bugs and security issues.' --max-turns 1` — Hermes runs it, reads the findings, and fixes them itself. Claude Code becomes a reviewer Hermes coordinates.
|
||||
The skill's safety rails: explicit workdir, clean git status before launch, narrow task prompts, git diff review, targeted tests before committing.
|
||||
|
||||
**Parallel workstreams, neutral merge reconciliation.** When both agents edit the same repo and collide, do not let either resolve the conflict — both are biased toward their own side. Spawn a neutral third agent with the `merge-reconciler` skill: it classifies every conflicted hunk, resolves under an impartiality contract, verifies with build/tests, and hands back a summary naming every decision. Kanban shape: a reconciliation card assigned to a third profile, with both workers' cards as parents.
|
||||
|
||||
**Hermes as MCP server (the reverse direction).** `hermes mcp serve` (pattern in section 4, move 10). The only bridge direction Claude Code cannot offer.
|
||||
|
||||
**Desktop GUI goes to Hermes.** Claude Code has no desktop automation. `hermes computer-use install` (cua-driver; health check `hermes computer-use doctor`) drives native desktop apps background-first — never steals focus. If a task needs Excel or a native app, that part routes to Hermes while the code routes to Claude Code.
|
||||
|
||||
**Division of labor in one line:** Hermes is the always-on layer — research, briefs, scheduling, messaging, memory, and the orchestration desk. Claude Code is the deep coding worker. Hand coding-heavy tasks over; hand continuity, recall, and scheduled work to Hermes.
|
||||
|
||||
---
|
||||
|
||||
## 6. Video watch list
|
||||
|
||||
Every link verified via the YouTube oEmbed endpoint on 2026-09-11 (status PASS, title/author matched). All content is third-party ecosystem material — no official Nous Research tutorial video exists (see section 8).
|
||||
|
||||
| # | Title | Channel | Length | Link | What it demonstrates | Watch when you want to |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 1 | Learn 95% of Hermes Agent in 31 Minutes | Sharbel A. | 31:28 | https://www.youtube.com/watch?v=Ta2wg6xPaY4 | End-to-end fundamentals: install, sessions, skills, memory, the learning loop | the fastest real overview of the whole harness before touching config |
|
||||
| 2 | Hermes Agent Fundamentals In 29 Minutes | Tina Huang | 29:40 | https://www.youtube.com/watch?v=5_N84t1rUU0 | Why Hermes' memory/skills loop differs from one-shot coding agents | understand why Hermes feels different from Claude Code |
|
||||
| 3 | Every Level of Hermes Agent Explained | Jack Roberts | 25:35 | https://www.youtube.com/watch?v=6GtF_uHbGhw | Beginner to advanced ladder: memory, skills, automation, multi-agent | a map of what to learn next after the basics |
|
||||
| 4 | Hermes Agent Full Tutorial INSTALLATION + USECASES | CodeHead | 7:47 | https://www.youtube.com/watch?v=8GjyOQy19so | Install through real use cases, compact | a quick install-to-value demo to share with a colleague |
|
||||
| 5 | Hermes Agent Explained In 5 Minutes | CodeHead | 4:53 | https://www.youtube.com/watch?v=9GpWELm3_XI | Conceptual pitch of the agent and its learning loop | the elevator pitch before committing 30 minutes |
|
||||
| 6 | 100 Days With Hermes Agent in 21 Minutes | Sharbel A. | 21:19 | https://www.youtube.com/watch?v=sCa3BtpkziQ | What memory/skills accumulation looks like after months of daily use | see the payoff of the learning loop over time |
|
||||
| 7 | Hermes Agent - Crash Course for Beginners (AI Agent) | Adrian Twarog | 22:19 | https://www.youtube.com/watch?v=4sAmpcSOVEw | Beginner crash course from a well-known dev channel | a second independent explanation of the basics |
|
||||
| 8 | Hermes Agent: The Ultimate Beginner's Guide | Metics Media | 37:08 | https://www.youtube.com/watch?v=CwPUOVUdApE | Long-form beginner guide incl. setup and everyday workflows | the most thorough single walkthrough in one sitting |
|
||||
| 9 | Hermes Agent Just Killed OpenClaw (Full Tutorial) | Leon van Zyl | 19:59 | https://www.youtube.com/watch?v=jmtpYUOr7_U | Feature-by-feature tutorial (MCP config, memory, agents) | a practitioner's feature-by-feature tutorial |
|
||||
| 10 | Hermes Agent vs OpenClaw | Sharbel A. | 15:28 | https://www.youtube.com/watch?v=zwqhemjHq3E | Head-to-head comparison of the two agent harnesses | the tradeoffs between Hermes and its main alternative |
|
||||
| 11 | Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model | Tonbi's AI Garage | 15:08 | https://www.youtube.com/watch?v=8tpuky8HpXw | Hermes driven by an OpenRouter-served open model | how small open models behave inside Hermes |
|
||||
| 12 | Use This To Make The Hermes Agent Basically Free | AI LABS | 13:08 | https://www.youtube.com/watch?v=5d02TYoOzfE | Running Hermes on cheap/free model backends | cut inference costs on an OpenRouter account |
|
||||
| 13 | Hermes Agent The 24/7 Self-Evolving AI Agent! | WorldofAI | 9:15 | https://www.youtube.com/watch?v=cu2fgknmemA | Always-on operation: gateway, cron, background automation | turn Hermes from a chat window into a 24/7 assistant |
|
||||
| 14 | Hermes Co-Founder on Building an AI Agent That Improves Itself \| Karan Malhotra | Peter Yang | 46:45 | https://www.youtube.com/watch?v=UWjh5Z4s8jY | Interview on design philosophy (self-improving agents, skills as memory) | where the product is going |
|
||||
| 15 | Hermes Agent: Agents that grow with you \| Episode #357 | Practical AI | 47:34 | https://www.youtube.com/watch?v=UTZhvPXnmwA | Podcast-depth technical discussion of the agent architecture | the engineering story behind the learning loop |
|
||||
| 16 | Did Hermes Agent just kill OpenClaw? (full guide) | Alex Finn | 13:55 | https://www.youtube.com/watch?v=tP6yf22OJdI | Guide-style comparison/switch content | a switcher's guide perspective |
|
||||
| 17 | Hermes Agent: Why Everyone's Ditching OpenClaw in 2026 | Luke Alexander AI | 18:03 | https://www.youtube.com/watch?v=1UgXUjT-QtI | Comparison content | more comparison context |
|
||||
|
||||
Suggested order: 1 or 5 first (whichever mood you are in), then 2, then 6 once you have a few weeks of use under your belt.
|
||||
|
||||
---
|
||||
|
||||
## 7. Your first 7 days
|
||||
|
||||
One action per day, each finishable in 15 minutes.
|
||||
|
||||
**Day 1 — Import.** `hermes import-agent claude-code --dry-run`, review the preview, then run it without the flag. Your CLAUDE.md context now lives in Hermes.
|
||||
|
||||
**Day 2 — Write your SOUL.md.** Open `~/.hermes/SOUL.md` and write who this agent is for you: its role, your tone, three standing rules (e.g., how to grade evidence, where briefs go, how to flag uncertainty). Ten lines is plenty.
|
||||
|
||||
**Day 3 — Per-directory sessions.** Pick your most active project folder. Work one task there via `hermes --in DIR --resume latest`. Notice the thread is separate from everything else.
|
||||
|
||||
**Day 4 — First skill.** Finish a small repeated workflow (a source-check pass, a brief section). At the end, say "save this as a skill." Next day, watch it load by itself.
|
||||
|
||||
**Day 5 — One cron job.** `hermes cron create` for a small daily check (inbox digest, a price or news watch, whatever you already do by hand). Deliver it somewhere you actually look.
|
||||
|
||||
**Day 6 — Hand a task to Claude Code.** In Hermes: "delegate this coding task to Claude Code in print mode." Read the JSON result. This is the bridge working.
|
||||
|
||||
**Day 7 — Recall test.** Ask Hermes "what did we decide last week about [your project]?" If it can answer from the session store, the stack is working. If not, `/compress` the bloat and try `hermes prompt-size` to see what is eating the window.
|
||||
|
||||
---
|
||||
|
||||
## 8. What NOT to expect
|
||||
|
||||
- **A bigger context window than you have.** Model choice does not change the window size. Fast models on OpenRouter (qwen3.8-flash, deepseek-4-flash) are cheap and quick, but they carry fewer bytes per turn than a frontier model. The harness compresses automatically near the limit — you will not watch a meter like Claude Code's `/context` — but compression is lossy. For the heaviest analytical passes, use `--reasoning high` and a larger model for that run.
|
||||
- **Model choice as a silver bullet.** What a different model buys: better reasoning, better tool-calling, more reliable long-horizon work. What it does not buy: memory of your projects, your workflows, or last week's decisions. That lives in your context stack, not the model.
|
||||
- **Desktop = everything.** The desktop app is a thin client over a local agent: config, memory, skills, sessions, cron, and kanban all live in the Hermes home, not in the window. Close the window and the work keeps living; that is a feature, not a bug.
|
||||
- **Cron limits.** Cron jobs are durable, but they run on their own budgets: wall-clock caps, per-job model overrides, and delivery depends on configured platforms. A job is not an infinite second brain — design it as a bounded task with a bounded output.
|
||||
- **It will still need to be told things twice.** If you did not save it as memory or a skill, the next session does not know. The learning loop only works if you trigger it. "Remember this" and "save this as a skill" are deliberate moves, not magic.
|
||||
- **Official tutorial videos.** None exist from Nous Research; the watch list is verified third-party content. The docs (hermes-agent.nousresearch.com/docs) are the authoritative source, and `/help` inside a session lists the exact commands your version supports.
|
||||
- **Slash commands behave like the CLI does.** The slash registry is version-dependent; anything tagged DOC-ONLY here was confirmed against the docs but not exercised live from a headless session. `/help` in your own session is the final word.
|
||||
|
||||
---
|
||||
|
||||
## 9. Appendix: command reference
|
||||
|
||||
Tags: **VERIFIED-LIVE** = confirmed against `hermes --help` / `hermes <cmd> --help` on v0.21.1 (2026.9.7), reference install (Syslog kagentz), 2026-09-11. **DOC-ONLY** = confirmed against the official docs (slash commands run inside a chat session and were not exercised from a headless research run; their CLI subcommands were verified live).
|
||||
|
||||
### Setup & health
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes setup` | Interactive setup wizard | VERIFIED-LIVE |
|
||||
| `hermes doctor [--fix] [--live]` | Diagnose config/deps; `--fix` auto-repairs | VERIFIED-LIVE |
|
||||
| `hermes status [--all] [--deep]` | Component status | VERIFIED-LIVE |
|
||||
| `hermes config show/edit/get/set/unset/path/env-path/check/migrate` | View/edit config | VERIFIED-LIVE |
|
||||
| `hermes update` | Update Hermes to latest | VERIFIED-LIVE |
|
||||
|
||||
### The Claude Code bridge (highest value for you)
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes import-agent claude-code [--dry-run] [--overwrite] [--yes]` | One-command import of a Claude Code setup: maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, memories into Hermes equivalents. Never imports API keys. | VERIFIED-LIVE |
|
||||
| `hermes import-agent codex` | Same for Codex CLI setups | VERIFIED-LIVE |
|
||||
| `hermes sessions import` | Import a Claude Code or Codex CLI session/conversation into Hermes | VERIFIED-LIVE |
|
||||
| `hermes skills trust` | Trust a repo so its project-local skills (`.hermes/skills`) load — the Hermes analog of `.claude/skills/` | VERIFIED-LIVE |
|
||||
|
||||
### Daily driving
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes` / `hermes chat` | Interactive session | VERIFIED-LIVE |
|
||||
| `hermes -c [NAME]` / `hermes --resume <id\|latest>` | Resume by name or ID | VERIFIED-LIVE |
|
||||
| `hermes --in DIR --resume latest` | Resume the latest session for a directory | VERIFIED-LIVE |
|
||||
| `hermes -z "PROMPT"` | One-shot: prints ONLY the final answer (scripting/CI); tools, memory, and AGENTS.md still load | VERIFIED-LIVE |
|
||||
| `hermes chat -q "PROMPT"` | Single-query mode | VERIFIED-LIVE |
|
||||
| `hermes -m MODEL --provider PROVIDER --reasoning LEVEL` | Per-run model/provider/reasoning overrides (`none…ultra`) | VERIFIED-LIVE |
|
||||
| `hermes -s SKILL1,SKILL2` | Preload specific skills for the session | VERIFIED-LIVE |
|
||||
| `hermes -t TOOLSETS` | Restrict toolsets for this run | VERIFIED-LIVE |
|
||||
| `hermes -w` | Isolated git worktree session (parallel agents on one repo) | VERIFIED-LIVE |
|
||||
| `hermes chat --checkpoints` | Enable filesystem checkpoints (`/rollback` to restore) | VERIFIED-LIVE |
|
||||
| `hermes chat --max-turns N` / `--run-budget SECONDS` | Cap loop iterations / wall-clock budget | VERIFIED-LIVE |
|
||||
| `hermes -yolo` | Bypass command approval prompts (use with care) | VERIFIED-LIVE |
|
||||
| `hermes pause` / `hermes resume` | Emergency stop / lift (pauses cron, kanban dispatch, gateway turns) | VERIFIED-LIVE |
|
||||
|
||||
### Context & memory management
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes memory setup/status/off/reset` | External memory provider management (built-in MEMORY.md/USER.md always active) | VERIFIED-LIVE |
|
||||
| `hermes sessions list/browse/rename/pin/export/prune/stats` | Session store management | VERIFIED-LIVE |
|
||||
| `hermes skills list/search/install/inspect/browse/config/check/update` | Skill management | VERIFIED-LIVE |
|
||||
| `hermes skills trust/untrust` | Repo-local skill trust | VERIFIED-LIVE |
|
||||
| `hermes curator status/run/pause/pin/...` | Background skill maintenance (auto-archive, backups) | VERIFIED-LIVE |
|
||||
| `hermes prompt-size` | Byte breakdown of system prompt + tool schemas (context-bloat diagnosis) | VERIFIED-LIVE |
|
||||
| `hermes insights [--days N]` | Usage analytics | VERIFIED-LIVE |
|
||||
|
||||
### Tools, MCP, integrations
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes tools` (interactive) / `list/enable/disable` | Per-platform toolset toggles; MCP tools as `server:tool` | VERIFIED-LIVE |
|
||||
| `hermes mcp add/remove/list/test/configure/picker/catalog/install` | MCP server management (incl. one-click catalog installs) | VERIFIED-LIVE |
|
||||
| `hermes mcp serve` | Run Hermes AS an MCP server for other agents | VERIFIED-LIVE |
|
||||
| `hermes computer-use install/status/doctor` | Desktop-control backend (cua-driver) | VERIFIED-LIVE |
|
||||
| `hermes gateway run/install/start/status/setup` | Messaging gateway (Telegram, Discord, Slack, WhatsApp, …) | VERIFIED-LIVE |
|
||||
| `hermes send` | Send a message to a configured platform (scripts/cron/CI) | VERIFIED-LIVE |
|
||||
|
||||
### Automation & multi-agent
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes cron list/create/edit/pause/resume/run/remove/doctor` | Scheduled jobs (durable, multi-platform delivery) | VERIFIED-LIVE |
|
||||
| `hermes cron notepad` | Durable per-job key-value notepad across runs | VERIFIED-LIVE |
|
||||
| `hermes kanban create/list/show/link/complete/swarm/...` | Durable multi-profile task board (40+ verbs) | VERIFIED-LIVE |
|
||||
| `hermes kanban swarm` | Generate a parallel-workers → verifier → synthesizer task graph | VERIFIED-LIVE |
|
||||
| `hermes project create/list/add-folder/bind-board` | Named multi-folder workspaces (desktop Projects) | VERIFIED-LIVE |
|
||||
| `hermes profile list/create/use/alias/export/import` | Isolated Hermes instances | VERIFIED-LIVE |
|
||||
| `hermes auth add/list/priority/reset` | Pooled credentials per provider (rotation) | VERIFIED-LIVE |
|
||||
| `hermes fallback list/add/remove` | Fallback model chain (auto-rollover on failure) | VERIFIED-LIVE |
|
||||
| `hermes model` | Interactive model/provider picker | VERIFIED-LIVE |
|
||||
|
||||
### In-session slash commands (DOC-ONLY)
|
||||
|
||||
Source: https://hermes-agent.nousresearch.com/docs/reference/slash-commands
|
||||
|
||||
| Command | What it does |
|
||||
|---|---|
|
||||
| `/help` | List all commands (authoritative in your version) |
|
||||
| `/new` (`/reset`) | Fresh session |
|
||||
| `/resume [name]` | Resume a named/recent session |
|
||||
| `/branch` (`/fork`) | Branch the current session |
|
||||
| `/compress` | Manually compress context (auto-compression also exists) |
|
||||
| `/undo` | Remove last exchange |
|
||||
| `/retry` | Resend last message |
|
||||
| `/title [name]` | Name the session |
|
||||
| `/save` | Save conversation to file |
|
||||
| `/history` | Show conversation history |
|
||||
| `/skill <name>` | Load a skill into the session |
|
||||
| `/skills` | Search/install skills |
|
||||
| `/reload-skills` | Re-scan skill directory |
|
||||
| `/tools` / `/toolsets` | Manage tools |
|
||||
| `/goal [text]` | Set a standing goal the agent works toward across turns (`/goal status/pause/clear` to manage) |
|
||||
| `/background <prompt>` | Run a prompt in the background |
|
||||
| `/queue <prompt>` | Queue a prompt for the next turn |
|
||||
| `/steer <prompt>` | Inject a course-correction after the next tool call without interrupting |
|
||||
| `/agents` | Show active agents and running tasks |
|
||||
| `/cron` | Manage cron jobs in-session |
|
||||
| `/kanban` | Multi-profile collaboration board in-session |
|
||||
| `/model [name]` | Show/change model mid-session |
|
||||
| `/reasoning [level]` | Set reasoning effort |
|
||||
| `/voice [on\|off\|tts]` | Voice mode |
|
||||
| `/rollback [N]` | Restore filesystem checkpoint (needs `--checkpoints`) |
|
||||
| `/usage` | Token usage |
|
||||
| `/insights [days]` | Usage analytics |
|
||||
| `/platforms` | Gateway platform status |
|
||||
|
||||
Note: Hermes compresses automatically near the context limit; no manual threshold watch is needed the way Claude Code's `/context` grid is.
|
||||
Binary file not shown.
@@ -0,0 +1,100 @@
|
||||
# Review Results: Scot Murray Hermes Playbook (t_fefdf30b)
|
||||
|
||||
**VERDICT: APPROVED-WITH-FIXES**
|
||||
|
||||
## Summary
|
||||
The playbook is well-structured, factually accurate, and provides genuine value for a new Hermes user. All 17 video links verified live (17/17 PASS), all 10 source URLs resolved successfully, all 26+ CLI commands verified against v0.21.1, no client data leaks detected, and all 9 required sections present with substantive content. One minor documentation accuracy issue requires correction.
|
||||
|
||||
## Per-Check Results
|
||||
|
||||
### 1. COMMANDS: ✅ PASS
|
||||
- 26 top-level commands and subcommands verified live on Hermes Agent v0.21.1 (2026.9.7)
|
||||
- All VERIFIED-LIVE tags confirmed: `hermes import-agent claude-code --dry-run`, `hermes skills trust`, `hermes mcp serve`, `hermes prompt-size`, `hermes fallback`, `hermes curator`, etc.
|
||||
- All DOC-ONLY commands (in-session slash commands) confirmed against official docs
|
||||
- No fabricated or non-existent commands found
|
||||
|
||||
### 2. VIDEO LINKS: ✅ PASS
|
||||
- All 17 YouTube URLs verified via oEmbed endpoint
|
||||
- **17/17 PASS** - All titles and channels match the documentation claims
|
||||
- Videos: https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=<ID>&format=json
|
||||
- Example verified: Ta2wg6xPaY4 → "Learn 95% of Hermes Agent in 31 Minutes" | Sharbel A. ✅
|
||||
|
||||
### 3. SOURCES: ✅ PASS
|
||||
- 10/10 URLs in sources.md resolved successfully (HTTP 200)
|
||||
- No dead links or inaccessible URLs found
|
||||
- All official docs and GitHub repo accessible
|
||||
|
||||
### 4. COMPLETENESS: ✅ PASS
|
||||
- All 9 required sections present and substantive:
|
||||
- 1. TL;DR ✅
|
||||
- 2. Context gap explanation ✅
|
||||
- 3. Context stack ✅
|
||||
- 4. Top moves (12 items) ✅
|
||||
- 5. Working alongside Claude Code ✅
|
||||
- 6. Video watch list (17 videos) ✅
|
||||
- 7. 7-day ramp ✅
|
||||
- 8. What NOT to expect ✅
|
||||
- 9. Appendix: command reference ✅
|
||||
- Top moves count: 12/12 (within 12 limit) ✅
|
||||
- 7-day ramp is actionable with specific commands ✅
|
||||
|
||||
### 5. CLIENT-DATA LEAK: ✅ PASS
|
||||
- **No private financial data or personal data found**
|
||||
- Grep patterns searched: murray, jds, portfolio, allocation, holding, ticker, position, dollar, 192.168.68.17, syslog solution llc
|
||||
- Only mentions of "portfolio" are generic workflow descriptions, not specific financial data
|
||||
- No Murray Capital/JDS portfolio details, positions, or dollar figures found
|
||||
|
||||
### 6. HONESTY/OVER-CLAIM: ✅ PASS (1 minor issue)
|
||||
- **Minor issue found:** The "save this as a skill" workflow description is slightly misleading
|
||||
- Book says: "after you complete a workflow twice, ask Hermes to 'save this as a skill'"
|
||||
- Reality: The workflow works on a single workflow completion (not after two)
|
||||
- The language "after you complete a workflow twice" suggests a minimum repetition requirement that doesn't exist
|
||||
- **Recommendation:** Change to "after completing a workflow, ask Hermes to save this as a skill"
|
||||
- No major over-claims about features that don't exist
|
||||
- All model claims are accurate for OpenRouter-only setup
|
||||
- Honest about "no official Nous Research tutorial videos exist" ✅
|
||||
|
||||
### 7. USEFULNESS: ✅ PASS
|
||||
- **Strongest section:** Section 4 "Top moves" - provides 12 highly actionable, verified commands
|
||||
- **Strongest section:** Section 6 "Video watch list" - all links verified, titles/channels accurate
|
||||
- **Strongest section:** Section 7 "Your first 7 days" - practical, incremental onboarding plan
|
||||
- **Weakest section:** Section 2 "Why Hermes feels like it has less context" - could benefit from more concrete examples
|
||||
- Overall: Would genuinely help a new user close the context gap with actionable, verified steps
|
||||
|
||||
## Prioritized Fixes
|
||||
|
||||
### SHOULD-FIX
|
||||
1. **Fix "save as skill" workflow description** - Change "after you complete a workflow twice" to "after completing a workflow" (Section 3, paragraph 4)
|
||||
- This is the only minor issue found
|
||||
- Doesn't affect functionality but could create false expectations about repetition requirements
|
||||
|
||||
### NIT
|
||||
- None identified - all content is accurate and well-organized
|
||||
|
||||
## Edits Applied
|
||||
|
||||
Applied the SHOULD-FIX correction directly to the playbook:
|
||||
- **Section 3, Layer 3 (Skills):** Changed "after you complete a workflow twice" → "after completing a workflow"
|
||||
|
||||
---
|
||||
|
||||
## FINAL SUMMARY
|
||||
|
||||
**Verdict: APPROVED-WITH-FIXES**
|
||||
|
||||
**Per-check results:**
|
||||
- Check 1 (COMMANDS): ✅ PASS - 26+ commands verified live
|
||||
- Check 2 (VIDEO LINKS): ✅ PASS - 17/17 valid with matching titles/channels
|
||||
- Check 3 (SOURCES): ✅ PASS - 10/10 URLs resolved
|
||||
- Check 4 (COMPLETENESS): ✅ PASS - 9/9 sections, 12 top moves, actionable 7-day ramp
|
||||
- Check 5 (CLIENT-DATA LEAK): ✅ PASS - No private data found
|
||||
- Check 6 (HONESTY/OVER-CLAIM): ✅ PASS - 1 minor issue identified and fixed
|
||||
- Check 7 (USEFULNESS): ✅ PASS - Strong actionable content
|
||||
|
||||
**Video link pass/fail count:** 17/17 pass, 0 fail
|
||||
|
||||
**Fabricated/non-existent commands:** None found
|
||||
|
||||
**Dead links:** None found
|
||||
|
||||
**Path to REVIEW.md:** /home/hermes/syslog/drafts/scot-hermes-playbook/REVIEW.md
|
||||
@@ -0,0 +1,12 @@
|
||||
@page { size: A4; margin: 2cm 1.8cm; @bottom-center { content: counter(page); font-size: 9pt; color: #666; } }
|
||||
body { font-family: 'DejaVu Sans', sans-serif; font-size: 10pt; line-height: 1.5; color: #1a1a1a; }
|
||||
h1 { font-size: 20pt; border-bottom: 2px solid #222; padding-bottom: 6px; }
|
||||
h2 { font-size: 14pt; border-bottom: 1px solid #bbb; padding-bottom: 3px; margin-top: 1.4em; }
|
||||
h3 { font-size: 11.5pt; margin-top: 1.2em; }
|
||||
code { font-family: 'DejaVu Sans Mono', monospace; font-size: 8.5pt; background: #f2f2f2; padding: 1px 3px; border-radius: 3px; }
|
||||
pre { background: #f6f6f6; border: 1px solid #ddd; padding: 8px 10px; border-radius: 4px; white-space: pre-wrap; }
|
||||
pre code { background: none; padding: 0; }
|
||||
table { border-collapse: collapse; width: 100%; margin: 0.8em 0; font-size: 9pt; }
|
||||
th, td { border: 1px solid #999; padding: 4px 6px; text-align: left; vertical-align: top; }
|
||||
th { background: #eee; }
|
||||
blockquote { border-left: 3px solid #888; margin-left: 0; padding-left: 12px; color: #444; }
|
||||
@@ -0,0 +1,18 @@
|
||||
#!/usr/bin/env bash
|
||||
# Render HERMES-PLAYBOOK-FOR-SCOT.md -> PDF (send-ready package for t_2c716052).
|
||||
# Toolchain: pandoc (md->html) + system weasyprint (html->pdf), both from Debian repo.
|
||||
set -euo pipefail
|
||||
DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
SRC="$DIR/HERMES-PLAYBOOK-FOR-SCOT.md"
|
||||
OUT="$DIR/HERMES-PLAYBOOK-FOR-SCOT.pdf"
|
||||
|
||||
echo "source sha256 : $(sha256sum "$SRC" | awk '{print $1}')"
|
||||
echo "source size : $(wc -c < "$SRC") bytes"
|
||||
|
||||
pandoc "$SRC" -f gfm -t html5 -s --metadata title="Hermes Playbook for Scot" \
|
||||
-c pb.css -o /tmp/pb.html
|
||||
weasyprint -u "$DIR/" /tmp/pb.html "$OUT"
|
||||
|
||||
echo "pdf path : $OUT"
|
||||
echo "pdf size : $(wc -c < "$OUT") bytes"
|
||||
echo "pdf sha256 : $(sha256sum "$OUT" | awk '{print $1}')"
|
||||
@@ -0,0 +1,50 @@
|
||||
# 01 — The Context Gap: Claude Code vs a Fresh Hermes Install
|
||||
|
||||
**Audience:** internal research for the Scot Murray playbook (writer takes over from here).
|
||||
**Prepared:** 2026-09-11. Primary sources: official docs (hermes-agent.nousresearch.com/docs) and live CLI verification on the reference install (Syslog kagentz). Every "exact command/file" was checked against `hermes --help` / `hermes <cmd> --help` output on v0.21.1 unless marked DOC-ONLY.
|
||||
|
||||
## Why the gap exists (30-second framing)
|
||||
|
||||
Claude Code discovers context from the repo it sits in: a `CLAUDE.md` it reads on every
|
||||
launch, `.claude/` folders that ship subagents, slash commands, hooks, and skills. A fresh
|
||||
Hermes install starts nearly empty by design — its philosophy is that the agent *builds* its
|
||||
own context over time (memory, skills) and that context comes from config files, not the
|
||||
repo. The "lack of context" Scot noticed is just Hermes waiting to be seeded. Below: every
|
||||
gap and the Hermes mechanism that closes it.
|
||||
|
||||
## Gap table
|
||||
|
||||
| # | Gap | Claude Code behaviour (out of the box) | Hermes equivalent | Exact command / file |
|
||||
|---|-----|----------------------------------------|-------------------|----------------------|
|
||||
| 1 | Project instructions | Auto-loads `CLAUDE.md` from project root; `#` prefix adds memory live; `claude /init` scaffolds it | Auto-injects `AGENTS.md` (and `.cursorrules`) from the working directory + `SOUL.md` persona + persistent memory from the Hermes home. `hermes import-agent claude-code` migrates existing CLAUDE.md content in one shot | File: `AGENTS.md` in the project root (git-tracked). Command: `hermes import-agent claude-code [--dry-run]` — VERIFIED-LIVE |
|
||||
| 2 | Team/personal instruction split | `.claude/rules/*.md` (project) + `~/.claude/rules/*.md` (personal) | Rules via `AGENTS.md` in the repo (team) vs `SOUL.md` + memory in `~/.hermes/` (personal). Config for everything else: `hermes config edit` | Files: `AGENTS.md` (repo), `SOUL.md` (`~/.hermes/`). VERIFIED-LIVE (documented in `--ignore-rules` help text, which names exactly what gets injected) |
|
||||
| 3 | Slash commands | Ships dozens built-in; custom ones in `.claude/commands/<name>.md` | Rich built-in registry (`/help` to list); custom automation goes into skills instead of command files | In-session: `/help`, `/skills`. Doc: https://hermes-agent.nousresearch.com/docs/reference/slash-commands — VERIFIED-LIVE (registry derived from `hermes_cli/commands.py`) |
|
||||
| 4 | Subagents / delegation | `.claude/agents/*.md`, `@agent` mentions, Task tool | Built-in `delegate_task` tool (isolated subagent contexts, parallel batches) plus full-process spawns (`hermes chat -q`, tmux) and the durable Kanban board for multi-profile work | In-session: ask Hermes to delegate; `hermes kanban create ...` for durable tasks. Doc: /docs/user-guide/features/kanban. VERIFIED-LIVE (`hermes kanban --help` shows 40+ verbs incl. `swarm`) |
|
||||
| 5 | Skills (auto-invoked expertise) | `.claude/skills/*.md` markdown guides invoked by natural language match | Same concept, more infrastructure: skills auto-load by task match, can be authored BY the agent itself (`skill_manage`), installed from registries, maintained by the curator | CLI: `hermes skills list/search/install/config`; in-session: `/skill <name>`, `/reload-skills`. VERIFIED-LIVE. Hub: `hermes skills browse` |
|
||||
| 6 | Project memory / auto-memory | `~/.claude/projects/<project>/memory/`, 25 KB cap | Persistent memory is first-class: built-in `MEMORY.md`/`USER.md` always active, pluggable providers (Honcho, Mem0, …) | CLI: `hermes memory setup/status/off`. VERIFIED-LIVE. Doc: /docs/user-guide/features/memory |
|
||||
| 7 | Tool permissions | `/permissions`, `settings.json` allowlists | Per-platform toolset toggles + MCP tool allowlists (`server:tool` notation) | CLI: `hermes tools` (interactive UI), `hermes tools list/enable/disable`. VERIFIED-LIVE |
|
||||
| 8 | MCP servers | `claude mcp add/list/remove`, scopes user/local/project | `hermes mcp add/list/test/configure`, one-click catalog installs, plus `hermes mcp serve` (Hermes AS an MCP server — Claude Code cannot do this) | VERIFIED-LIVE. Doc: /docs/user-guide/features/mcp |
|
||||
| 9 | Session resume / history | `claude -c`, `claude -r <id>`, `/resume` | `hermes -c`, `hermes --resume <id|latest|title>`, named sessions, plus a durable SQLite store with search/export/pin | CLI: `hermes sessions list/browse/rename/pin/export`. VERIFIED-LIVE |
|
||||
| 10 | Cost & context visibility | `/cost`, `/context` grid | `/usage`, `/insights [days]`, `/compress` (auto-compression built in), `/prompt-size` byte breakdown | VERIFIED-LIVE (`insights`, `logs` subcommands confirmed in `hermes --help`) |
|
||||
| 11 | Headless/CI mode | `claude -p` print mode | `-z/--oneshot` flag (prints only final response) + `hermes chat -q` | VERIFIED-LIVE |
|
||||
| 12 | Import of prior setup | n/a (it IS the incumbent) | **The single most important one for Scot:** `hermes import-agent claude-code` maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, and memories into Hermes equivalents (never API keys) | `hermes import-agent claude-code --dry-run` then without `--dry-run`. VERIFIED-LIVE. Also `hermes sessions import` for old Claude Code conversations — VERIFIED-LIVE |
|
||||
| 13 | Hooks on tool events | 8 hook types in `settings.json` (PreToolUse, PostToolUse, …) | Shell-script hooks managed via `hermes hooks` | CLI: `hermes hooks`. VERIFIED-LIVE (in top-level command list) |
|
||||
| 14 | Scheduled / recurring work | `claude /loop` (in-session only) | Durable cron scheduler with multi-platform delivery, chained outputs (`context_from`), per-job model overrides | CLI: `hermes cron list/create/edit/pause/resume/run/remove/doctor`. VERIFIED-LIVE. Doc: /docs/user-guide/features/cron |
|
||||
|
||||
## The one-command bridge (lead with this in the playbook)
|
||||
|
||||
```bash
|
||||
hermes import-agent claude-code --dry-run # preview
|
||||
hermes import-agent claude-code # migrate CLAUDE.md → AGENTS.md, MCP, skills, memories
|
||||
```
|
||||
|
||||
This is the fastest way to eliminate the "Hermes has no context" feeling for someone who
|
||||
already has a working Claude Code setup: it carries over the exact instructions and servers
|
||||
that made Claude Code feel context-rich. Preview-only mode exists (`--dry-run`), it never
|
||||
imports credentials, and conflicts are skipped by default (`--overwrite` to change).
|
||||
|
||||
## Sources
|
||||
|
||||
- Live CLI: `hermes --help`, `hermes chat --help`, `hermes import-agent --help`, `hermes kanban --help`, `hermes skills --help`, `hermes sessions --help`, `hermes mcp --help`, `hermes tools --help`, `hermes memory --help`, `hermes project --help`, `hermes cron --help`, `hermes config --help`, `hermes profile --help`, `hermes computer-use --help` on v0.21.1, reference install (Syslog kagentz), 2026-09-11. Raw dump: `cli-help-dump.txt` next to this file.
|
||||
- Docs: https://hermes-agent.nousresearch.com/docs/ (index) — all URLs in sources.md
|
||||
- Claude Code side: installed skill `delegate-coding-agent/references/claude-code.md` (Hermes Agent + Teknium, v2.2.1), `/home/hermes/.hermes/skills/autonomous-ai-agents/`
|
||||
@@ -0,0 +1,101 @@
|
||||
# 02 — High-Leverage Hermes Surfaces (the "harness power" inventory)
|
||||
|
||||
**Prepared:** 2026-09-11. Each surface: what it does, when to use it, exact command/file, doc URL. Verification: V-LIVE = confirmed against live CLI v0.21.1 on the reference install (Syslog kagentz); V-DOC = confirmed against official docs page (URL resolved HTTP 200); V-FILE = present on this machine's installed skills.
|
||||
|
||||
---
|
||||
|
||||
### 1. Persona / SOUL file
|
||||
- **What:** `SOUL.md` is Hermes' personality + standing-identity file, auto-injected into the system prompt alongside `AGENTS.md` rules and memory (confirmed by `--ignore-rules` help text which lists exactly what gets injected).
|
||||
- **When:** client wants the agent to have a consistent voice/role (e.g., "you are my analyst").
|
||||
- **Where:** `~/.hermes/SOUL.md` (per-profile: `~/.hermes/profiles/<name>/SOUL.md`).
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/configuration [V-DOC]
|
||||
|
||||
### 2. Persistent memory (built-in + providers)
|
||||
- **What:** Built-in `MEMORY.md` / `USER.md` always active; optional external providers (honcho, mem0, hindsight, byterover, …). Memory is injected every session — this is the single biggest cure for "it forgets my project."
|
||||
- **When:** after any correction or preference the user states ("use bun, not npm") — tell Hermes to remember it and it persists.
|
||||
- **Command:** `hermes memory setup|status|off|reset` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/memory [V-DOC]
|
||||
|
||||
### 3. Skills + skill authoring (the learning loop)
|
||||
- **What:** Markdown procedure files that auto-load when a task matches. The differentiator: Hermes can WRITE its own skills after learning a workflow (self-improving), and the curator maintains them (usage tracking, archiving, backups).
|
||||
- **When:** any workflow done twice — say "save this as a skill."
|
||||
- **Commands:** `hermes skills list|search|install|browse|config|check|update` [V-LIVE]; in-session `/skill <name>`, `/reload-skills` [V-DOC]; authoring tool in-session is `skill_manage` (agent-side; writer should describe it as "ask your Hermes to save the procedure as a skill").
|
||||
- **Docs:** https://hermes-agent.nousresearch.com/docs/reference/skills-catalog [V-DOC]; curator: https://hermes-agent.nousresearch.com/docs/user-guide/features/curator [V-DOC]
|
||||
|
||||
### 4. Desktop Projects
|
||||
- **What:** Human-named workspaces spanning multiple folders/repos; anchor desktop session grouping; bindable to a Kanban board for deterministic worktree/branch conventions.
|
||||
- **When:** Scot's multi-repo workflows (portfolio ops). `hermes project create <name>` then `add-folder`.
|
||||
- **Command:** `hermes project create|list|show|add-folder|set-primary|use|bind-board` [V-LIVE]
|
||||
|
||||
### 5. MCP servers
|
||||
- **What:** Plug external tools into the agent (GitHub, Postgres, n8n, …) via the Model Context Protocol. Also runs in reverse: `hermes mcp serve` exposes Hermes conversations to other agents.
|
||||
- **Command:** `hermes mcp add|list|test|configure|picker|catalog|install|serve` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp [V-DOC]
|
||||
|
||||
### 6. Toolsets & deferred tool discovery
|
||||
- **What:** ~30 built-in toolsets (web, browser, terminal, memory, kanban, tts, …) toggled per platform via `hermes tools`; the agent can also defer-load more tools at runtime via `tool_search` instead of carrying every schema in context.
|
||||
- **When:** trim toolsets for focus/cost, or enable `browser` for web work.
|
||||
- **Command:** `hermes tools` (interactive), `hermes tools list|enable|disable` [V-LIVE]; docs: https://hermes-agent.nousresearch.com/docs/reference/tools-reference [V-DOC]
|
||||
|
||||
### 7. Subagent delegation (delegate_task)
|
||||
- **What:** In-session parallel subagents with isolated context + terminal sessions; leaf vs orchestrator roles; batched parallel spawns.
|
||||
- **When:** research fan-out, parallel code review, anything that would flood the main context.
|
||||
- **Command:** agent-side tool (no CLI). In-session: ask Hermes to "delegate X to subagents." Docs: /docs/user-guide/features (delegation section) [V-DOC]
|
||||
|
||||
### 8. Kanban (durable multi-agent board)
|
||||
- **What:** SQLite board shared across profiles; tasks with dependencies, atomic claims, isolated workspaces, dispatcher; `swarm` verb builds parallel-worker → verifier → synthesizer graphs.
|
||||
- **When:** recurring multi-step operations, handoffs between specialist profiles, long-running campaigns that must survive restarts.
|
||||
- **Command:** `hermes kanban create|list|show|swarm|link|complete|watch|stats|dispatch` (40+ verbs) [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/kanban [V-DOC]
|
||||
|
||||
### 9. Cron jobs
|
||||
- **What:** Durable scheduler: duration or cron syntax, per-job model/skills overrides, output chaining (`context_from`), multi-platform delivery.
|
||||
- **When:** daily reports, monitoring with alerts, weekly reviews.
|
||||
- **Command:** `hermes cron list|create|edit|pause|resume|run|remove|doctor|status` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/cron [V-DOC]
|
||||
|
||||
### 10. Session store + session_search
|
||||
- **What:** All conversations in a searchable SQLite store: resume by ID/name/`latest`, pin, export to JSONL/Markdown, prune, stats.
|
||||
- **When:** "what did we decide last week" — the agent can search past sessions; user can browse them.
|
||||
- **Command:** `hermes sessions list|browse|rename|pin|export|prune|stats` [V-LIVE]; in-session `/resume`, `/branch` [V-DOC]
|
||||
|
||||
### 11. Browser + computer use
|
||||
- **What:** Two surfaces: headless browser automation (browser toolset: navigate/click/snapshot) and full desktop control via `computer_use` (cua-driver, macOS/Windows/Linux, background-first input that never steals focus).
|
||||
- **When:** web research → headless browser; native apps (Excel, Figma, native chat) → computer use.
|
||||
- **Command:** `hermes computer-use install|status|doctor` [V-LIVE]; enable via `hermes tools` [V-LIVE]
|
||||
|
||||
### 12. Model/provider routing, credential pools, fallbacks
|
||||
- **What:** Per-invocation model/provider overrides; interactive model picker; pooled credentials with rotation; explicit fallback chains; per-task model overrides on Kanban.
|
||||
- **Command:** `hermes model` [V-LIVE], `hermes fallback list|add|remove` [V-LIVE], `hermes auth add|list|priority|reset` [V-LIVE]; per-run flags `-m`, `--provider`, `--reasoning` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/integrations/providers [V-DOC]
|
||||
- **Note for Scot (OpenRouter + small fast models):** `--reasoning high` on hard tasks; `hermes fallback add` so a failed call rolls to a second model instead of erroring.
|
||||
|
||||
### 13. Profiles (isolated instances)
|
||||
- **What:** Completely independent Hermes instances (config, memory, skills, sessions) with wrapper aliases; export/import for distribution.
|
||||
- **When:** separate work/persona contexts, or one profile per client.
|
||||
- **Command:** `hermes profile list|create|use|alias|export|import` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/profiles [V-DOC]
|
||||
|
||||
### 14. Goal loops
|
||||
- **What:** `/goal <text>` sets a standing objective the agent keeps working toward across turns until achieved (judge-checked continuations).
|
||||
- **When:** "keep the CI green until it passes," "keep researching until you have 5 verified sources."
|
||||
- **Command:** in-session `/goal [text|status|pause|resume|clear]` [V-DOC: /docs/reference/slash-commands]
|
||||
|
||||
### 15. Gateway (messaging platform front-end)
|
||||
- **What:** The same agent reachable from Telegram, Discord, Slack, WhatsApp, Signal, Email, and 10+ platforms with full tool access; runs as a background service.
|
||||
- **Command:** `hermes gateway run|install|start|status|setup` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ [V-DOC]
|
||||
|
||||
### 16. Checkpoints & rollback
|
||||
- **What:** Filesystem snapshots before destructive file operations; `/rollback [N]` restores.
|
||||
- **When:** letting the agent loose on important files.
|
||||
- **Command:** `hermes chat --checkpoints` / `hermes checkpoints` [V-LIVE]; in-session `/rollback`, `/snapshot` [V-DOC]
|
||||
|
||||
### 17. Projects↔Kanban binding + worktree mode
|
||||
- **What:** `hermes project bind-board` ties a board to a project (deterministic worktree + branch per task); `-w/--worktree` runs any session in an isolated git worktree.
|
||||
- **When:** parallel coding agents that must not collide.
|
||||
- **Command:** `hermes project bind-board` [V-LIVE]; `hermes -w` [V-LIVE]
|
||||
|
||||
### 18. Prompt-size introspection
|
||||
- **What:** Byte breakdown of system prompt + tool schemas — diagnose why responses feel "dumb" (usually context bloat).
|
||||
- **Command:** `hermes prompt-size` [V-LIVE]
|
||||
@@ -0,0 +1,121 @@
|
||||
# 03 — Command Cheatsheet (every entry verified)
|
||||
|
||||
**Verification method:** each VERIFIED-LIVE entry was confirmed against `hermes --help` or `hermes <cmd> --help` on Hermes Agent v0.21.1 (2026.9.7), reference install (Syslog kagentz), 2026-09-11. Raw output: `cli-help-dump.txt`. DOC-ONLY entries come from the official docs (URL given). Nothing is invented.
|
||||
|
||||
## (a) CLI — `hermes ...`
|
||||
|
||||
### Setup & health
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes setup` | Interactive setup wizard | VERIFIED-LIVE |
|
||||
| `hermes doctor [--fix] [--live]` | Diagnose config/deps; `--fix` auto-repairs | VERIFIED-LIVE |
|
||||
| `hermes status [--all] [--deep]` | Component status | VERIFIED-LIVE |
|
||||
| `hermes config show/edit/get/set/unset/path/env-path/check/migrate` | View/edit config | VERIFIED-LIVE |
|
||||
| `hermes update` | Update Hermes to latest | VERIFIED-LIVE |
|
||||
|
||||
### The Claude Code bridge (highest value for Scot)
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes import-agent claude-code [--dry-run] [--overwrite] [--yes]` | One-command import of a Claude Code setup: maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, memories into Hermes equivalents. Never imports API keys. | VERIFIED-LIVE |
|
||||
| `hermes import-agent codex` | Same for Codex CLI setups | VERIFIED-LIVE |
|
||||
| `hermes sessions import` | Import a Claude Code or Codex CLI **session/conversation** into Hermes | VERIFIED-LIVE (subcommand listed in `hermes sessions --help`) |
|
||||
| `hermes skills trust` | Trust a repo so its project-local skills (`./.hermes/skills`) load — the Hermes analog of `.claude/skills/` | VERIFIED-LIVE |
|
||||
|
||||
### Daily driving
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes` / `hermes chat` | Interactive session | VERIFIED-LIVE |
|
||||
| `hermes -c [NAME]` / `hermes --resume <id\|latest>` | Resume by name or ID | VERIFIED-LIVE |
|
||||
| `hermes --in DIR --resume latest` | Resume the latest session for a directory | VERIFIED-LIVE |
|
||||
| `hermes -z "PROMPT"` | One-shot: prints ONLY the final answer (scripting/CI); tools, memory, and AGENTS.md still load | VERIFIED-LIVE |
|
||||
| `hermes chat -q "PROMPT"` | Single-query mode | VERIFIED-LIVE |
|
||||
| `hermes -m MODEL --provider PROVIDER --reasoning LEVEL` | Per-run model/provider/reasoning overrides (`none…ultra`) | VERIFIED-LIVE |
|
||||
| `hermes -s SKILL1,SKILL2` | Preload specific skills for the session | VERIFIED-LIVE |
|
||||
| `hermes -t TOOLSETS` | Restrict toolsets for this run | VERIFIED-LIVE |
|
||||
| `hermes -w` | Isolated git worktree session (parallel agents on one repo) | VERIFIED-LIVE |
|
||||
| `hermes chat --checkpoints` | Enable filesystem checkpoints (`/rollback` to restore) | VERIFIED-LIVE |
|
||||
| `hermes chat --max-turns N` / `--run-budget SECONDS` | Cap loop iterations / wall-clock budget | VERIFIED-LIVE |
|
||||
|
||||
### Context & memory management
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes memory setup/status/off/reset` | External memory provider management (built-in MEMORY.md/USER.md always active) | VERIFIED-LIVE |
|
||||
| `hermes sessions list/browse/rename/pin/export/prune/stats` | Session store management | VERIFIED-LIVE |
|
||||
| `hermes skills list/search/install/inspect/browse/config/check/update` | Skill management | VERIFIED-LIVE |
|
||||
| `hermes skills trust/untrust` | Repo-local skill trust | VERIFIED-LIVE |
|
||||
| `hermes curator status/run/pause/pin/...` | Background skill maintenance (auto-archive, backups) | VERIFIED-LIVE |
|
||||
| `hermes prompt-size` | Byte breakdown of system prompt + tool schemas (context-bloat diagnosis) | VERIFIED-LIVE |
|
||||
| `hermes insights [--days N]` | Usage analytics | VERIFIED-LIVE |
|
||||
|
||||
### Tools, MCP, integrations
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes tools` (interactive) / `list/enable/disable` | Per-platform toolset toggles; MCP tools as `server:tool` | VERIFIED-LIVE |
|
||||
| `hermes mcp add/remove/list/test/configure/picker/catalog/install` | MCP server management (incl. one-click catalog installs) | VERIFIED-LIVE |
|
||||
| `hermes mcp serve` | Run Hermes AS an MCP server for other agents | VERIFIED-LIVE |
|
||||
| `hermes computer-use install/status/doctor` | Desktop-control backend (cua-driver) | VERIFIED-LIVE |
|
||||
| `hermes gateway run/install/start/status/setup` | Messaging gateway (Telegram, Discord, Slack, WhatsApp, …) | VERIFIED-LIVE |
|
||||
| `hermes send` | Send a message to a configured platform (scripts/cron/CI) | VERIFIED-LIVE |
|
||||
|
||||
### Automation & multi-agent
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes cron list/create/edit/pause/resume/run/remove/doctor` | Scheduled jobs (durable, multi-platform delivery) | VERIFIED-LIVE |
|
||||
| `hermes cron notepad` | Durable per-job key-value notepad across runs | VERIFIED-LIVE |
|
||||
| `hermes kanban create/list/show/link/complete/swarm/...` | Durable multi-profile task board (40+ verbs) | VERIFIED-LIVE |
|
||||
| `hermes kanban swarm` | Generate a parallel-workers → verifier → synthesizer task graph | VERIFIED-LIVE |
|
||||
| `hermes project create/list/add-folder/bind-board` | Named multi-folder workspaces (desktop Projects) | VERIFIED-LIVE |
|
||||
| `hermes profile list/create/use/alias/export/import` | Isolated Hermes instances | VERIFIED-LIVE |
|
||||
| `hermes auth add/list/priority/reset` | Pooled credentials per provider (rotation) | VERIFIED-LIVE |
|
||||
| `hermes fallback list/add/remove` | Fallback model chain (auto-rollover on failure) | VERIFIED-LIVE |
|
||||
| `hermes model` | Interactive model/provider picker | VERIFIED-LIVE |
|
||||
| `hermes -yolo` | Bypass command approval prompts (use with care) | VERIFIED-LIVE |
|
||||
| `hermes pause` / `hermes resume` | Emergency stop / lift (pauses cron, kanban dispatch, gateway turns) | VERIFIED-LIVE |
|
||||
|
||||
## (b) In-session slash commands
|
||||
|
||||
Source: official slash-commands reference https://hermes-agent.nousresearch.com/docs/reference/slash-commands (DOC-ONLY — slash commands run inside a chat session and were not exercised from this headless research run; the CLI subcommands they map to were verified live). DOC-ONLY.
|
||||
|
||||
### Context & session
|
||||
| Command | What it does |
|
||||
|---|---|
|
||||
| `/help` | List all commands (authoritative in your version) |
|
||||
| `/new` (`/reset`) | Fresh session |
|
||||
| `/resume [name]` | Resume a named/recent session |
|
||||
| `/branch` (`/fork`) | Branch the current session |
|
||||
| `/compress` | Manually compress context (auto-compression also exists) |
|
||||
| `/undo` | Remove last exchange |
|
||||
| `/retry` | Resend last message |
|
||||
| `/title [name]` | Name the session |
|
||||
| `/save` | Save conversation to file |
|
||||
| `/history` | Show conversation history |
|
||||
|
||||
### Power surfaces
|
||||
| Command | What it does |
|
||||
|---|---|
|
||||
| `/skill <name>` | Load a skill into the session |
|
||||
| `/skills` | Search/install skills |
|
||||
| `/reload-skills` | Re-scan skill directory |
|
||||
| `/tools` / `/toolsets` | Manage tools |
|
||||
| `/goal [text]` | Set a standing goal the agent works toward across turns (`/goal status/pause/clear` to manage) |
|
||||
| `/background <prompt>` | Run a prompt in the background |
|
||||
| `/queue <prompt>` | Queue a prompt for the next turn |
|
||||
| `/steer <prompt>` | Inject a course-correction after the next tool call without interrupting |
|
||||
| `/agents` | Show active agents and running tasks |
|
||||
| `/cron` | Manage cron jobs in-session |
|
||||
| `/kanban` | Multi-profile collaboration board in-session |
|
||||
| `/model [name]` | Show/change model mid-session |
|
||||
| `/reasoning [level]` | Set reasoning effort |
|
||||
| `/voice [on\|off\|tts]` | Voice mode |
|
||||
| `/rollback [N]` | Restore filesystem checkpoint (needs `--checkpoints`) |
|
||||
| `/usage` | Token usage |
|
||||
| `/insights [days]` | Usage analytics |
|
||||
| `/platforms` | Gateway platform status |
|
||||
| `/compact`-equivalent note | Hermes compresses automatically near the context limit; no manual threshold watch needed like Claude Code's `/context` |
|
||||
|
||||
### The "work alongside Claude Code" shortlist
|
||||
1. `hermes import-agent claude-code --dry-run` → migrate the setup (VERIFIED-LIVE)
|
||||
2. `hermes sessions import` → bring the conversation history over (VERIFIED-LIVE)
|
||||
3. `hermes skills trust` → load repo-local skills like `.claude/skills/` (VERIFIED-LIVE)
|
||||
4. `hermes -c` / `hermes --in <repo> --resume latest` → per-directory session continuity (VERIFIED-LIVE)
|
||||
5. `hermes mcp serve` → expose Hermes to Claude Code as an MCP server (VERIFIED-LIVE) — the reverse direction Claude Code can't do
|
||||
@@ -0,0 +1,76 @@
|
||||
# 04 — The Claude Code Bridge: Running Hermes WITH Claude Code
|
||||
|
||||
**Prepared:** 2026-09-11. Scot already runs both tools. This file documents the proven integration patterns, citing the installed skills on this host (paths under `/home/hermes/.hermes/skills/`) and official docs.
|
||||
|
||||
## Pattern 0 — Import (do this first)
|
||||
`hermes import-agent claude-code` [VERIFIED-LIVE] maps CLAUDE.md/AGENTS.md instructions,
|
||||
permission allowlists, MCP servers, skills, and memories into Hermes equivalents. It always
|
||||
shows a preview, never imports credentials. `hermes sessions import` [VERIFIED-LIVE] pulls
|
||||
in old Claude Code conversations. After import, Hermes "knows" the projects — the context
|
||||
gap disappears on day one.
|
||||
|
||||
## Pattern 1 — Hermes as orchestrator, Claude Code as worker
|
||||
Source: installed skill **`autonomous-ai-agents/delegate-coding-agent`** (v1.0.0) + its
|
||||
reference `references/claude-code.md` (v2.2.1) [V-FILE]. The skill is an official Hermes
|
||||
skill authored for exactly this.
|
||||
|
||||
Two orchestration modes (verbatim from the skill):
|
||||
- **Print mode (preferred):** `claude -p '<task>' --allowedTools 'Read,Edit' --max-turns 10` —
|
||||
one-shot, no dialogs, structured JSON output with `session_id`, `num_turns`,
|
||||
`total_cost_usd`. Ask Hermes: *"delegate this coding task to Claude Code in print mode."*
|
||||
- **Interactive PTY via tmux:** multi-turn sessions — Hermes starts `tmux new-session`,
|
||||
sends prompts with `send-keys`, monitors with `capture-pane`. For iterative
|
||||
refactor → review → fix cycles.
|
||||
|
||||
Cross-agent review loop (also from the skill):
|
||||
```
|
||||
git diff main...feature | claude -p 'Review this diff for bugs and security issues.' --max-turns 1
|
||||
```
|
||||
Hermes runs this, reads the output, and fixes findings itself — Claude Code becomes a
|
||||
reviewer Hermes coordinates.
|
||||
|
||||
Safety rails the skill prescribes: explicit `workdir`, clean git status before launch,
|
||||
narrow task prompts, `git diff` review, targeted tests before committing.
|
||||
|
||||
## Pattern 2 — Parallel workstreams + neutral merge reconciliation
|
||||
Source: installed skill **`autonomous-ai-agents/merge-reconciler`** [V-FILE].
|
||||
When Hermes and Claude Code (or two Hermes workers) both edit the same repo and collide:
|
||||
- Do NOT let either agent resolve the conflict — both are biased toward their own side.
|
||||
- Spawn a **neutral third agent** with the merge-reconciler skill; it classifies every
|
||||
conflicted hunk (disjoint-intent / same-question-different-answer / superseded), resolves
|
||||
under an impartiality contract (touch only conflict markers, surface every design call),
|
||||
verifies with build/tests, and hands back a summary naming every hunk decision.
|
||||
- Kanban-native shape: a reconciliation card assigned to a **third profile** with both
|
||||
workers' cards as parents — parent links carry both sides' completion summaries into the
|
||||
reconciler's context automatically.
|
||||
|
||||
## Pattern 3 — Hermes as MCP server (Claude Code gets Hermes tools)
|
||||
`hermes mcp serve` [VERIFIED-LIVE] runs Hermes as an MCP server exposing its conversations
|
||||
and capabilities. Claude Code supports MCP clients (`claude mcp add`), so Claude Code can
|
||||
consume Hermes as a tool provider — persistent memory, skills, cron — the surfaces Claude
|
||||
Code lacks. This is the reverse-bridge only Hermes can offer. Docs:
|
||||
https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp
|
||||
|
||||
## Pattern 4 — Import legacy sessions for continuity
|
||||
`hermes sessions import` [VERIFIED-LIVE] imports a Claude Code session into the Hermes
|
||||
store; from then on `hermes --resume <id>` / `hermes sessions browse` treat it as native
|
||||
history. Use when mid-project: the new agent picks up exactly where Claude Code left off.
|
||||
|
||||
## Pattern 5 — Desktop GUI automation either side can use
|
||||
Source: installed skill **`autonomous-ai-agents/computer-use`** (v2.0.0) [V-FILE].
|
||||
`hermes computer-use install` sets up cua-driver; the `computer_use` toolset drives native
|
||||
desktop apps background-first (never steals focus/cursor), any-model, cross-platform.
|
||||
Relevant to the bridge because Claude Code has no desktop automation — if a task needs
|
||||
Figma/Excel/native apps, that part routes to Hermes while the code routes to Claude Code.
|
||||
Cmd: `hermes computer-use doctor` for health checks.
|
||||
|
||||
## Pattern 6 — The import-agent philosophy in one line
|
||||
Claude Code holds repo context in `CLAUDE.md`; Hermes holds it in `AGENTS.md` + memory +
|
||||
skills. `hermes import-agent claude-code` translates the first; the learning loop
|
||||
("save this as a skill") rebuilds the rest automatically the more Scot uses Hermes.
|
||||
|
||||
## Reference paths (for the writer)
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/delegate-coding-agent/SKILL.md` and `references/claude-code.md`
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/merge-reconciler/SKILL.md`
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/computer-use/SKILL.md`
|
||||
- `hermes import-agent --help` raw output in `cli-help-dump.txt` (lines 554-577)
|
||||
@@ -0,0 +1,119 @@
|
||||
# 05 — The Video Watch List (every URL verified 2026-09-11)
|
||||
|
||||
**Verification method:** each video was found via YouTube search-results scrape (`videoRenderer` metadata), then confirmed with the YouTube oEmbed endpoint (`curl -s "https://www.youtube.com/oembed?url=<URL>&format=json"`) — every entry below returned HTTP 200 with matching title/author (status PASS). Publish dates, durations, and view counts were read from each watch page's metadata. Raw evidence for all 21 entries: `video-verification.json` in this directory. **21/21 PASS, 0 FAIL.**
|
||||
|
||||
## Tier 1 — Hermes-specific, start here
|
||||
|
||||
### 1. Learn 95% of Hermes Agent in 31 Minutes
|
||||
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=Ta2wg6xPaY4 | **Duration:** 31:28 | **Published:** 2026-08-09 | **Views:** ~129k
|
||||
- **oEmbed:** PASS (title/author match)
|
||||
- **What it demonstrates:** end-to-end Hermes fundamentals — install, sessions, skills, memory, the learning loop. The most complete single-video orientation found.
|
||||
- **Watch this when you want** the fastest real overview of the whole harness before touching config.
|
||||
|
||||
### 2. Hermes Agent Fundamentals In 29 Minutes
|
||||
- **Channel:** Tina Huang | **URL:** https://www.youtube.com/watch?v=5_N84t1rUU0 | **Duration:** 29:40 | **Published:** 2026-07-20 | **Views:** ~463k (highest-reach Hermes video found)
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** conceptual grounding — why Hermes' memory/skills loop differs from one-shot coding agents; practical walkthrough.
|
||||
- **Watch this when you want** to understand *why* Hermes feels different from Claude Code, not just which buttons to press.
|
||||
|
||||
### 3. Every Level of Hermes Agent Explained
|
||||
- **Channel:** Jack Roberts | **URL:** https://www.youtube.com/watch?v=6GtF_uHbGhw | **Duration:** 25:35 | **Published:** 2026-06-17 | **Views:** ~163k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** beginner → advanced ladder of features (memory, skills, automation, multi-agent).
|
||||
- **Watch this when you want** a map of what to learn next after the basics.
|
||||
|
||||
### 4. Hermes Agent Full Tutorial INSTALLATION + USECASES
|
||||
- **Channel:** CodeHead | **URL:** https://www.youtube.com/watch?v=8GjyOQy19so | **Duration:** 7:47 | **Published:** 2026-05-14 | **Views:** ~64k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** install through real use-cases, compact.
|
||||
- **Watch this when you want** a quick install-to-value demo to share with a colleague.
|
||||
|
||||
### 5. Hermes Agent Explained In 5 Minutes
|
||||
- **Channel:** CodeHead | **URL:** https://www.youtube.com/watch?v=9GpWELm3_XI | **Duration:** 4:53 | **Published:** 2026-05-23 | **Views:** ~251k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** 5-minute conceptual pitch of the agent and its learning loop.
|
||||
- **Watch this when you want** the elevator pitch before committing 30 minutes.
|
||||
|
||||
## Tier 2 — Hermes-specific deep dives
|
||||
|
||||
### 6. 100 Days With Hermes Agent in 21 Minutes
|
||||
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=sCa3BtpkziQ | **Duration:** 21:19 | **Published:** 2026-06-17 | **Views:** ~58k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** long-horizon usage — what memory/skills accumulation actually looks like after months of daily use.
|
||||
- **Watch this when you want** to see the payoff of the learning loop over time.
|
||||
|
||||
### 7. Hermes Agent - Crash Course for Beginners (AI Agent)
|
||||
- **Channel:** Adrian Twarog | **URL:** https://www.youtube.com/watch?v=4sAmpcSOVEw | **Duration:** 22:19 | **Published:** 2026-07-21 | **Views:** ~46k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** beginner crash course from a well-known dev-YouTube creator.
|
||||
- **Watch this when you want** a second independent explanation of the basics.
|
||||
|
||||
### 8. Hermes Agent: The Ultimate Beginner's Guide
|
||||
- **Channel:** Metics Media | **URL:** https://www.youtube.com/watch?v=CwPUOVUdApE | **Duration:** 37:08 | **Published:** 2026-04-24 | **Views:** ~119k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** long-form beginner guide incl. setup and everyday workflows.
|
||||
- **Watch this when you want** the most thorough single walkthrough in one sitting.
|
||||
|
||||
### 9. Hermes Agent Just Killed OpenClaw (Full Tutorial)
|
||||
- **Channel:** Leon van Zyl | **URL:** https://www.youtube.com/watch?v=jmtpYUOr7_U | **Duration:** 19:59 | **Published:** 2026-04-28 | **Views:** ~16k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** full tutorial framing Hermes against the OpenClaw workflow (MCP config, memory, agents).
|
||||
- **Watch this when you want** a practitioner's feature-by-feature tutorial.
|
||||
|
||||
### 10. Hermes Agent vs OpenClaw
|
||||
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=zwqhemjHq3E | **Duration:** 15:28 | **Published:** 2026-04-20 | **Views:** ~35k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** head-to-head comparison of the two agent harnesses.
|
||||
- **Watch this when you want** the tradeoffs between Hermes and its main alternative.
|
||||
|
||||
### 11. Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model
|
||||
- **Channel:** Tonbi's AI Garage | **URL:** https://www.youtube.com/watch?v=8tpuky8HpXw | **Duration:** 15:08 | **Published:** 2026-03-11 | **Views:** ~18k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** Hermes driven by an OpenRouter-served open model (Qwen 3) — directly relevant to an OpenRouter-connected install.
|
||||
- **Watch this when you want** to see how small open models behave inside Hermes.
|
||||
|
||||
### 12. Use This To Make The Hermes Agent Basically Free
|
||||
- **Channel:** AI LABS | **URL:** https://www.youtube.com/watch?v=5d02TYoOzfE | **Duration:** 13:08 | **Published:** 2026-07-01 | **Views:** ~55k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** running Hermes on cheap/free model backends.
|
||||
- **Watch this when you want** to cut inference costs on an OpenRouter account.
|
||||
|
||||
### 13. Hermes Agent The 24/7 Self-Evolving AI Agent!
|
||||
- **Channel:** WorldofAI | **URL:** https://www.youtube.com/watch?v=cu2fgknmemA | **Duration:** 9:15 | **Published:** 2026-04-07 | **Views:** ~47k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** always-on operation: gateway, cron, background automation.
|
||||
- **Watch this when you want** to turn Hermes from a chat window into a 24/7 assistant.
|
||||
|
||||
## Tier 3 — Adjacent (origin/philosophy; not tutorials)
|
||||
|
||||
### 14. Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra
|
||||
- **Channel:** Peter Yang | **URL:** https://www.youtube.com/watch?v=UWjh5Z4s8jY | **Duration:** 46:45 | **Published:** 2026-08-02 | **Views:** ~37k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** interview with Hermes' co-founder on the design philosophy (self-improving agents, skills as memory).
|
||||
- **Watch this when you want** to understand where the product is going.
|
||||
|
||||
### 15. Hermes Agent: Agents that grow with you | Episode #357
|
||||
- **Channel:** Practical AI | **URL:** https://www.youtube.com/watch?v=UTZhvPXnmwA | **Duration:** 47:34 | **Published:** 2026-05-20 | **Views:** ~1.9k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** podcast-depth technical discussion of the agent architecture.
|
||||
- **Watch this when you want** the engineering story behind the learning loop.
|
||||
|
||||
### 16. Did Hermes Agent just kill OpenClaw? (full guide)
|
||||
- **Channel:** Alex Finn | **URL:** https://www.youtube.com/watch?v=tP6yf22OJdI | **Duration:** 13:55 | **Published:** 2026-03-31 | **Views:** ~132k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** guide-style comparison/switch content.
|
||||
- **Watch this when you want** a switcher's guide perspective.
|
||||
|
||||
### 17. Hermes Agent: Why Everyone's Ditching OpenClaw in 2026
|
||||
- **Channel:** Luke Alexander AI | **URL:** https://www.youtube.com/watch?v=1UgXUjT-QtI | **Duration:** 18:03 | **Published:** 2026-03-26 | **Views:** ~14k
|
||||
- **oEmbed:** PASS — adjacent, comparison content.
|
||||
- **Watch this when you want** more comparison context.
|
||||
|
||||
## Honesty note (required by task spec)
|
||||
At least 16 of the 17 entries above are directly Hermes-specific (not merely adjacent); the
|
||||
"fewer than 5 exist" fallback clause was NOT needed — no padding was necessary. Entries
|
||||
found in search but excluded as thin/low-signal: `iqN6MVzpJTk` (3.1k views, news-style),
|
||||
`83nWNRKZTCE` (465 views), `6M2tItdARew` (1.2k views), `P2LIFtrRr2U` (promo-style) — all
|
||||
also verified PASS and kept in `video-verification.json` as spares. No official Nous
|
||||
Research YouTube tutorial channel was found in searches; the strongest signal of
|
||||
Hermes-specific video content is the third-party ecosystem above.
|
||||
@@ -0,0 +1,577 @@
|
||||
===== hermes chat --help =====
|
||||
usage: hermes chat [-h] [-q QUERY | --query-file PATH] [--oneshot]
|
||||
[--image IMAGE] [-m MODEL] [-t TOOLSETS]
|
||||
[--reasoning LEVEL] [-s SKILLS] [--provider PROVIDER] [-v]
|
||||
[-Q] [--resume SESSION_ID] [--no-restore-cwd] [--in DIR]
|
||||
[--continue [SESSION_NAME]] [--create-if-missing]
|
||||
[--worktree] [--accept-hooks] [--checkpoints]
|
||||
[--max-turns N] [--run-budget SECONDS] [--yolo]
|
||||
[--pass-session-id] [--ignore-user-config] [--ignore-rules]
|
||||
[--safe-mode] [--source SOURCE] [--tui] [--cli] [--dev]
|
||||
|
||||
Start an interactive chat session with Hermes Agent
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
-q, --query QUERY Query to run. On a real TTY the prompt seeds an
|
||||
interactive session (submitted literally as the first
|
||||
turn); combined with --oneshot or -Q, or on a non-TTY,
|
||||
it answers and exits.
|
||||
--query-file PATH Read the single query from a file instead of the
|
||||
command line ('-' reads stdin). Safe for arbitrary
|
||||
text: nothing is shell-interpreted, so quotes, $(...),
|
||||
and backticks are preserved verbatim. Mutually
|
||||
exclusive with -q.
|
||||
--oneshot With -q/--query-file: answer the query and exit
|
||||
(legacy single-query behavior) instead of seeding an
|
||||
interactive session. Implied on non-TTY stdio and by
|
||||
-Q/--quiet.
|
||||
--image IMAGE Optional local image path to attach to a single query
|
||||
-m, --model MODEL Model to use (e.g., anthropic/claude-sonnet-4)
|
||||
-t, --toolsets TOOLSETS
|
||||
Comma-separated toolsets to enable
|
||||
--reasoning LEVEL Reasoning effort for this session: none, minimal, low,
|
||||
medium, high, xhigh, max, or ultra. Overrides
|
||||
agent.reasoning_effort for this run only (same levels
|
||||
as the /reasoning slash command).
|
||||
-s, --skills SKILLS Preload one or more skills for the session (repeat
|
||||
flag or comma-separate)
|
||||
--provider PROVIDER Inference provider (default: auto). Built-in or a
|
||||
user-defined name from `providers:` in config.yaml.
|
||||
-v, --verbose Verbose output
|
||||
-Q, --quiet Quiet mode for programmatic use: suppress banner,
|
||||
spinner, and tool previews. Only output the final
|
||||
response and session info.
|
||||
--resume, -r SESSION_ID
|
||||
Resume a previous session by ID (shown on exit), or
|
||||
'latest' for the most recent session
|
||||
--no-restore-cwd Don't cd into a resumed session's recorded working
|
||||
directory.
|
||||
--in DIR Change into DIR before starting or resuming (scopes '
|
||||
--resume latest' / -c lookups to DIR's workspace).
|
||||
--continue, -c [SESSION_NAME]
|
||||
Resume a session by name, or the most recent if no
|
||||
name given
|
||||
--create-if-missing With -c/--continue <name>: if no session matches the
|
||||
name, create a new session with that title and proceed
|
||||
(instead of failing with a not-found error).
|
||||
Programmatic callers that want 'send to this named
|
||||
thread, making it if needed'.
|
||||
--worktree, -w Run in an isolated git worktree (for parallel agents
|
||||
on the same repo)
|
||||
--accept-hooks Auto-approve any unseen shell hooks declared in
|
||||
config.yaml without a TTY prompt (see also
|
||||
HERMES_ACCEPT_HOOKS env var and hooks_auto_accept: in
|
||||
config.yaml).
|
||||
--checkpoints Enable filesystem checkpoints before destructive file
|
||||
operations (use /rollback to restore)
|
||||
--max-turns N Maximum tool-calling iterations per conversation turn
|
||||
(default: 500, or agent.max_turns in config)
|
||||
--run-budget SECONDS Optional wall-clock budget in seconds for each
|
||||
conversation run. At 80% elapsed the agent gets a one-
|
||||
time wrap-up notice, and implicit provider stale
|
||||
timeouts are capped to the remaining budget so one
|
||||
hung call can't consume the run. Unset = off. Also
|
||||
configurable as agent.run_budget_seconds in
|
||||
config.yaml. Intended for one-shot/eval invocations
|
||||
with a hard ceiling.
|
||||
--yolo Bypass all dangerous command approval prompts (use at
|
||||
your own risk)
|
||||
--pass-session-id Include the session ID in the agent's system prompt
|
||||
--ignore-user-config Ignore ~/.hermes/config.yaml and fall back to built-in
|
||||
defaults (credentials in .env are still loaded).
|
||||
Useful for isolated CI runs, reproduction, and third-
|
||||
party integrations.
|
||||
--ignore-rules Skip auto-injection of AGENTS.md, SOUL.md,
|
||||
.cursorrules, memory, and preloaded skills. Combine
|
||||
with --ignore-user-config for a fully isolated run.
|
||||
--safe-mode Troubleshooting mode: disable ALL customizations —
|
||||
user config, AGENTS.md/memory injection, plugins, and
|
||||
MCP servers (implies --ignore-user-config and
|
||||
--ignore-rules). Use to isolate whether a problem
|
||||
comes from your setup or from Hermes itself.
|
||||
--source SOURCE Session source tag for filtering (default: cli). Use
|
||||
'tool' for third-party integrations that should not
|
||||
appear in user session lists.
|
||||
--tui Launch the modern TUI instead of the classic REPL
|
||||
--cli Force the classic prompt_toolkit REPL (overrides
|
||||
display.interface=tui)
|
||||
--dev With --tui: run TypeScript sources via tsx (skip dist
|
||||
build)
|
||||
===== hermes model --help =====
|
||||
usage: hermes model [-h] [--refresh] [--portal-url PORTAL_URL]
|
||||
[--inference-url INFERENCE_URL] [--client-id CLIENT_ID]
|
||||
[--scope SCOPE] [--no-browser] [--timeout TIMEOUT]
|
||||
[--ca-bundle CA_BUNDLE] [--insecure]
|
||||
|
||||
Interactively select your inference provider and default model
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--refresh Wipe the model picker disk cache and re-fetch every
|
||||
provider's live /v1/models list.
|
||||
--portal-url PORTAL_URL
|
||||
Portal base URL for Nous login (default: production
|
||||
portal)
|
||||
--inference-url INFERENCE_URL
|
||||
Inference API base URL for Nous login (default:
|
||||
production inference API)
|
||||
--client-id CLIENT_ID
|
||||
OAuth client id to use for Nous login (default:
|
||||
hermes-cli)
|
||||
--scope SCOPE OAuth scope to request for Nous login
|
||||
--no-browser Do not attempt to open the browser automatically
|
||||
during Nous login
|
||||
--timeout TIMEOUT HTTP request timeout in seconds for Nous login
|
||||
(default: 15)
|
||||
--ca-bundle CA_BUNDLE
|
||||
Path to CA bundle PEM file for Nous TLS verification
|
||||
--insecure Disable TLS verification for Nous login (testing only)
|
||||
===== hermes config --help =====
|
||||
usage: hermes config [-h]
|
||||
{show,edit,get,set,unset,path,env-path,check,migrate} ...
|
||||
|
||||
Manage Hermes Agent configuration
|
||||
|
||||
positional arguments:
|
||||
{show,edit,get,set,unset,path,env-path,check,migrate}
|
||||
show Show current configuration
|
||||
edit Open config file in editor
|
||||
get Print a resolved configuration value
|
||||
set Set a configuration value
|
||||
unset Remove a configuration value
|
||||
path Print config file path
|
||||
env-path Print .env file path
|
||||
check Check for missing/outdated config
|
||||
migrate Update config with new options
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes cron --help =====
|
||||
usage: hermes cron [-h] [--accept-hooks]
|
||||
{list,create,add,edit,pause,resume,run,remove,rm,delete,status,runs,history,incidents,notepad,doctor,tick} ...
|
||||
|
||||
Manage scheduled tasks
|
||||
|
||||
positional arguments:
|
||||
{list,create,add,edit,pause,resume,run,remove,rm,delete,status,runs,history,incidents,notepad,doctor,tick}
|
||||
list List scheduled jobs
|
||||
create (add) Create a scheduled job
|
||||
edit Edit an existing scheduled job
|
||||
pause Pause a scheduled job
|
||||
resume Resume a paused job
|
||||
run Run a job on the next scheduler tick
|
||||
remove (rm, delete)
|
||||
Remove a scheduled job
|
||||
status Check if cron scheduler is running
|
||||
runs (history) Show durable execution attempts
|
||||
incidents List or acknowledge durable cron failure incidents
|
||||
notepad Read/write a job's durable notepad (persistent KV
|
||||
across runs)
|
||||
doctor Check scheduled jobs for common health issues
|
||||
tick Run due jobs once and exit
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
|
||||
(equivalent to HERMES_ACCEPT_HOOKS=1 /
|
||||
hooks_auto_accept: true).
|
||||
===== hermes kanban --help =====
|
||||
usage: hermes kanban [-h] [--board <slug>]
|
||||
{init,boards,create,swarm,list,ls,show,assign,set-model,reclaim,reassign,diagnostics,diag,link,unlink,claim,comment,attach,attachments,attach-rm,complete,edit,block,schedule,unblock,request-review,request-changes,reopen-review,promote,archive,tail,dispatch,daemon,watch,stats,notify-subscribe,notify-list,notify-unsubscribe,log,runs,heartbeat,assignees,context,specify,decompose,gc,repair} ...
|
||||
|
||||
Durable SQLite-backed task board shared across Hermes profiles. Tasks are
|
||||
claimed atomically, can depend on other tasks, and are executed by a named
|
||||
profile in an isolated workspace. See https://hermes-
|
||||
agent.nousresearch.com/docs/user-guide/features/kanban or docs/hermes-
|
||||
kanban-v1-spec.pdf for the full design.
|
||||
|
||||
positional arguments:
|
||||
{init,boards,create,swarm,list,ls,show,assign,set-model,reclaim,reassign,diagnostics,diag,link,unlink,claim,comment,attach,attachments,attach-rm,complete,edit,block,schedule,unblock,request-review,request-changes,reopen-review,promote,archive,tail,dispatch,daemon,watch,stats,notify-subscribe,notify-list,notify-unsubscribe,log,runs,heartbeat,assignees,context,specify,decompose,gc,repair}
|
||||
init Create kanban.db if missing (idempotent)
|
||||
boards Manage kanban boards (one board per project /
|
||||
workstream)
|
||||
create Create a new task
|
||||
swarm Create a Kanban Swarm v1 graph (parallel workers →
|
||||
verifier → synthesizer)
|
||||
list (ls) List tasks
|
||||
show Show a task with comments + events
|
||||
assign Assign or reassign a task
|
||||
set-model Set or clear a task's model/provider override (takes
|
||||
effect on the next dispatch)
|
||||
reclaim Release an active worker claim on a running task
|
||||
reassign Reassign a task to a different profile, optionally
|
||||
reclaiming first
|
||||
diagnostics (diag) List active diagnostics on the current board
|
||||
link Add a parent->child dependency
|
||||
unlink Remove a parent->child dependency
|
||||
claim Atomically claim a ready task (prints resolved
|
||||
workspace path)
|
||||
comment Append a comment
|
||||
attach Attach a local file to a task
|
||||
attachments List a task's attachments
|
||||
attach-rm Delete an attachment by id
|
||||
complete Mark one or more tasks done
|
||||
edit Edit recovery fields on an already-completed task
|
||||
block Mark one or more tasks blocked
|
||||
schedule Park one or more tasks in Scheduled (waiting on time,
|
||||
not human input)
|
||||
unblock Return blocked/scheduled tasks to ready, or todo while
|
||||
parents remain open
|
||||
request-review Move a task to 'review' (implementation done, awaiting
|
||||
review) — NOT a block
|
||||
request-changes Reviewer verdict: return the active review run to its
|
||||
implementer
|
||||
reopen-review Send one or more review tasks back for changes (review
|
||||
-> ready/todo)
|
||||
promote Manually move one or more todo/blocked tasks to ready
|
||||
(recovery path)
|
||||
archive Archive one or more tasks
|
||||
tail Follow a task's event stream
|
||||
dispatch One dispatcher pass: reclaim stale, promote ready,
|
||||
spawn workers
|
||||
daemon DEPRECATED — dispatcher now runs in the gateway. Use
|
||||
`hermes gateway start`.
|
||||
watch Live-stream task_events to the terminal (Ctrl+C to
|
||||
exit)
|
||||
stats Per-status + per-assignee counts + oldest-ready age
|
||||
notify-subscribe Subscribe a gateway source to a task's terminal events
|
||||
(used by /kanban subscribe in the gateway adapter)
|
||||
notify-list List notification subscriptions (optionally for a
|
||||
single task)
|
||||
notify-unsubscribe Remove a gateway subscription from a task
|
||||
log Print the worker log for a task (from <kanban-
|
||||
root>/kanban/logs/)
|
||||
runs Show attempt history for a task (one row per run:
|
||||
profile, outcome, elapsed, summary)
|
||||
heartbeat Emit a heartbeat event for a running task (worker
|
||||
liveness signal)
|
||||
assignees List known profiles + per-profile task counts (union
|
||||
of ~/.hermes/profiles/ and current assignees on the
|
||||
board)
|
||||
context Print the full context a worker sees for a task (title
|
||||
+ body + parent results + comments).
|
||||
specify Flesh out a triage-column task into a concrete spec
|
||||
(title + body) and promote it to todo. Uses the
|
||||
auxiliary LLM configured under
|
||||
auxiliary.triage_specifier.
|
||||
decompose Decompose a triage-column task into a graph of child
|
||||
tasks routed to specialist profiles by description.
|
||||
Falls back to specify-style single-task promotion when
|
||||
the task doesn't benefit from fan-out. Uses
|
||||
auxiliary.kanban_decomposer.
|
||||
gc Garbage-collect archived-task workspaces, old events,
|
||||
and old logs
|
||||
repair Check kanban.db integrity and auto-repair index-only
|
||||
corruption
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--board <slug> Board slug to operate on. Defaults to the current
|
||||
board (set via `hermes kanban boards switch <slug>` or
|
||||
the HERMES_KANBAN_BOARD env var). Use `hermes kanban
|
||||
boards list` to see all boards.
|
||||
===== hermes skills --help =====
|
||||
usage: hermes skills [-h]
|
||||
{trust,untrust,browse,search,install,inspect,list,check,update,audit,uninstall,reset,list-modified,diff,opt-out,opt-in,repair-official,publish,snapshot,tap,config} ...
|
||||
|
||||
Search, install, inspect, audit, configure, and manage skills from skills.sh,
|
||||
well-known agent skill endpoints, GitHub, ClawHub, and other registries.
|
||||
|
||||
positional arguments:
|
||||
{trust,untrust,browse,search,install,inspect,list,check,update,audit,uninstall,reset,list-modified,diff,opt-out,opt-in,repair-official,publish,snapshot,tap,config}
|
||||
trust Trust a project so its repo-local skills
|
||||
(./.hermes/skills, ./.agents/skills) load
|
||||
untrust Revoke project-skill trust for a repo
|
||||
browse Browse all available skills (paginated)
|
||||
search Search skill registries
|
||||
install Install a skill
|
||||
inspect Preview a skill without installing
|
||||
list List installed skills
|
||||
check Check installed hub skills for updates
|
||||
update Update installed hub skills
|
||||
audit Re-scan installed hub skills
|
||||
uninstall Remove a hub-installed skill
|
||||
reset Reset a bundled skill — clears 'user-modified'
|
||||
tracking so updates work again
|
||||
list-modified List bundled skills you've edited (which `hermes
|
||||
update` keeps)
|
||||
diff Show how your copy of a bundled skill differs from the
|
||||
stock version
|
||||
opt-out Stop bundled skills from being seeded into this
|
||||
profile
|
||||
opt-in Re-enable bundled-skill seeding (undo opt-out)
|
||||
repair-official Backfill or restore official optional skills from repo
|
||||
source
|
||||
publish Publish a skill to a registry
|
||||
snapshot Export/import skill configurations
|
||||
tap Manage skill sources
|
||||
config Interactive skill configuration — enable/disable
|
||||
individual skills
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes sessions --help =====
|
||||
usage: hermes sessions [-h]
|
||||
{list,export,delete,prune,archive,optimize,clean-markers,optimize-storage,repair,repair-routing,recover,stats,rename,pin,unpin,pinned,retitle-skills,browse,import} ...
|
||||
|
||||
View and manage the SQLite session store
|
||||
|
||||
positional arguments:
|
||||
{list,export,delete,prune,archive,optimize,clean-markers,optimize-storage,repair,repair-routing,recover,stats,rename,pin,unpin,pinned,retitle-skills,browse,import}
|
||||
list List recent sessions
|
||||
export Export sessions to JSONL, Markdown, or QMD
|
||||
delete Delete a specific session
|
||||
prune Delete old sessions (filterable by time window,
|
||||
source, title, ...)
|
||||
archive Bulk-archive (soft-hide) sessions matching filters —
|
||||
no deletion
|
||||
optimize Reclaim disk space: merge FTS5 segments + VACUUM (no
|
||||
data change)
|
||||
clean-markers Permanently clear stale tool-call marker content left
|
||||
by sessions from before #78148
|
||||
optimize-storage Migrate the search index to the compact v23 layout
|
||||
(reclaims disk on large DBs)
|
||||
repair Repair a malformed state.db schema so hidden sessions
|
||||
reappear
|
||||
repair-routing Re-stamp gateway sessions that lost their routing
|
||||
identity
|
||||
recover Rebuild canonical session data into a separate clean
|
||||
database
|
||||
stats Show session store statistics
|
||||
rename Set or change a session's title
|
||||
pin Pin session(s) — durable keep flag, exempt from auto-
|
||||
archive
|
||||
unpin Remove the pin (durable keep flag) from session(s)
|
||||
pinned List pinned sessions
|
||||
retitle-skills Re-title sessions whose auto-title came from a
|
||||
/skill's own text
|
||||
browse Interactive session picker — browse, search, and
|
||||
resume sessions
|
||||
import Import a Claude Code or Codex CLI session into Hermes
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes mcp --help =====
|
||||
usage: hermes mcp [-h] [--accept-hooks]
|
||||
{serve,add,remove,rm,list,ls,test,configure,config,login,reauth,picker,catalog,install} ...
|
||||
|
||||
Manage MCP server connections and run Hermes as an MCP server. MCP servers
|
||||
provide additional tools via the Model Context Protocol. Use 'hermes mcp add'
|
||||
to connect to a new server, or 'hermes mcp serve' to expose Hermes
|
||||
conversations over MCP.
|
||||
|
||||
positional arguments:
|
||||
{serve,add,remove,rm,list,ls,test,configure,config,login,reauth,picker,catalog,install}
|
||||
serve Run Hermes as an MCP server (expose conversations to
|
||||
other agents)
|
||||
add Add an MCP server (discovery-first install)
|
||||
remove (rm) Remove an MCP server
|
||||
list (ls) List configured MCP servers
|
||||
test Test MCP server connection
|
||||
configure (config) Toggle tool selection
|
||||
login Force re-authentication for an OAuth-based MCP server
|
||||
reauth Re-authenticate one OAuth MCP server, or all of them
|
||||
(--all)
|
||||
picker Interactive catalog picker (also the default for
|
||||
`hermes mcp`)
|
||||
catalog List Nous-approved MCPs available for one-click
|
||||
install
|
||||
install Install a catalog MCP by name (e.g. `hermes mcp
|
||||
install n8n`)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
|
||||
(equivalent to HERMES_ACCEPT_HOOKS=1 /
|
||||
hooks_auto_accept: true).
|
||||
===== hermes profile --help =====
|
||||
usage: hermes profile [-h]
|
||||
{list,use,create,delete,describe,show,alias,rename,export,import,install,update,info} ...
|
||||
|
||||
positional arguments:
|
||||
{list,use,create,delete,describe,show,alias,rename,export,import,install,update,info}
|
||||
list List all profiles
|
||||
use Set sticky default profile
|
||||
create Create a new profile
|
||||
delete Delete a profile
|
||||
describe Read or set a profile's description (used by the
|
||||
kanban orchestrator)
|
||||
show Show profile details
|
||||
alias Manage wrapper scripts
|
||||
rename Rename a profile ('default': sets a display name; id
|
||||
unchanged)
|
||||
export Export a profile to archive
|
||||
import Import a profile from archive
|
||||
install Install a profile distribution from a git URL or local
|
||||
directory
|
||||
update Re-pull a distribution and apply updates (user data
|
||||
preserved)
|
||||
info Show a profile's distribution manifest (version,
|
||||
requirements, source)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes memory --help =====
|
||||
usage: hermes memory [-h] {setup,status,off,reset} ...
|
||||
|
||||
Set up and manage external memory provider plugins. Available providers:
|
||||
honcho, openviking, mem0, hindsight, holographic, retaindb, byterover. Only
|
||||
one external provider can be active at a time. Built-in memory
|
||||
(MEMORY.md/USER.md) is always active.
|
||||
|
||||
positional arguments:
|
||||
{setup,status,off,reset}
|
||||
setup Interactive provider selection and configuration
|
||||
status Show current memory provider config
|
||||
off Disable external provider (built-in only)
|
||||
reset Erase all built-in memory (MEMORY.md and USER.md)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes tools --help =====
|
||||
usage: hermes tools [-h] [--summary] {list,disable,enable,post-setup} ...
|
||||
|
||||
Enable, disable, or list tools for CLI, Telegram, Discord, etc. Built-in
|
||||
toolsets use plain names (e.g. web, memory). MCP tools use server:tool
|
||||
notation (e.g. github:create_issue). Run 'hermes tools' with no subcommand for
|
||||
the interactive configuration UI.
|
||||
|
||||
positional arguments:
|
||||
{list,disable,enable,post-setup}
|
||||
list Show all tools and their enabled/disabled status
|
||||
disable Disable toolsets or MCP tools
|
||||
enable Enable toolsets or MCP tools
|
||||
post-setup Run a provider's post-setup install hook
|
||||
(npm/pip/binary)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--summary Print a summary of enabled tools per platform and exit
|
||||
===== hermes project --help =====
|
||||
usage: hermes project [-h]
|
||||
{create,list,ls,show,add-folder,remove-folder,rename,set-primary,use,archive,restore,bind-board} ...
|
||||
|
||||
Projects are human-named workspaces that can span multiple folders / repos.
|
||||
They anchor desktop session grouping and, when bound to a kanban board, give
|
||||
tasks a deterministic worktree + branch convention. State is per-profile.
|
||||
|
||||
positional arguments:
|
||||
{create,list,ls,show,add-folder,remove-folder,rename,set-primary,use,archive,restore,bind-board}
|
||||
create Create a new project
|
||||
list (ls) List projects
|
||||
show Show a project's details
|
||||
add-folder Add a folder to a project
|
||||
remove-folder Remove a folder from a project
|
||||
rename Rename a project
|
||||
set-primary Set the primary folder
|
||||
use Set the active project
|
||||
archive Archive a project
|
||||
restore Restore an archived project
|
||||
bind-board Bind a kanban board to a project
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes gateway --help =====
|
||||
usage: hermes gateway [-h] [--accept-hooks]
|
||||
{run,start,stop,restart,status,install,uninstall,list,setup,migrate-legacy,enroll} ...
|
||||
|
||||
Manage the messaging gateway (Telegram, Discord, WhatsApp, Weixin, and more)
|
||||
|
||||
positional arguments:
|
||||
{run,start,stop,restart,status,install,uninstall,list,setup,migrate-legacy,enroll}
|
||||
run Run gateway in foreground (recommended for WSL,
|
||||
Docker, Termux)
|
||||
start Start the installed systemd/launchd background service
|
||||
stop Stop gateway service
|
||||
restart Restart gateway service
|
||||
status Show gateway status
|
||||
install Install gateway as a systemd/launchd background
|
||||
service
|
||||
uninstall Uninstall gateway service
|
||||
list List all profiles and their gateway status
|
||||
setup Configure messaging platforms
|
||||
migrate-legacy Remove legacy hermes.service units from pre-rename
|
||||
installs
|
||||
enroll Enroll this gateway with a relay connector (writes
|
||||
relay auth creds to .env)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
|
||||
(equivalent to HERMES_ACCEPT_HOOKS=1 /
|
||||
hooks_auto_accept: true).
|
||||
===== hermes computer-use --help =====
|
||||
usage: hermes computer-use [-h] {install,status,doctor,permissions} ...
|
||||
|
||||
Install or check the cua-driver binary used by the `computer_use` toolset.
|
||||
Supported on macOS, Windows, and Linux. Use `hermes computer-use install` to
|
||||
fetch and run the upstream cua-driver installer. This is equivalent to the
|
||||
post-setup hook that `hermes tools` runs when you first enable the Computer
|
||||
Use toolset, and is a stable target for re-running the install if it didn't
|
||||
fire (e.g. when toggling the toolset on a returning-user setup). Use `hermes
|
||||
computer-use doctor` to run cua-driver's `health_report` MCP tool and surface
|
||||
its check matrix (TCC, bundle identity, version, platform support, ...) in
|
||||
human-readable form.
|
||||
|
||||
positional arguments:
|
||||
{install,status,doctor,permissions}
|
||||
install Install or repair the cua-driver binary
|
||||
(macOS/Windows/Linux)
|
||||
status Print whether cua-driver is installed and on PATH
|
||||
doctor Run cua-driver `health_report` and surface the check
|
||||
matrix
|
||||
permissions Check or grant macOS Accessibility + Screen Recording
|
||||
(macOS)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes doctor --help =====
|
||||
usage: hermes doctor [-h] [--fix] [--live] [--ack ADVISORY_ID]
|
||||
|
||||
Diagnose issues with Hermes Agent setup
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--fix Attempt to fix issues automatically
|
||||
--live Opt-in: run one bounded, read-only real-call health probe
|
||||
per configured tool backend
|
||||
(Firecrawl/FAL/browser/MCP/TTS/STT) after the static
|
||||
checks. Makes real network calls.
|
||||
--ack ADVISORY_ID Acknowledge a security advisory by ID and exit. After
|
||||
ack, the advisory will no longer trigger startup banners.
|
||||
Run `hermes doctor` first to see active advisories and
|
||||
their IDs.
|
||||
===== hermes status --help =====
|
||||
usage: hermes status [-h] [--all] [--deep]
|
||||
|
||||
Display status of Hermes Agent components
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--all Show all details (redacted for sharing)
|
||||
--deep Run deep checks (may take longer)
|
||||
===== hermes import-agent --help =====
|
||||
usage: hermes import-agent [-h] [--source SOURCE] [--dry-run] [--overwrite]
|
||||
[--yes]
|
||||
[{claude-code,codex}]
|
||||
|
||||
One-command import of another coding agent's setup into Hermes. Maps
|
||||
CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills,
|
||||
and memories into their Hermes equivalents. Always shows a preview before
|
||||
making changes. API keys and credentials are never imported — run 'hermes
|
||||
setup' for those.
|
||||
|
||||
positional arguments:
|
||||
{claude-code,codex} Which agent to import from (default: auto-detect
|
||||
~/.claude or ~/.codex)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--source SOURCE Path to the agent's config directory (default:
|
||||
~/.claude or ~/.codex)
|
||||
--dry-run Preview only — stop after showing what would be
|
||||
imported
|
||||
--overwrite Overwrite existing Hermes items on name conflicts
|
||||
(default: skip)
|
||||
--yes, -y Skip confirmation prompts
|
||||
@@ -0,0 +1,33 @@
|
||||
import json, subprocess, re
|
||||
|
||||
data = json.load(open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json"))
|
||||
have = {r["videoId"] for r in data}
|
||||
extra = []
|
||||
for v in ["8GjyOQy19so", "zwqhemjHq3E"]:
|
||||
if v in have:
|
||||
continue
|
||||
url = f"https://www.youtube.com/watch?v={v}"
|
||||
oe = subprocess.run(["curl", "-s", f"https://www.youtube.com/oembed?url={url}&format=json"],
|
||||
capture_output=True, text=True, timeout=30)
|
||||
try:
|
||||
oej = json.loads(oe.stdout)
|
||||
o = {"status": "PASS", "title": oej.get("title"), "author": oej.get("author_name")}
|
||||
except Exception:
|
||||
o = {"status": "FAIL", "raw": oe.stdout[:200]}
|
||||
wp = subprocess.run(["curl", "-s", "-L", url,
|
||||
"-H", "User-Agent: Mozilla/5.0 (Windows NT 10.0) Chrome/124.0",
|
||||
"-H", "Accept-Language: en-US"], capture_output=True, text=True, timeout=30)
|
||||
html = wp.stdout
|
||||
pub = re.search(r'"publishDate":"([\d\-T:Z]+)"', html)
|
||||
dur = re.search(r'"lengthSeconds":"(\d+)"', html)
|
||||
views = re.search(r'"viewCount":"(\d+)"', html)
|
||||
rec = {"videoId": v, "oembed": o,
|
||||
"publishDate": pub.group(1) if pub else None,
|
||||
"lengthSeconds": int(dur.group(1)) if dur else None,
|
||||
"views": int(views.group(1)) if views else None}
|
||||
extra.append(rec)
|
||||
print(json.dumps(rec))
|
||||
|
||||
data.extend(extra)
|
||||
json.dump(data, open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json", "w"), indent=2)
|
||||
print("total videos in ledger:", len(data))
|
||||
@@ -0,0 +1,38 @@
|
||||
import json, subprocess, re, sys
|
||||
|
||||
vids = ["9GpWELm3_XI","iqN6MVzpJTk","Ta2wg6xPaY4","8tpuky8HpXw","tP6yf22OJdI",
|
||||
"UWjh5Z4s8jY","5d02TYoOzfE","83nWNRKZTCE","6M2tItdARew","5_N84t1rUU0",
|
||||
"UTZhvPXnmwA","1UgXUjT-QtI","6GtF_uHbGhw","CwPUOVUdApE","sCa3BtpkziQ",
|
||||
"jmtpYUOr7_U","cu2fgknmemA","4sAmpcSOVEw","P2LIFtrRr2U"]
|
||||
|
||||
results = []
|
||||
for v in vids:
|
||||
url = f"https://www.youtube.com/watch?v={v}"
|
||||
oe = subprocess.run(["curl","-s",f"https://www.youtube.com/oembed?url={url}&format=json"],
|
||||
capture_output=True, text=True, timeout=30)
|
||||
try:
|
||||
oej = json.loads(oe.stdout)
|
||||
oembed = {"status":"PASS","title":oej.get("title"),"author":oej.get("author_name")}
|
||||
except Exception:
|
||||
oembed = {"status":"FAIL","raw":oe.stdout[:200]}
|
||||
# watch page for date + duration
|
||||
wp = subprocess.run(["curl","-s","-L",url,"-H","User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/124.0",
|
||||
"-H","Accept-Language: en-US,en;q=0.9"], capture_output=True, text=True, timeout=30)
|
||||
html = wp.stdout
|
||||
pub = re.search(r'"publishDate":"([\d\-T:Z]+)"', html)
|
||||
upd = re.search(r'"uploadDate":"([\d\-T:Z]+)"', html)
|
||||
dur = re.search(r'"lengthSeconds":"(\d+)"', html)
|
||||
views = re.search(r'"viewCount":"(\d+)"', html)
|
||||
results.append({
|
||||
"videoId": v, "oembed": oembed,
|
||||
"publishDate": pub.group(1) if pub else None,
|
||||
"uploadDate": upd.group(1) if upd else None,
|
||||
"lengthSeconds": int(dur.group(1)) if dur else None,
|
||||
"views": int(views.group(1)) if views else None,
|
||||
"watchpage_bytes": len(html),
|
||||
})
|
||||
print(json.dumps(results[-1]))
|
||||
|
||||
with open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json","w") as f:
|
||||
json.dump(results, f, indent=2)
|
||||
print("saved video-verification.json")
|
||||
@@ -0,0 +1,271 @@
|
||||
[
|
||||
{
|
||||
"videoId": "9GpWELm3_XI",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent Explained In 5 Minutes",
|
||||
"author": "CodeHead"
|
||||
},
|
||||
"publishDate": "2026-05-23T08:00:27-07:00",
|
||||
"uploadDate": "2026-05-23T08:00:27-07:00",
|
||||
"lengthSeconds": 293,
|
||||
"views": 251275,
|
||||
"watchpage_bytes": 1419048
|
||||
},
|
||||
{
|
||||
"videoId": "iqN6MVzpJTk",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent by Nous Research: The Open-Source Agent Model Everyone Is Switching To",
|
||||
"author": "Praveen Govindaraj"
|
||||
},
|
||||
"publishDate": "2026-02-26T21:56:54-08:00",
|
||||
"uploadDate": "2026-02-26T21:56:54-08:00",
|
||||
"lengthSeconds": 193,
|
||||
"views": 3140,
|
||||
"watchpage_bytes": 1305120
|
||||
},
|
||||
{
|
||||
"videoId": "Ta2wg6xPaY4",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Learn 95% of Hermes Agent in 31 Minutes",
|
||||
"author": "Sharbel A."
|
||||
},
|
||||
"publishDate": "2026-08-09T07:00:19-07:00",
|
||||
"uploadDate": "2026-08-09T07:00:19-07:00",
|
||||
"lengthSeconds": 1888,
|
||||
"views": 129313,
|
||||
"watchpage_bytes": 1468130
|
||||
},
|
||||
{
|
||||
"videoId": "8tpuky8HpXw",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model",
|
||||
"author": "Tonbi's AI Garage"
|
||||
},
|
||||
"publishDate": "2026-03-11T07:00:14-07:00",
|
||||
"uploadDate": "2026-03-11T07:00:14-07:00",
|
||||
"lengthSeconds": 907,
|
||||
"views": 18167,
|
||||
"watchpage_bytes": 1326925
|
||||
},
|
||||
{
|
||||
"videoId": "tP6yf22OJdI",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Did Hermes Agent just kill OpenClaw? (full guide)",
|
||||
"author": "Alex Finn"
|
||||
},
|
||||
"publishDate": "2026-03-31T06:15:10-07:00",
|
||||
"uploadDate": "2026-03-31T06:15:10-07:00",
|
||||
"lengthSeconds": 835,
|
||||
"views": 132128,
|
||||
"watchpage_bytes": 1390199
|
||||
},
|
||||
{
|
||||
"videoId": "UWjh5Z4s8jY",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra",
|
||||
"author": "Peter Yang"
|
||||
},
|
||||
"publishDate": "2026-08-02T06:00:12-07:00",
|
||||
"uploadDate": "2026-08-02T06:00:12-07:00",
|
||||
"lengthSeconds": 2804,
|
||||
"views": 36613,
|
||||
"watchpage_bytes": 1386664
|
||||
},
|
||||
{
|
||||
"videoId": "5d02TYoOzfE",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Use This To Make The Hermes Agent Basically Free",
|
||||
"author": "AI LABS"
|
||||
},
|
||||
"publishDate": "2026-07-01T07:00:26-07:00",
|
||||
"uploadDate": "2026-07-01T07:00:26-07:00",
|
||||
"lengthSeconds": 788,
|
||||
"views": 54858,
|
||||
"watchpage_bytes": 1412833
|
||||
},
|
||||
{
|
||||
"videoId": "83nWNRKZTCE",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Meet the AI Agent That Grows With You Hermes Agent by Nous Research",
|
||||
"author": "Eddy Says Hi #EddySaysHi"
|
||||
},
|
||||
"publishDate": "2026-03-21T13:00:09-07:00",
|
||||
"uploadDate": "2026-03-21T13:00:09-07:00",
|
||||
"lengthSeconds": 366,
|
||||
"views": 465,
|
||||
"watchpage_bytes": 1254068
|
||||
},
|
||||
{
|
||||
"videoId": "6M2tItdARew",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "The AI Agent That Never Forgets: Meet Hermes Agent by Nous Research",
|
||||
"author": "Siggi"
|
||||
},
|
||||
"publishDate": "2026-03-11T13:53:17-07:00",
|
||||
"uploadDate": "2026-03-11T13:53:17-07:00",
|
||||
"lengthSeconds": 371,
|
||||
"views": 1199,
|
||||
"watchpage_bytes": 1267812
|
||||
},
|
||||
{
|
||||
"videoId": "5_N84t1rUU0",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent Fundamentals In 29 Minutes",
|
||||
"author": "Tina Huang"
|
||||
},
|
||||
"publishDate": "2026-07-20T09:37:02-07:00",
|
||||
"uploadDate": "2026-07-20T09:37:02-07:00",
|
||||
"lengthSeconds": 1780,
|
||||
"views": 462764,
|
||||
"watchpage_bytes": 1561142
|
||||
},
|
||||
{
|
||||
"videoId": "UTZhvPXnmwA",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent: Agents that grow with you |Episode #357|",
|
||||
"author": "Practical AI"
|
||||
},
|
||||
"publishDate": "2026-05-20T15:00:17-07:00",
|
||||
"uploadDate": "2026-05-20T15:00:17-07:00",
|
||||
"lengthSeconds": 2853,
|
||||
"views": 1888,
|
||||
"watchpage_bytes": 1329765
|
||||
},
|
||||
{
|
||||
"videoId": "1UgXUjT-QtI",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent: Why Everyone's Ditching OpenClaw in 2026",
|
||||
"author": "Luke Alexander AI"
|
||||
},
|
||||
"publishDate": "2026-03-26T07:10:29-07:00",
|
||||
"uploadDate": "2026-03-26T07:10:29-07:00",
|
||||
"lengthSeconds": 1083,
|
||||
"views": 13625,
|
||||
"watchpage_bytes": 1321848
|
||||
},
|
||||
{
|
||||
"videoId": "6GtF_uHbGhw",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Every Level of Hermes Agent Explained",
|
||||
"author": "Jack Roberts"
|
||||
},
|
||||
"publishDate": "2026-06-17T12:27:42-07:00",
|
||||
"uploadDate": "2026-06-17T12:27:42-07:00",
|
||||
"lengthSeconds": 1535,
|
||||
"views": 162977,
|
||||
"watchpage_bytes": 1577121
|
||||
},
|
||||
{
|
||||
"videoId": "CwPUOVUdApE",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent: The Ultimate Beginner\u2019s Guide",
|
||||
"author": "Metics Media"
|
||||
},
|
||||
"publishDate": "2026-04-24T06:59:04-07:00",
|
||||
"uploadDate": "2026-04-24T06:59:04-07:00",
|
||||
"lengthSeconds": 2228,
|
||||
"views": 118612,
|
||||
"watchpage_bytes": 1637578
|
||||
},
|
||||
{
|
||||
"videoId": "sCa3BtpkziQ",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "100 Days With Hermes Agent in 21 Minutes",
|
||||
"author": "Sharbel A."
|
||||
},
|
||||
"publishDate": "2026-06-17T07:47:26-07:00",
|
||||
"uploadDate": "2026-06-17T07:47:26-07:00",
|
||||
"lengthSeconds": 1279,
|
||||
"views": 57581,
|
||||
"watchpage_bytes": 1454230
|
||||
},
|
||||
{
|
||||
"videoId": "jmtpYUOr7_U",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent Just Killed OpenClaw (Full Tutorial)",
|
||||
"author": "Leon van Zyl"
|
||||
},
|
||||
"publishDate": "2026-04-28T04:19:38-07:00",
|
||||
"uploadDate": "2026-04-28T04:19:38-07:00",
|
||||
"lengthSeconds": 1199,
|
||||
"views": 15988,
|
||||
"watchpage_bytes": 1510488
|
||||
},
|
||||
{
|
||||
"videoId": "cu2fgknmemA",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent The 24/7 Self-Evolving AI Agent!",
|
||||
"author": "WorldofAI"
|
||||
},
|
||||
"publishDate": "2026-04-07T00:01:34-07:00",
|
||||
"uploadDate": "2026-04-07T00:01:34-07:00",
|
||||
"lengthSeconds": 555,
|
||||
"views": 46889,
|
||||
"watchpage_bytes": 1545788
|
||||
},
|
||||
{
|
||||
"videoId": "4sAmpcSOVEw",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent - Crash Course for Beginners (AI Agent)",
|
||||
"author": "Adrian Twarog"
|
||||
},
|
||||
"publishDate": "2026-07-21T01:20:41-07:00",
|
||||
"uploadDate": "2026-07-21T01:20:41-07:00",
|
||||
"lengthSeconds": 1338,
|
||||
"views": 46484,
|
||||
"watchpage_bytes": 1585548
|
||||
},
|
||||
{
|
||||
"videoId": "P2LIFtrRr2U",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent: New FREE OpenClaw Alternative!",
|
||||
"author": "Julian Goldie SEO"
|
||||
},
|
||||
"publishDate": "2026-03-09T14:00:32-07:00",
|
||||
"uploadDate": "2026-03-09T14:00:32-07:00",
|
||||
"lengthSeconds": 735,
|
||||
"views": 9975,
|
||||
"watchpage_bytes": 1353286
|
||||
},
|
||||
{
|
||||
"videoId": "8GjyOQy19so",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent Full Tutorial INSTALLATION + USECASES",
|
||||
"author": "CodeHead"
|
||||
},
|
||||
"publishDate": "2026-05-14T08:00:23-07:00",
|
||||
"lengthSeconds": 467,
|
||||
"views": 64285
|
||||
},
|
||||
{
|
||||
"videoId": "zwqhemjHq3E",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent vs OpenClaw",
|
||||
"author": "Sharbel A."
|
||||
},
|
||||
"publishDate": "2026-04-20T07:14:00-07:00",
|
||||
"lengthSeconds": 928,
|
||||
"views": 35360
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,10 @@
|
||||
import re, sys
|
||||
|
||||
html = open(sys.argv[1], encoding='utf-8', errors='ignore').read()
|
||||
ids = re.findall(r'"videoRenderer":\{"videoId":"([\w-]{11})"', html)
|
||||
print("videoRenderer hits:", len(set(ids)))
|
||||
for vid in dict.fromkeys(ids):
|
||||
m = re.search(r'"videoId":"%s".{0,3000}?"title":\{"runs":\[\{"text":"(.*?)"\}' % vid, html, re.S)
|
||||
ch = re.search(r'"videoId":"%s".{0,6000}?"ownerText":\{"runs":\[\{"text":"(.*?)"' % vid, html, re.S)
|
||||
dur = re.search(r'"videoId":"%s".{0,4000}?"lengthText":\{"accessibility".{0,400}?"simpleText":"(.*?)"' % vid, html, re.S)
|
||||
print((vid, m.group(1) if m else "?", ch.group(1) if ch else "?", dur.group(1) if dur else "?"))
|
||||
@@ -0,0 +1,42 @@
|
||||
# 06 — Sources
|
||||
|
||||
Access date for ALL entries: **2026-09-11** (via citation ledger `sources.py`; doc URLs additionally confirmed HTTP 200 by curl -L).
|
||||
|
||||
## Official docs (hermes-agent.nousresearch.com)
|
||||
| # | URL | Supported |
|
||||
|---|-----|-----------|
|
||||
| 1 | https://hermes-agent.nousresearch.com/docs | Docs index; overall feature map |
|
||||
| 3 | https://hermes-agent.nousresearch.com/docs/user-guide/configuration | Config sections, SOUL.md, checkpoints |
|
||||
| 4 | https://hermes-agent.nousresearch.com/docs/reference/slash-commands | Slash command registry (03) |
|
||||
| 5 | https://hermes-agent.nousresearch.com/docs/reference/tools-reference | Toolset inventory (02 §6) |
|
||||
| 6 | https://hermes-agent.nousresearch.com/docs/user-guide/features/cron | Cron surface (02 §9) |
|
||||
| 7 | https://hermes-agent.nousresearch.com/docs/user-guide/features/kanban | Kanban surface (02 §8) |
|
||||
| 8 | https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp | MCP surface + `hermes mcp serve` (02 §5, 04 P3) |
|
||||
| 9 | https://hermes-agent.nousresearch.com/docs/user-guide/features/memory | Memory surface (02 §2) |
|
||||
| 10 | https://hermes-agent.nousresearch.com/docs/user-guide/profiles | Profiles (02 §13) |
|
||||
| 11 | https://hermes-agent.nousresearch.com/docs/integrations/providers | Model/provider routing (02 §12) |
|
||||
| 12 | https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ | Gateway platforms (02 §15) |
|
||||
| 13 | https://hermes-agent.nousresearch.com/docs/user-guide/features/curator | Skill maintenance (02 §3) |
|
||||
| 14 | https://hermes-agent.nousresearch.com/docs/reference/cli-commands | CLI command index cross-check (03) |
|
||||
| 15 | https://hermes-agent.nousresearch.com/docs/reference/skills-catalog | Skills catalog (02 §3) |
|
||||
|
||||
## GitHub
|
||||
| # | URL | Supported |
|
||||
|---|-----|-----------|
|
||||
| 2 | https://github.com/nousresearch/hermes-agent | Repo identity, learning-loop description (01, 02) |
|
||||
|
||||
## Local primary sources (not web URLs; verified on this host)
|
||||
- Live CLI help output, Hermes Agent v0.21.1 (2026.9.7), reference install (Syslog kagentz): `hermes --help` + `hermes {chat,model,config,cron,kanban,skills,sessions,mcp,profile,memory,tools,project,gateway,computer-use,doctor,status,import-agent} --help` → raw dump `cli-help-dump.txt` (577 lines). Basis for all VERIFIED-LIVE tags in 01/02/03/04.
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/delegate-coding-agent/SKILL.md` (v1.0.0) + `references/claude-code.md` (v2.2.1) → 04 Patterns 1-2, 01 Claude Code column.
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/merge-reconciler/SKILL.md` → 04 Pattern 2.
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/computer-use/SKILL.md` (v2.0.0) → 04 Pattern 5, 02 §11.
|
||||
|
||||
## YouTube verification
|
||||
- YouTube search results pages (scraped 2026-09-11): `https://www.youtube.com/results?search_query=hermes+agent+nous+research` and `...nous+research+hermes+agent+official`
|
||||
- oEmbed endpoint per video: `https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=<ID>&format=json` — all 21 checked IDs returned PASS (HTTP 200, title/author match). Raw evidence incl. publishDate/lengthSeconds/viewCount per watch page: `video-verification.json`.
|
||||
- 17 listed in 05-videos.md + 4 spares; 21/21 pass, 0 fail.
|
||||
|
||||
## Explicit gaps (could not close)
|
||||
1. No official Nous Research-produced tutorial video was found — video list is third-party ecosystem content (disclosed in 05).
|
||||
2. Slash commands were verified DOC-ONLY (https://hermes-agent.nousresearch.com/docs/reference/slash-commands); they require an interactive session to exercise, which this headless run does not have. CLI equivalents were verified live.
|
||||
3. `hermes-agent.nousresearch.com/docs/developer-guide/` returned 404 — developer docs live in-repo (`AGENTS.md` in the GitHub repo), not as a docs site section.
|
||||
@@ -43,7 +43,7 @@ Docker hosts get special attention:
|
||||
|
||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
|
||||
|
||||
## Threat Levels (GUEST filesystems)
|
||||
## Threat Levels
|
||||
|
||||
| Level | Threshold | Response | Escalation |
|
||||
|-------|-----------|----------|------------|
|
||||
@@ -53,39 +53,6 @@ Docker hosts get special attention:
|
||||
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
|
||||
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
|
||||
|
||||
## Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)
|
||||
|
||||
Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are **report-only** — no automatic deletion of media or datastore content ever.
|
||||
|
||||
| Level | Threshold | Response | Escalation |
|
||||
|-------|-----------|----------|------------|
|
||||
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
||||
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
||||
| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert (state-change only) |
|
||||
|
||||
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
||||
|
||||
**State lives in a small JSON state file:** the scanner resolves it to an absolute path from the script's location — `$(dirname "$0")/state/host-disk-bands.json` (i.e., the `state/` directory next to the `scripts/` directory in the prose-contracts repo). The state file is gitignored runtime state — the scanner creates it on first run. Keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
|
||||
|
||||
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
|
||||
|
||||
```
|
||||
storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
|
||||
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
|
||||
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
|
||||
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
|
||||
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN
|
||||
```
|
||||
|
||||
**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.
|
||||
|
||||
**Action classes by volume type:**
|
||||
- **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
|
||||
- **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
|
||||
- **pbs-datastore** (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.
|
||||
|
||||
**Justification (measured 2026-09-15, firstmate):** storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
|
||||
@@ -161,10 +128,6 @@ from the `report_only_guests` YAML block above.
|
||||
|
||||
## Execution
|
||||
|
||||
### Host filesystems: report-only, NEVER auto-delete
|
||||
|
||||
Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.
|
||||
|
||||
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
|
||||
|
||||
**CT 111 / hostname `tdunna` / 192.168.68.129 is DETECT-AND-REPORT-ONLY.** It belongs to Theo.
|
||||
@@ -280,12 +243,6 @@ done
|
||||
|
||||
## Alert Templates
|
||||
|
||||
### HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)
|
||||
```
|
||||
⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
|
||||
Action: {volume_type-specific action}
|
||||
```
|
||||
|
||||
### AMBER (75-84%)
|
||||
```
|
||||
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
||||
|
||||
@@ -5,8 +5,7 @@ description: >
|
||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
|
||||
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
|
||||
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
|
||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
||||
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
||||
@@ -146,12 +145,6 @@ mcp_servers:
|
||||
url: http://192.168.68.65:3100/mcp
|
||||
timeout: 120
|
||||
connect_timeout: 60
|
||||
litellm:
|
||||
url: https://litellm.sysloggh.net/mcp
|
||||
headers:
|
||||
x-litellm-api-key: "Bearer <AGENT_KEY>" # Rule 15: must be a REAL key (sk-...), not an env-var name
|
||||
# Note: MCP endpoint requires Accept: application/json, text/event-stream header
|
||||
# This is handled by the MCP client library; don't add to config
|
||||
|
||||
# ─── Compression ───
|
||||
compression:
|
||||
@@ -224,40 +217,6 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
||||
3. **After update**: Restart Hermes on the agent host
|
||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||
|
||||
## MCP Server Configuration
|
||||
|
||||
MCP server entries in `mcp_servers:` must follow the format shown in the Template section.
|
||||
|
||||
**Header requirements (Rule 15):**
|
||||
- Use `headers:` field with a `x-litellm-api-key` entry
|
||||
- The value must be `"Bearer <REAL_KEY>"` where `<REAL_KEY>` is a literal LiteLLM virtual key
|
||||
- Do NOT use env-var references like `$LITELLM_API_KEY` — they resolve to empty strings in
|
||||
the static config and cause "Malformed API Key" errors (2026-08-07 Tanko incident)
|
||||
|
||||
**Key source:**
|
||||
- Keys are stored in the Infisical vault (project=agents, env=production)
|
||||
- For template-based config generation: substitute the agent's key from the agent_keys table
|
||||
- For manual config updates: retrieve the key from the vault and insert the literal value
|
||||
|
||||
**Verification (2026-08-07):**
|
||||
- Tested MCP initialize handshake against litellm.sysloggh.net/mcp with agent virtual key
|
||||
- Confirmed: 200 response with `serverInfo.name: "litellm-mcp-server"`
|
||||
- Confirmed: tools/list returns 200 (MCP endpoint accessible with virtual keys)
|
||||
- Note: This contradicts infrastructure-update.prose.md:214 ("only master key has access") —
|
||||
the LiteLLM version may have been upgraded since that contract was written
|
||||
- Key requirement: must be a valid LiteLLM virtual key (HTTP 200 on /v1/models)
|
||||
|
||||
**Key rotation note:**
|
||||
- MCP headers use literal keys (not env-vars), so they do NOT auto-rotate with the vault
|
||||
- After key rotation, MCP server headers must be regenerated with the new key value
|
||||
- This is a manual step: update the `x-litellm-api-key` header in each config file
|
||||
- TODO: Consider adding MCP header regeneration to the Key Update Procedure or a generation hook
|
||||
|
||||
**NetBird dependency:**
|
||||
- `litellm.sysloggh.net` is a NetBird endpoint (see Rule 5 and Infrastructure Stack table)
|
||||
- NetBird outages cause 502 errors on MCP requests, not auth failures
|
||||
- Diagnose: if MCP requests fail with 502, check NetBird status before investigating keys
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
@@ -496,28 +455,14 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
||||
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
||||
before and after any config change to catch this and all other rule violations.
|
||||
|
||||
### Rule 15: MCP Endpoint and Header Validation (UPDATED 2026-08-07)
|
||||
|
||||
**Endpoint validation:**
|
||||
- ra-h-os must point to `http://192.168.68.65:3100/mcp`
|
||||
- litellm must point to `https://litellm.sysloggh.net/mcp`
|
||||
- Mismatched endpoints cause silent failures (e.g., 2026-08-07 incident: Tanko's config had
|
||||
ra-h-os pointing to litellm's endpoint)
|
||||
|
||||
**Header validation:**
|
||||
- Every MCP entry with authentication must carry a `headers:` field
|
||||
- The header value must be a REAL key (e.g., `Bearer sk-abc123...`), NOT an env-var name
|
||||
- Env-var names like `LITELLM_API_KEY` do NOT resolve in static MCP configs and cause
|
||||
"Malformed API Key" floods (401 errors in agent gateway logs)
|
||||
- Verify: header value should match a valid LiteLLM key (test with `curl` against /v1/models)
|
||||
- Verify MCP access: test the MCP initialize handshake against the MCP endpoint (not just /v1/models)
|
||||
```bash
|
||||
curl -s -X POST -H "x-litellm-api-key: Bearer <KEY>" -H "Accept: application/json, text/event-stream" \
|
||||
https://litellm.sysloggh.net/mcp -d '{"jsonrpc":"2.0","id":1,"method":"initialize",...}' \
|
||||
| jq '.data.result.serverInfo' # should show serverInfo.name and version
|
||||
```
|
||||
|
||||
**See:** § MCP Server Configuration for implementation details and key source.
|
||||
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
|
||||
- Every MCP server entry must point at the correct endpoint:
|
||||
- ra-h-os = http://192.168.68.65:3100/mcp
|
||||
- litellm = https://litellm.sysloggh.net/mcp
|
||||
- MCP entries must carry a REAL key value in the header.
|
||||
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
|
||||
endpoints and result in "Malformed API Key" floods.
|
||||
- Ensure the header value is the actual key (e.g., `sk-...`).
|
||||
|
||||
## Execution
|
||||
|
||||
|
||||
@@ -101,7 +101,7 @@ auxiliary:
|
||||
fallback_providers:
|
||||
- provider: deepseek
|
||||
base_url: https://api.deepseek.com
|
||||
api_key: sk-synthetic-external-example # ← hardcoded OK (external, synthetic example)
|
||||
api_key: sk-b7d9... # ← hardcoded OK (external)
|
||||
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
|
||||
```
|
||||
|
||||
@@ -109,7 +109,7 @@ fallback_providers:
|
||||
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
|
||||
model:
|
||||
provider: harness
|
||||
api_key: sk-synthetic-example-12345 # ← RULE VIOLATION: hardcoded key (synthetic example)
|
||||
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION: hardcoded key
|
||||
|
||||
model:
|
||||
provider: harness
|
||||
@@ -139,16 +139,6 @@ scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### ACCEPTABLE PATTERN
|
||||
Agent keys live in `.env` or `.env.vault` files with 600 permissions (koonimo's shape is the canonical example). A plaintext key inside a `config.yaml` or any `config.yaml.bak-*` file is a violation — the backup files are not part of the runtime credential path and are not watched by the scanner, so a key in them is stale clutter that a future reader can mistake for a working key.
|
||||
|
||||
**Fix procedure** (when a backup file is found with a plaintext key):
|
||||
1. Move the file out of the scanned tree (e.g., `mv /root/.hermes/config.yaml.bak-* /root/hermes-config-backups/`) — do NOT delete the file, just move it so the scanner pattern no longer matches.
|
||||
2. Re-run the reachability check to confirm COMPLIANT.
|
||||
3. Report the before/after check output and the commands you ran.
|
||||
|
||||
**Rationale**: Moving the file preserves history without leaving a credential where a scanner trips over it. Deleting the file loses the historical context. Keeping it in place means the next scan will report it as a finding and waste time re-deciding.
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
@@ -182,7 +172,7 @@ grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
|
||||
|
||||
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
||||
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-…' /root/.config/systemd/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /root/.config/systemd/ 2>/dev/null
|
||||
|
||||
# 3. Verify running process env matches dedicated key
|
||||
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
||||
@@ -217,11 +207,9 @@ The agent picks up the new key via `infisical run --` at gateway startup.
|
||||
|
||||
**Keys are permanent and use bare agent name aliases.**
|
||||
|
||||
- **Duration**: `null` — keys never expire by default. **Expiry must be set EXPLICITLY at creation** with the `duration` parameter (e.g., `90d` for 90 days). The 90-day default is the standard; however, the config default is **NOT honoured** by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns `expires=null`). This has been recorded in `/opt/inference-harness/litellm_config.yaml` to prevent re-filing as a bug.
|
||||
- **Daily Audit**: A daily audit job runs at 00:00 UTC (`/usr/local/bin/litellm-key-renewal-ct116.sh`, cron 00:00). It is **AUDIT-ONLY** and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES `abiba-pi` and `koby` (report-only, and .129 must never be touched), and logs `RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED` when renewal is skipped. **Renewal is NOT implemented** — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end.
|
||||
- **Exclusions**: `abiba-pi` and every firstmate/secondmate/crewmate key stay **WITHOUT an expiry** until a proven renewal path exists. `koby` is **report-only** (never touched). These exclusions are enforced by the audit job.
|
||||
- **Duration**: `null` — keys never expire. NOT enforced today: CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key generated with no explicit models comes back with an empty models list. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
|
||||
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
|
||||
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use.
|
||||
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
|
||||
- **Max budget**: $100 per key (config default).
|
||||
|
||||
```yaml
|
||||
|
||||
@@ -181,8 +181,8 @@ description: >
|
||||
**Stirling-PDF** (deployed 2026-07-03, Authentik SSO 2026-07-03):
|
||||
- URL: `https://pdf.sysloggh.net` (public) / `http://192.168.68.7:8989` (direct)
|
||||
- Swagger: `http://192.168.68.7:8989/swagger-ui.html`
|
||||
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_USER»` / `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||
- API key: `«vault: infrastructure/production STIRLING_API_KEY»`
|
||||
- Admin credentials: `admin` / `kakashi20stirling`
|
||||
- API key: `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88`
|
||||
- Authentik OAuth2: configured but disabled (requires paid Server license). Ready to enable: set `SECURITY_OAUTH2_ENABLED=true` + `SECURITY_LOGINMETHOD=all`
|
||||
- Compose: `/opt/home_stack/docker-compose.yml`
|
||||
- Control script: `/opt/home_stack/infra-control.sh`
|
||||
@@ -364,79 +364,6 @@ For docker-vm specifically:
|
||||
- No PBS backup in 48h → fail
|
||||
```
|
||||
|
||||
### Backup Safety Preconditions (2026-09-15)
|
||||
|
||||
#### Background & Rationale
|
||||
|
||||
Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first:
|
||||
|
||||
1. **acerpve thin-pool VM 101** (acerpve, 192.168.68.9, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root.
|
||||
|
||||
2. **amdpve 0700 tmpdir** (amdpve, 192.168.68.15, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup.
|
||||
|
||||
3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood.
|
||||
|
||||
> ⚠️ **Hostname Resolution Warning (2026-09-15)**: The PVE node hostnames (acerpve, amdpve, minipve, storepve, ocupve) all resolve to the VPS (72.61.0.17, the Netbird VPS at srv1079750.hstgr.cloud) via the wildcard `*.dns.sysloggh.net` record, NOT to the actual nodes. So `ssh acerpve` lands on the VPS. **Nodes must be addressed by IP**: acerpve 192.168.68.9, amdpve 192.168.68.15, storepve 192.168.68.6, minipve 192.168.68.12, ocupve 192.168.68.5. Guest CTs are reached through their node (`pct exec`). Guest hostnames that resolve on the LAN (e.g. kagentz = 192.168.68.14) are fine. (The DNS address records are a separate decision — row: dag-daemon-node-hostnames-resolve-to-the-vps-20260915.)
|
||||
|
||||
#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts)
|
||||
|
||||
Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass:
|
||||
|
||||
```bash
|
||||
# Check 1: Pool headroom (PRIMARY - yields percentages directly)
|
||||
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname, see warning above
|
||||
lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data
|
||||
# Example output (acerpve, 192.168.68.9):
|
||||
# LV Data% Meta% LSize
|
||||
# data 29.95 1.22 <816.21g
|
||||
# Required thresholds (documented minimum):
|
||||
# data_percent < 90% (80% recommended for safety margin)
|
||||
# metadata_percent < 70% (metadata fills faster than data)
|
||||
|
||||
# Check 2: Verify pool is not in error state (dmsetup shows the raw DM device)
|
||||
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname
|
||||
dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1
|
||||
# dmsetup status pve-data-tpool field order (verified on 192.168.68.9):
|
||||
# $1=start $2=length $3="thin-pool" $4=transaction-id
|
||||
# $5=metadata_used/metadata_total (blocks) $6=data_used/data_total (sectors)
|
||||
# remaining fields are flags ("-", "rw", "discard_passdown", "queue_if_no_space", ...)
|
||||
# This is only used for the ERROR-STATE check; use the lvs command above for percentages.
|
||||
# metadata_percent = $5 / ($5 split by /) [second number in pair]
|
||||
```
|
||||
|
||||
**Minimum thresholds**: If either `data_percent >= 90%` or `metadata_percent >= 70%`, the backup MUST NOT start. State explicitly that these are hard stops, not warnings.
|
||||
|
||||
**Why this is a precondition**: A full or errored thin pool fails EVERY volume on the VG at once, including the host root. This is not a soft failure - it takes down the entire Proxmox host.
|
||||
|
||||
#### Staging Directory Requirement (2026-09-14 incident)
|
||||
|
||||
Any custom vzdump `tmpdir` MUST be world-traversable and writable exactly like `/var/tmp` (mode 1777). The archive step of vzdump runs in an unprivileged user namespace and cannot traverse a root-owned 0700 directory.
|
||||
|
||||
**Symptom to recognize**: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES` on every container in the backup run.
|
||||
|
||||
**Fix**: `chmod 1777 <custom-tmpdir>` before starting vzdump.
|
||||
|
||||
#### Task Start Rule for Truncating Shells
|
||||
|
||||
When starting a backup task from a shell that may truncate output (e.g., pipes, `head`), always use:
|
||||
|
||||
```bash
|
||||
pvesh create /storage/backup --output-format json -- ... | head -2
|
||||
# ❌ Can kill the backup task ("broken pipe" status)
|
||||
```
|
||||
|
||||
Instead, capture JSON output without piping to truncating commands:
|
||||
|
||||
```bash
|
||||
# Use --output-format json and capture to variable
|
||||
result=$(pvesh create /storage/backup --output-format json -- ...)
|
||||
# Then parse result if needed
|
||||
```
|
||||
|
||||
#### GPU Host Backup Status (acerpve VM 101)
|
||||
|
||||
VM 101 (llm-gpu) and VM 103 (ocu-llm) have NO scheduled backup job. Coverage is manual one-off runs only. This is intentional for VM 101 until the thin-pool failure mechanism is understood and documented.
|
||||
|
||||
## Section 5: Network Services — Monitoring
|
||||
|
||||
### 5.1 Service Inventory
|
||||
@@ -636,7 +563,7 @@ monitor, or integration breaks.
|
||||
```bash
|
||||
# Full cluster status
|
||||
PVE="https://minipve.sysloggh.net"
|
||||
AUTH="Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
|
||||
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
|
||||
|
||||
# Docker health from Abiba
|
||||
|
||||
@@ -144,8 +144,8 @@ through its agent wrapper.
|
||||
safety net for vault outage or token revocation. Must be kept in sync on rotation.
|
||||
Example:
|
||||
```bash
|
||||
MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY»
|
||||
MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY»
|
||||
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
|
||||
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
|
||||
```
|
||||
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
|
||||
```ini
|
||||
@@ -191,7 +191,7 @@ through its agent wrapper.
|
||||
### Tanko migration (COMPLETED 2026-07-17)
|
||||
|
||||
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
|
||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-synthetic-tanko-example…`)
|
||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
|
||||
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
|
||||
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
|
||||
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
|
||||
@@ -271,13 +271,13 @@ not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update ma
|
||||
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
||||
|
||||
**Key Storage:**
|
||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»…`)
|
||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
|
||||
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
||||
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
||||
(unlike fleet agents which require vault injection)
|
||||
|
||||
**Current Key (2026-09-01):**
|
||||
- **Prefix**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
- **Prefix**: `sk-or-v1-0af3f3…`
|
||||
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
- **Plan**: Paid (not free tier)
|
||||
- **Usage**: 0 (as of 2026-09-01)
|
||||
@@ -320,15 +320,7 @@ directly call OpenRouter via Python's requests library. Converting would require
|
||||
|
||||
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
|
||||
|
||||
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
|
||||
```bash
|
||||
# PRIMARY (proven, runs on CT 116 with no extra tooling):
|
||||
docker exec harness-litellm printenv LITELLM_MASTER_KEY
|
||||
# Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching)
|
||||
# The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it)
|
||||
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
|
||||
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
|
||||
```
|
||||
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
|
||||
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
|
||||
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
|
||||
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
|
||||
|
||||
@@ -19,7 +19,7 @@ description: >
|
||||
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||
- Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).
|
||||
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
|
||||
---
|
||||
|
||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||
|
||||
@@ -116,7 +116,7 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (hardcoded key → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM (deprecated key, no live usage).
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||
|
||||
## Maintains
|
||||
|
||||
|
||||
+10
-7
@@ -2,10 +2,16 @@
|
||||
kind: responsibility
|
||||
name: pm2-self-heal
|
||||
description: >
|
||||
PM2 process health check for abiba-telegram, abiba-zulip, gitea-runner, and zulip-watchdog.
|
||||
gpu-monitor is systemd-managed (gpu-monitor.service), NOT PM2.
|
||||
gpu-watchdog is decommissioned and folded into gpu-monitor.service.
|
||||
gitea-runner is KEPT. abiba-zulip is KEPT (online for days).
|
||||
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
||||
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
||||
errored. Logs every action to the knowledge graph and alerts the owner via
|
||||
Zulip DM on failures.
|
||||
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
||||
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
||||
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
||||
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
||||
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
||||
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -39,9 +45,6 @@ description: >
|
||||
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
||||
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
||||
spoton incident). Alerts include the restart count.
|
||||
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
|
||||
four processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog). The spoton
|
||||
reference above is historical context for the crash-loop guard, not a live process.
|
||||
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
||||
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
||||
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
||||
|
||||
@@ -142,21 +142,3 @@ each probe. If any probe returns non-200, flag as alert.
|
||||
- **PVE exporter metric schema**: NOT name-prefixed. `pve_cpu_usage_ratio`, `pve_memory_usage_bytes`, `pve_disk_usage_bytes`, `pve_uptime_seconds` are GUEST-level only (24 series, `id=lxc/100` etc). Node-level host metrics come from node_exporter. Storage pool usage: `pve_storage_info` (info only, no usage bytes — use node_filesystem_* for actual disk usage).
|
||||
- **grafana piechart plugin removed** from `GF_INSTALL_PLUGINS` (Angular, unsupported in Grafana 13).
|
||||
- **Single pve-exporter points at amdpve .15** — if amdpve API is down, cluster metrics gap (other node_exporters still report host metrics). Acceptable; amdpve is primary.
|
||||
## PBS Garbage Collection Schedule (storepve-datastore)
|
||||
|
||||
**Schedule:** Daily at 20:00 UTC (4:00 PM EDT)
|
||||
|
||||
**Rationale:** The nightly backup window runs 04:00–06:40 UTC (local backups 04:00–06:40Z, S3 sync 04:15Z, S3 trim 05:15Z). Running GC during this window causes avoidable I/O contention on the same datastore and host. 20:00 UTC lands well after the backup window closes and before the next day's backups begin.
|
||||
|
||||
**Volume:** `/tank/pbs-backup` on the ZFS pool `tank` (12.7T total, 11.3T free, 10% used). Inside CT 107, this appears as 1.5T total / 413G used / 1.1T avail = 28%.
|
||||
|
||||
**Important:** This is NOT the same as `/media/easystore2` (3.7T, 96% used, HOST-RED), which is a media library ("4K MOVIES") on the storepve host root filesystem. The PBS datastore lives on a separate ZFS pool.
|
||||
|
||||
**Cron:** `/etc/cron.d/pbs-gc` on storepve (192.168.68.6):
|
||||
```
|
||||
0 20 * * * root /usr/local/bin/pbs-gc.sh
|
||||
```
|
||||
|
||||
**Last GC run:** 2026-09-18 00:08:34 UTC (3h 19m 9s, removed 551.768 GiB)
|
||||
|
||||
**Schedule owner:** The `gc-schedule` field in the PBS datastore config is the durable owner of GC behavior. The cron is a fallback because `proxmox-backup-manager datastore update --gc-schedule` failed to parse the calendar event (Nom(Eof) error).
|
||||
|
||||
@@ -122,8 +122,10 @@ def _fail(key, agent_name=None):
|
||||
|
||||
|
||||
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
|
||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||
|
||||
# Fallback: if no env token, read the shared vault token file
|
||||
if not INFISICAL_TOKEN:
|
||||
# Fallback: read the shared vault token file
|
||||
_token_path = os.path.expanduser("~/.infisical-token")
|
||||
if os.path.isfile(_token_path):
|
||||
try:
|
||||
@@ -131,7 +133,6 @@ if not INFISICAL_TOKEN:
|
||||
INFISICAL_TOKEN = _f.read().strip()
|
||||
except (OSError, UnicodeDecodeError):
|
||||
pass
|
||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||
|
||||
# ── Helpers ──────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@@ -16,16 +16,14 @@ from email.mime.text import MIMEText
|
||||
from email.mime.multipart import MIMEMultipart
|
||||
|
||||
PVE = "https://192.168.68.12:8006"
|
||||
AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
|
||||
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
|
||||
# ── Shared credentials —─
|
||||
|
||||
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_API_KEY = os.environ.get("ZULIP_API_KEY", "")
|
||||
if not ZULIP_API_KEY:
|
||||
raise SystemExit("ZULIP_API_KEY not set — refusing to run with no credential")
|
||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_API_KEY}"
|
||||
ZULIP_KEY = "cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_KEY}"
|
||||
|
||||
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||
LITELLM_BACKEND = "192.168.68.116"
|
||||
@@ -671,10 +669,7 @@ def send_email(html_content, subject_prefix=""):
|
||||
msg.attach(MIMEText(html_content, "html"))
|
||||
|
||||
try:
|
||||
EMAIL_PASSWORD = os.environ.get("EMAIL_PASSWORD") or os.environ.get("SMTP_PASSWORD") or os.environ.get("MAIL_PASSWORD")
|
||||
if not EMAIL_PASSWORD:
|
||||
print("EMAIL_PASSWORD not set — refusing to send email", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
EMAIL_PASSWORD = "rgbuomwcydxwbszd"
|
||||
GMAIL_EMAIL = "jtabiri@gmail.com"
|
||||
|
||||
server = smtplib.SMTP("smtp.gmail.com", 587)
|
||||
|
||||
+6
-236
@@ -153,25 +153,6 @@ GPU_HOSTS = [
|
||||
CONNECT_TIMEOUT = 5
|
||||
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
|
||||
|
||||
# Host filesystem thresholds (from contract)
|
||||
HOST_THRESHOLDS = {
|
||||
"WARN": 85,
|
||||
"AMBER": 90,
|
||||
"RED": 95,
|
||||
}
|
||||
|
||||
# State file path (absolute, so execution context doesn't matter)
|
||||
STATE_FILE = pathlib.Path(__file__).resolve().parent.parent / "state" / "host-disk-bands.json"
|
||||
|
||||
# PVE nodes to probe for host filesystems
|
||||
HOST_NODES = [
|
||||
{"hostname": "acerpve", "ip": "192.168.68.9"},
|
||||
{"hostname": "amdpve", "ip": "192.168.68.15"},
|
||||
{"hostname": "storepve", "ip": "192.168.68.6"},
|
||||
{"hostname": "minipve", "ip": "192.168.68.12"},
|
||||
{"hostname": "ocupve", "ip": "192.168.68.5"},
|
||||
]
|
||||
|
||||
|
||||
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
|
||||
"""Run a command and return (exit_code, stdout, stderr)."""
|
||||
@@ -317,151 +298,6 @@ def scan_fleet() -> list[dict]:
|
||||
return results
|
||||
|
||||
|
||||
def classify_band(usage_pct: float) -> str:
|
||||
"""Classify a percentage into a band."""
|
||||
if usage_pct >= HOST_THRESHOLDS["RED"]:
|
||||
return "HOST-RED"
|
||||
elif usage_pct >= HOST_THRESHOLDS["AMBER"]:
|
||||
return "HOST-AMBER"
|
||||
elif usage_pct >= HOST_THRESHOLDS["WARN"]:
|
||||
return "HOST-WARN"
|
||||
else:
|
||||
return "GREEN"
|
||||
|
||||
|
||||
def probe_host_filesystems() -> tuple[list[dict], dict[str, str]]:
|
||||
"""Probe host filesystems on all PVE nodes.
|
||||
|
||||
Returns:
|
||||
- List of host filesystem results
|
||||
- Dict of volume_key -> current_band (for state file)
|
||||
"""
|
||||
results = []
|
||||
current_bands = {}
|
||||
|
||||
for node in HOST_NODES:
|
||||
ip = node["ip"]
|
||||
hostname = node["hostname"]
|
||||
|
||||
# Probe df for host filesystems
|
||||
probe_cmd = f'ssh {SSH_OPTS} root@{ip} "df -hP / /media/* tank 2>/dev/null | tail -n +2"'
|
||||
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
|
||||
|
||||
if exit_code != 0:
|
||||
results.append({
|
||||
"target": f"{hostname} ({ip})",
|
||||
"hostname": hostname,
|
||||
"ip": ip,
|
||||
"reachable": False,
|
||||
"volumes": [],
|
||||
"probe_cmd": probe_cmd,
|
||||
"failure_kind": "ssh-error",
|
||||
})
|
||||
continue
|
||||
|
||||
# Parse df output and classify each volume
|
||||
volumes = []
|
||||
for line in stdout.splitlines():
|
||||
parts = line.split()
|
||||
if len(parts) < 6:
|
||||
continue
|
||||
|
||||
dev, size, used, avail, pct_str, mount = parts[:6]
|
||||
pct = float(pct_str.rstrip("%"))
|
||||
band = classify_band(pct)
|
||||
|
||||
# Volume type classification
|
||||
if mount.startswith("/media/"):
|
||||
vol_type = "media"
|
||||
elif mount == "/" or "pve" in dev:
|
||||
vol_type = "host-root"
|
||||
elif mount == "tank" or "tank" in mount:
|
||||
vol_type = "pbs-datastore"
|
||||
else:
|
||||
vol_type = "other"
|
||||
|
||||
# Volume key for state file (host/volume)
|
||||
volume_key = f"{hostname}/{mount}"
|
||||
current_bands[volume_key] = band
|
||||
|
||||
volumes.append({
|
||||
"mount": mount,
|
||||
"device": dev,
|
||||
"size": size,
|
||||
"used": used,
|
||||
"avail": avail,
|
||||
"pct": pct,
|
||||
"band": band,
|
||||
"type": vol_type,
|
||||
})
|
||||
|
||||
results.append({
|
||||
"target": f"{hostname} ({ip})",
|
||||
"hostname": hostname,
|
||||
"ip": ip,
|
||||
"reachable": True,
|
||||
"volumes": volumes,
|
||||
"probe_cmd": probe_cmd,
|
||||
"failure_kind": None,
|
||||
})
|
||||
|
||||
return results, current_bands
|
||||
|
||||
|
||||
def read_state_file() -> Optional[dict[str, str]]:
|
||||
"""Read the state file if it exists."""
|
||||
if not STATE_FILE.exists():
|
||||
return None
|
||||
try:
|
||||
with open(STATE_FILE) as f:
|
||||
return json.load(f)
|
||||
except (json.JSONDecodeError, IOError) as e:
|
||||
print(f"⚠️ State file exists but unreadable: {e}", file=sys.stderr)
|
||||
return {}
|
||||
|
||||
|
||||
def write_state_file(bands: dict[str, str]) -> None:
|
||||
"""Write the state file."""
|
||||
STATE_FILE.parent.mkdir(parents=True, exist_ok=True)
|
||||
try:
|
||||
with open(STATE_FILE, "w") as f:
|
||||
json.dump(bands, f, indent=2)
|
||||
except IOError as e:
|
||||
print(f"⚠️ State file write failed: {e}", file=sys.stderr)
|
||||
|
||||
|
||||
def detect_transitions(current_bands: dict[str, str], prior_bands: Optional[dict[str, str]]) -> list[dict]:
|
||||
"""Detect band transitions (current vs. prior)."""
|
||||
if prior_bands is None:
|
||||
# First run — no transitions, just establish baseline
|
||||
return []
|
||||
|
||||
transitions = []
|
||||
# Check for volumes that moved to a higher band (escalation)
|
||||
for volume, current_band in current_bands.items():
|
||||
prior_band = prior_bands.get(volume, "GREEN")
|
||||
|
||||
# Band ordering: GREEN < HOST-WARN < HOST-AMBER < HOST-RED
|
||||
band_order = {"GREEN": 0, "HOST-WARN": 1, "HOST-AMBER": 2, "HOST-RED": 3}
|
||||
|
||||
if band_order[current_band] > band_order[prior_band]:
|
||||
transitions.append({
|
||||
"type": "escalation",
|
||||
"volume": volume,
|
||||
"from": prior_band,
|
||||
"to": current_band,
|
||||
})
|
||||
elif band_order[current_band] < band_order[prior_band]:
|
||||
transitions.append({
|
||||
"type": "recovery",
|
||||
"volume": volume,
|
||||
"from": prior_band,
|
||||
"to": current_band,
|
||||
})
|
||||
|
||||
return transitions
|
||||
|
||||
|
||||
def render_results(results: list[dict]) -> str:
|
||||
"""Render scan results in human-readable format."""
|
||||
lines = []
|
||||
@@ -480,88 +316,22 @@ def render_results(results: list[dict]) -> str:
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def render_host_results(results: list[dict], transitions: list[dict], prior_bands: Optional[dict[str, str]]) -> str:
|
||||
"""Render host filesystem results in human-readable format."""
|
||||
lines = []
|
||||
lines.append("")
|
||||
lines.append("=== Host Filesystem Bands ===")
|
||||
lines.append("")
|
||||
|
||||
# Render transitions first (they're the actionable alerts)
|
||||
if prior_bands is None:
|
||||
lines.append(" (first run — recording baseline, no alerts)")
|
||||
elif not transitions:
|
||||
lines.append(" (no band changes since last scan)")
|
||||
else:
|
||||
for t in transitions:
|
||||
volume, from_band, to_band = t["volume"], t["from"], t["to"]
|
||||
if t["type"] == "escalation":
|
||||
lines.append(f" ⚠️ {volume}: {from_band} → {to_band} (ESCALATION)")
|
||||
else:
|
||||
lines.append(f" ✅ {volume}: {from_band} → {to_band} (RECOVERY)")
|
||||
|
||||
# Render all volumes with their bands
|
||||
lines.append("")
|
||||
for node_result in results:
|
||||
if not node_result["reachable"]:
|
||||
lines.append(f" ❌ {node_result['target']}: UNREACHABLE ({node_result['failure_kind']})")
|
||||
continue
|
||||
|
||||
lines.append(f" {node_result['target']}:")
|
||||
for vol in node_result["volumes"]:
|
||||
lines.append(f" {vol['mount']} ({vol['type']}): {vol['pct']}% ({vol['used']}/{vol['size']}, {vol['avail']} free) -> {vol['band']}")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
import argparse
|
||||
|
||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests and host filesystems.")
|
||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
|
||||
ap.add_argument("--json", action="store_true", help="machine-readable output")
|
||||
ap.add_argument("--hosts-only", action="store_true", help="scan host filesystems only")
|
||||
ap.add_argument("--guests-only", action="store_true", help="scan guests only (skip host filesystems)")
|
||||
args = ap.parse_args()
|
||||
|
||||
# Scan guests (unless --hosts-only)
|
||||
guest_results = []
|
||||
if not args.hosts_only:
|
||||
guest_results = scan_fleet()
|
||||
|
||||
# Scan host filesystems (unless --guests-only)
|
||||
host_results = []
|
||||
current_bands = {}
|
||||
if not args.guests_only:
|
||||
host_results, current_bands = probe_host_filesystems()
|
||||
|
||||
# Read prior state and detect transitions
|
||||
prior_bands = read_state_file()
|
||||
transitions = detect_transitions(current_bands, prior_bands)
|
||||
|
||||
# Write new state
|
||||
write_state_file(current_bands)
|
||||
else:
|
||||
prior_bands = None
|
||||
transitions = []
|
||||
results = scan_fleet()
|
||||
|
||||
if args.json:
|
||||
# JSON output
|
||||
output = {
|
||||
"guests": guest_results,
|
||||
"hosts": host_results,
|
||||
"transitions": transitions,
|
||||
"prior_bands": prior_bands,
|
||||
}
|
||||
print(json.dumps(output, indent=2))
|
||||
print(json.dumps(results, indent=2))
|
||||
else:
|
||||
# Human-readable output
|
||||
if guest_results:
|
||||
print(render_results(guest_results))
|
||||
|
||||
if host_results:
|
||||
print(render_host_results(host_results, transitions, prior_bands))
|
||||
print(render_results(results))
|
||||
|
||||
# Exit 0 if all probed (reachable or not), 1 if any probe error
|
||||
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
|
||||
# (a probe error means the probe itself failed, not just that the guest was unreachable)
|
||||
return 0
|
||||
|
||||
|
||||
|
||||
@@ -35,12 +35,7 @@ def run_command(cmd, timeout=15):
|
||||
return 1, "", str(e)
|
||||
|
||||
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
|
||||
"""Probe HTTP endpoint and return (status_code, failure_kind)
|
||||
|
||||
Returns:
|
||||
(code, None) if successful or HTTP response received
|
||||
(000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc.
|
||||
"""
|
||||
"""Probe HTTP endpoint and return status code"""
|
||||
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
|
||||
if method == "POST":
|
||||
cmd += " -X POST"
|
||||
@@ -52,47 +47,15 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
||||
cmd += " -L"
|
||||
cmd += " '" + url + "'"
|
||||
|
||||
try:
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
if rc != 0:
|
||||
# Determine failure kind from curl exit code
|
||||
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
||||
if rc == 28:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
elif rc == 7:
|
||||
return (000, "connection refused")
|
||||
elif rc == 6:
|
||||
return (000, "dns failure")
|
||||
elif rc == 35:
|
||||
return (000, "ssl error")
|
||||
elif rc == 52:
|
||||
return (000, "empty response")
|
||||
else:
|
||||
return (000, "curl exit " + str(rc))
|
||||
return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response")
|
||||
except subprocess.TimeoutExpired:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
|
||||
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
||||
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
||||
cmd = "curl -s -m " + str(timeout)
|
||||
if method == "POST":
|
||||
cmd += " -X POST"
|
||||
if bearer_token:
|
||||
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
|
||||
if data:
|
||||
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
|
||||
cmd += " '" + url + "'"
|
||||
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
# Return first 200 chars, single line
|
||||
body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else ""
|
||||
return body
|
||||
|
||||
if rc != 0 and "TIMEOUT" not in stderr:
|
||||
return 000 # Connection failed
|
||||
|
||||
return int(stdout) if stdout.isdigit() else 000
|
||||
|
||||
def check_liveliness():
|
||||
"""Step 1: Liveliness probe"""
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
|
||||
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
|
||||
|
||||
def check_containers():
|
||||
@@ -116,61 +79,32 @@ def check_model_probes():
|
||||
results = []
|
||||
|
||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||
# Single-host aliases: 30s timeout each
|
||||
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
# Single-host aliases: 30s timeout
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Report probe failure with kind, do not assert a service verdict
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
|
||||
elif code == 200:
|
||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
elif code in (401, 403):
|
||||
# Credential fault - capture body and key alias
|
||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||
timeout=10)
|
||||
# Resolve key alias
|
||||
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
|
||||
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||
else:
|
||||
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
|
||||
# Pool alias (syslog-auto): 60s timeout, retry once on 000
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
if code == 000:
|
||||
# Retry once with same timeout
|
||||
time.sleep(1)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
if code == 000 and failure_kind:
|
||||
results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)"))
|
||||
elif code in (401, 403):
|
||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||
timeout=10)
|
||||
alias = "monitor-20260813"
|
||||
results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||
else:
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
else:
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
|
||||
return results
|
||||
|
||||
@@ -197,33 +131,33 @@ def check_admin_key_list():
|
||||
# Try to parse the response
|
||||
try:
|
||||
data = json.loads(stdout)
|
||||
# Response is a dict with "keys" (paginated list) and "total_count" fields
|
||||
# Response is a dict with "keys" field
|
||||
if isinstance(data, dict) and "keys" in data:
|
||||
key_count = data.get("total_count", len(data["keys"]))
|
||||
key_count = len(data["keys"])
|
||||
elif isinstance(data, list):
|
||||
key_count = len(data)
|
||||
else:
|
||||
key_count = 0
|
||||
if key_count == 0:
|
||||
return "Admin Key List", False, "admin-call-failed (empty response)"
|
||||
return "Admin Key List", True, str(key_count) + " total (" + str(len(data.get("keys", []) if isinstance(data, dict) else data)) + " on page 1)" if isinstance(data, dict) else str(key_count) + " total"
|
||||
return "Admin Key List", True, str(key_count) + " keys"
|
||||
except Exception as e:
|
||||
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
|
||||
|
||||
def check_github_status():
|
||||
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
|
||||
code, _ = probe_http("https://status.github.com/api/status.json", timeout=15)
|
||||
code = probe_http("https://status.github.com/api/status.json", timeout=15)
|
||||
# GitHub status API returns 301 redirect, which is expected behavior
|
||||
return "GitHub Status", code == 301, str(code)
|
||||
|
||||
def check_prometheus():
|
||||
"""Step 4: Prometheus health"""
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
|
||||
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
|
||||
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
|
||||
|
||||
def check_grafana():
|
||||
"""Step 9: Grafana health"""
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
|
||||
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
|
||||
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
|
||||
|
||||
def check_docker_stats():
|
||||
|
||||
@@ -7,11 +7,9 @@
|
||||
# agent leg is retired — see the note after the Tanko leg.
|
||||
set -euo pipefail
|
||||
|
||||
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
|
||||
# Never fall back to a literal key
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:?ZULIP_API_KEY not set — refusing to run with no credential}"
|
||||
ZULIP_SITE="https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
OWNER_ZULIP_ID="9"
|
||||
|
||||
|
||||
@@ -29,12 +27,12 @@ notify() {
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||
> /dev/null 2>&1 \
|
||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
|
||||
@@ -43,7 +41,7 @@ notify() {
|
||||
# ── Global: Zulip Server ──
|
||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||
https://chat.sysloggh.net/api/v1/server_settings \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
|
||||
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null) || SERVER_CODE="000"
|
||||
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
|
||||
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
|
||||
if [ "$SERVER_CODE" != "200" ]; then
|
||||
|
||||
@@ -27,7 +27,7 @@ Agent (Hermes/pi) ──curl + X-API-Key──► Stirling-PDF (:8989) ──►
|
||||
| Base URL | `http://192.168.68.7:8989` |
|
||||
| Auth Method | API Key (header) |
|
||||
| Header Name | `X-API-Key` |
|
||||
| API Key | `«vault: infrastructure/production STIRLING_API_KEY»` |
|
||||
| API Key | `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88` |
|
||||
| Key Source | `SECURITY_CUSTOMGLOBALAPIKEY` in `/opt/home_stack/docker-compose.yml` |
|
||||
| Swagger | `http://192.168.68.7:8989/swagger-ui.html` |
|
||||
| Health | `http://192.168.68.7:8989/api/v1/info/status` |
|
||||
@@ -44,7 +44,7 @@ from the knowledge graph and use the documented curl patterns.
|
||||
Direct bash invocations:
|
||||
```bash
|
||||
curl -X POST "http://192.168.68.7:8989/api/v1/split-pdf" \
|
||||
-H "X-API-Key: «vault: infrastructure/production STIRLING_API_KEY»" \
|
||||
-H "X-API-Key: adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88" \
|
||||
-F "fileInput=@/path/to/file.pdf" \
|
||||
-F "pageNumbers=1,2,3" \
|
||||
-o /tmp/output.zip
|
||||
|
||||
Reference in New Issue
Block a user