Compare commits

..
Author SHA1 Message Date
root 937afc0baf Ship: monitoring-contract-fixes-20260809
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Apply monitoring contract updates to match verified as-built reality:

- litellm-health.prose.md: Updated PVE node count from 5 to 6 (.4/.5/.6/.9/.12/.15:9100)
- hermes-key-enforcement.prose.md: Updated key expiry policy (explicit 90d duration required) and audit-only status
- litellm-api-keys.prose.md: Updated master key retrieval path (dynamic, not hardcoded)

Verified facts:
- LiteLLM /metrics mounts ONLY with success_callback: [prometheus]
- All 3 GPU exporters deployed and verified 200 on :9400
- Alertmanager + Zulip bridge delivering alerts to #agent-hub > alerts-infra
- Prometheus node job covers all 6 PVE nodes
- Agents set model.context_length: 131072 to prevent silent fallback to 256K

Note: The original patch does not apply cleanly because most changes were already applied in subsequent work. Only the PVE node count needed updating.
2026-09-15 03:58:01 +00:00
root 288d613876 fix(pm2-zulip): apply AS-BUILT note to pm2-self-heal description
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Apply the key change from the PM2/Zulip contract fixes patch (2026-08-09):
- Add the AS-BUILT 2026-08-09 note to the pm2-self-heal.prose.md description
  documenting the captain's ruling that gpu-monitor is systemd-managed,
  gpu-watchdog is decommissioned, and gitea-runner + abiba-zulip are kept.

The zulip-health.prose.md parameterized restart table from the patch is
not applied because the file has been restructured since the patch was
created: Mumuni is no longer monitored from this host (captain ruling
2026-09-10), so the per-agent restart command table is no longer needed.

The Maintains section of pm2-self-heal.prose.md already reflects the
correct process list (abiba-zulip, abiba-telegram, gitea-runner,
zulip-watchdog) from a previous PR.
2026-09-15 02:40:30 +00:00
83 changed files with 2357 additions and 6551 deletions
+10 -21
View File
@@ -1,18 +1,19 @@
name: PR Pipeline — Authorize → Validate → Review → Merge
# TRIGGER IS INTENTIONALLY UNFILTERED — DO NOT RE-ADD A `paths:` FILTER.
#
# This workflow previously carried `paths: ['**.prose.md', 'scripts/**.sh',
# '**.yaml', '**.yml']` on both `push` and `pull_request`. Any PR whose diff
# touched none of those patterns (for example a `deliverables/`-only PR, or a
# `scripts/*.py` / `bin/*` change) therefore produced NO Gitea Actions run at
# all: validation, lint, ai-review and the merge gate were silently skipped.
# Validation must run for every pull request and every push to master, so the
# trigger is deliberately unconditional.
on:
push:
branches: [master]
paths:
- '**.prose.md'
- 'scripts/**.sh'
- '**.yaml'
- '**.yml'
pull_request:
types: [opened, synchronize, reopened]
paths:
- '**.prose.md'
- 'scripts/**.sh'
- '**.yaml'
- '**.yml'
jobs:
auth:
@@ -71,18 +72,6 @@ jobs:
git fetch origin "${{ gitea.ref }}" --depth=50
git checkout "${{ gitea.sha }}"
- name: Committed-credential scan (secret guard)
run: |
# Fails the build on a credential-shaped string in the tree. Patterns
# live in scripts/secret-patterns.tsv; the only tolerated literal
# examples are in scripts/secret-allowlist.tsv, each with a reason.
# Do not turn this into a warning: a warning in a stream nobody reads
# is how six live credentials sat in this repo for weeks.
bash scripts/secret-scan.sh
- name: Secret guard self-test
run: bash tests/test_secret_scan.sh
- name: Structure + regression + consistency lint
run: bash scripts/prose-lint.sh
-1
View File
@@ -1,2 +1 @@
__pycache__/
state/host-disk-bands.json
-6
View File
@@ -51,12 +51,6 @@ Two incidents taught us this:
- `/grafana/` nginx route — was reverted Jul 2, must not reappear
- `CT 122` or `CT 123` as CT ID labels — don't exist in the cluster
- These rules are hardcoded in `scripts/prose-lint.sh`
- **Committed-credential guard:** `scripts/secret-scan.sh` FAILS the build on
credential-shaped strings (patterns in `scripts/secret-patterns.tsv`, prose
included). Tolerated literals are listed one-per-example with a reason in
`scripts/secret-allowlist.tsv`; never allowlist a live credential. It runs in
the CI lint job, in `scripts/prose-lint.sh`, and via
`bash scripts/secret-scan.sh --staged` before committing.
### Stage 3 — AI Review
- Diff is sent to `syslog-auto` model via LiteLLM
-142
View File
@@ -1,142 +0,0 @@
---
kind: responsibility
name: agent-health-check
description: >
Consolidated agent health verification for the LiteLLM + GPU + Zulip +
gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours
via cron and on-demand via "run contract: agent-health-check". Verifies:
LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway
liveness, CT liveness, gateway log health, config YAML integrity,
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
title: Agent Health Check — Consolidated
version: 1.0.0
runtime_contract: 2
agent: abiba
---
# Agent Health Check
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only.
**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
## Requires
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
- **SSH access** to GPU hosts — `llmuser` on .8 (owns `llama-server`), `root` on .110 and .15 — and agent CTs (.122, .129, .114, .24)
- **Python 3** for script execution
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
## Maintains
- last_check: timestamp — When the last full diagnostic ran
- overall_severity: "healthy" | "degraded" | "critical"
- liteLLM_keys: map of agent → key validity
- gpu_ports: map of host → port conflict status
- agents: map of agent → streaming health + gateway liveness
- ct_liveness: map of CT → active status
- config_integrity: map of config file → valid/invalid
## Execution
**Execution model**: This contract is executed by a host-scheduled cron job (see
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool
calls; never repeat a prior report unless a live probe fails.**
```bash
# Run the consolidated health check script
python3 /root/scripts/agent-health-check.py --json
```
**Report format**: Begin every report with the **absolute path the script
executed from** so a stale-consumer report is distinguishable from a real fault
at read time. Summarize actual results from each check. Apply the standing probe
rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
**Expected output**: JSON with `overall_severity` field. If `healthy`, report
"Agent health check: OK". If `degraded` or `critical`, report the specific
failures and their severity.
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
MUST include one clause per check leg, in every state: healthy, degraded/warn,
skipped, or failed. A missing leg must never look the same as a healthy leg.
Required legs and their templates in every state:
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
- The GPU leg has six non-healthy states the code can produce:
(i) `gpu-unreachable:{host}` — SSH probe failed;
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
(v) unit active, /health body contains "error" — error response;
(vi) unit active, /health body unrecognised — unknown health.
In every case the failing host and reason must be named.
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
- skipped: `CTs: SKIPPED (SSH access unavailable)`
- `Vault secrets: 3/3 present`
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
- skipped: `Vault secrets: SKIPPED (vault not configured)`
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
long as the host is identifiable from context; the full `probe-failed: <target>
<kind>` form is required when a leg reports a failure in the detail section.
### Probe Shape (per standing rules from 1150.msg)
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
service answered — report the code, never "down". A redirect is not a failure.
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
2. **A failed probe is never a service verdict.** Print
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
then report.
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
(retried at 25s: also timeout)" is actionable.
## Strategies
### When LiteLLM keys are invalid
Report the specific agent + key name. Do not attempt to fix — credential
rotation is a separate operation.
### When GPU port conflicts are detected
Report the conflicting ports and processes. Do not kill processes — that's a
destructive action requiring captain approval.
### When gateway liveness is degraded
Report the specific CT + gateway status. Do not restart unless the restart
debounce window has passed.
### When CT liveness is down
Report the specific CT. Do not restart — that's a destructive action.
### When config YAML is invalid
Report the specific file + parse error. Do not fix — that's a config change.
### When gateway log health is degraded
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
## Continuity
- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
- **On `agent-health` command**: Run on-demand and report to user
- **On critical alert**: Escalate to relay message immediately
+5 -5
View File
@@ -16,8 +16,8 @@ litellm.exceptions.AuthenticationError: OpenrouterException -
```
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
**Old Key**: `«vault: agents/production OPENROUTER_API_KEY»`
**New Key**: `«vault: agents/production OPENROUTER_API_KEY»`
**Old Key**: `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`
**New Key**: `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`
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
### 2. Telegram Bot Conflict (CRITICAL)
@@ -48,14 +48,14 @@ McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 sec
```bash
# Container .env update
sudo docker exec agent-zero bash -c '
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»|" /a0/usr/.env
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
'
```
**Verification**:
```bash
curl -s https://openrouter.ai/api/v1/auth/key \
-H "Authorization: Bearer «vault: agents/production OPENROUTER_API_KEY»" | python3 -m json.tool
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
```
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
@@ -140,7 +140,7 @@ Added section:
| Component | Status | Details |
|-----------|--------|---------|
| **OpenRouter Key** | ✅ Valid | `«vault: agents/production OPENROUTER_API_KEY»` user verified, |
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
| **MCP Services** | ✅ Working | No timeouts after key fix |
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
+4 -4
View File
@@ -54,7 +54,7 @@ description: >
```
4. **Return status**
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-synthetic...", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
- If vault secret is missing: `{ vault_synced: false }`
@@ -89,8 +89,8 @@ description: >
| Field | Value |
|-------|-------|
| **Key Prefix** | `«vault: agents/production OPENROUTER_API_KEY»` |
| **Full Key** | `«vault: agents/production OPENROUTER_API_KEY»` (in vault + /a0/usr/.env) |
| **Key Prefix** | `sk-or-v1-0af3f3` |
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
| **Free Tier** | No |
| **Monthly Usage** | 0 (as of 2026-09-01) |
@@ -101,7 +101,7 @@ description: >
| Date | Action | Notes |
|------|--------|-------|
| 2026-09-01 | fix-401 | Old key `«vault: agents/production OPENROUTER_API_KEY»…` returned 401 "User not found". Replaced with new key `«vault: agents/production OPENROUTER_API_KEY»…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
## Infrastructure References
+23 -90
View File
@@ -94,16 +94,7 @@ def audit(path):
cfg = yaml.safe_load(f)
model = cfg.get("model", {})
fb_raw = cfg.get("fallback_providers", {})
# Normalize: fallback_providers may be a dict (single provider) or a list of dicts
# (one entry per fallback). Both shapes are valid; we must handle both without crashing.
if isinstance(fb_raw, dict):
fb_entries = [fb_raw]
elif isinstance(fb_raw, list):
fb_entries = fb_raw
else:
fb_entries = [fb_raw] # Let it fail the check below as malformed
fb = fb_entries[0] if fb_entries else {}
fb = cfg.get("fallback_providers", {})
comp = cfg.get("compression", {})
aux = cfg.get("auxiliary", {})
deleg = cfg.get("delegation", {})
@@ -123,22 +114,12 @@ def audit(path):
)
# --- Rule 5: Main Config Base URL ---
# Canonical internal base (hermes-key-enforcement.prose.md:19/38/57/84/91/96)
# and public base (serves /v1 only, per 2026-09-19 probe from CT 116).
# The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only.
# FAIL anything else (do not widen to accept any path ending in /v1).
# Internal /v1 is non-canonical but working (authenticated via nginx), so WARN not FAIL.
canonical_internal = "http://192.168.68.116/litellm/v1"
public_host = "https://litellm.sysloggh.net/v1"
non_canonical_internal = "http://192.168.68.116/v1"
allowed_bases = (canonical_internal, public_host)
actual_base = model.get("base_url")
if actual_base in allowed_bases:
check(True, "Rule 5", f"model.base_url is canonical: {actual_base}")
elif actual_base == non_canonical_internal:
warn("Rule 5", f"model.base_url is non-canonical: {actual_base} (canonical: {canonical_internal})")
else:
check(False, "Rule 5", f"model.base_url must be one of {allowed_bases} (got {actual_base!r})")
expected_base = "http://192.168.68.116/v1"
check(
model.get("base_url") == expected_base,
"Rule 5",
f"model.base_url must be {expected_base} (got {model.get('base_url')!r}) — /v1 not /litellm/v1",
)
# --- Rule 6: max_tokens Is Required ---
check(
@@ -216,32 +197,22 @@ def audit(path):
"Rule 14",
f"delegation.provider must be 'harness' (got {deleg.get('provider')!r})",
)
# Check each fallback entry. A malformed entry (not a mapping) is a VIOLATION, not a crash.
for idx, entry in enumerate(fb_entries):
prefix = f"fallback_providers[{idx}]"
if not isinstance(entry, dict):
check(
False,
"Rule 14",
f"{prefix} must be a mapping (got {type(entry).__name__})",
)
continue
check(
entry.get("provider") == "deepseek",
"Rule 14",
f"{prefix}.provider must be 'deepseek' (got {entry.get('provider')!r}) — "
f"true fallback diversity, not same endpoint as primary",
)
check(
entry.get("model") == "deepseek-v4-flash",
"Rule 14",
f"{prefix}.model must be 'deepseek-v4-flash' (got {entry.get('model')!r})",
)
check(
entry.get("api_key_env") == "DEEPSEEK_API_KEY",
"Rule 14",
f"{prefix}.api_key_env must be DEEPSEEK_API_KEY (got {entry.get('api_key_env')!r})",
)
check(
fb.get("provider") == "deepseek",
"Rule 14",
f"fallback_providers.provider must be 'deepseek' (got {fb.get('provider')!r}) — "
f"true fallback diversity, not same endpoint as primary",
)
check(
fb.get("model") == "deepseek-v4-flash",
"Rule 14",
f"fallback_providers.model must be 'deepseek-v4-flash' (got {fb.get('model')!r})",
)
check(
fb.get("api_key_env") == "DEEPSEEK_API_KEY",
"Rule 14",
f"fallback_providers.api_key_env must be DEEPSEEK_API_KEY (got {fb.get('api_key_env')!r})",
)
# --- custom_providers sanity ---
check(
@@ -291,44 +262,6 @@ def audit(path):
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
)
# --- MCP Server Checks (Rule 15) ---
# Valid MCP server endpoints
VALID_MCP_ENDPOINTS = {
'ra-h-os': 'http://192.168.68.65:3100/mcp',
'litellm': 'https://litellm.sysloggh.net/mcp',
}
# Check MCP servers if they exist
mcp_servers = cfg.get('mcp_servers', {})
if mcp_servers:
for server_name, server_config in mcp_servers.items():
url = server_config.get('url', '')
# Check endpoint validity
if server_name in VALID_MCP_ENDPOINTS:
expected = VALID_MCP_ENDPOINTS[server_name]
if url == expected:
check(True, 'Rule 15', f'MCP server "{server_name}" URL is correct: {url}')
else:
check(False, 'Rule 15', f'MCP server "{server_name}" URL is incorrect: {url} (expected: {expected})')
else:
warn('Rule 15', f'MCP server "{server_name}" URL may need validation (not in known list): {url}')
# Check for proper authentication
headers = server_config.get('headers', {})
has_auth = False
for key, value in headers.items():
if 'key' in key.lower() or 'auth' in key.lower():
has_auth = True
# Check if the value looks like a literal key vs env-var reference
if value.startswith('Bearer ') and value[7:].startswith('sk-'):
check(True, 'Rule 15', f'MCP server "{server_name}" has valid auth header: {key}')
else:
warn('Rule 15', f'MCP server "{server_name}" header may use env-var instead of literal key: {key} = {value}')
break
if not has_auth:
warn('Rule 15', f'MCP server "{server_name}" has no authentication header')
# --- Report ---
print(f"{'=' * 60}")
print(f"Hermes Config Audit: {path}")
-173
View File
@@ -1,173 +0,0 @@
# Search ranking policy for the agent-consumption layer.
#
# Everything here is CONFIG, not code, so it is reviewable and changeable without
# touching the module. Read by scripts/search-agent-consume.py.
#
# Why this file exists: multi-engine aggregation returns results with no
# filtering, no dedupe and no reranking. On 2026-09-26 that put a shopping page
# and a dictionary definition into "best practices agent context management",
# and put four SEO blogs ABOVE the actual Proxmox forum threads on a precise
# technical query. Identical queries also ranked differently between runs, which
# is the strongest argument for a deterministic layer rather than hoping the
# engines behave.
version: 1
# ── Non-answers: dropped outright, never returned ────────────────────────────
# These are pages that cannot answer a question: navigational homepages,
# shopping/product pages, dictionary definitions, and login walls.
non_answer:
# URL path is empty -> it is a site's front door, not an answer. Still allowed
# when the host is explicitly preferred (see prefer_domains), because some
# docs/repo front doors ARE the answer.
host_root: true
path_patterns:
- '/dictionary/'
- '/dictionary?'
- '/wiki/Wiktionary:'
- '/search?'
- '/cart'
- '/checkout'
- '/login'
- '/signin'
- '/sign-in'
- '/account/login'
- '/shop/'
- '/store/'
- '/dp/' # Amazon-style product URL
- '/gp/product/'
- '/add-to-cart'
- '/checkout'
# NOTE: '/products/' and '/product/' were REMOVED as path patterns. They fired
# on docs.digitalocean.com/products/inference/... — a legitimate documentation
# page — which the 2026-09-26 before/after run caught. Shopping is caught by
# the shopping HOST list instead, which does not have that false positive.
# Query strings that betray a search/shopping surface rather than an article.
query_keys:
- 'q'
- 'query'
- 's'
- 'search'
- 'add-to-cart'
# Hosts that are shopping/retail and never answer a technical question.
hosts:
- bestbuy.com
- amazon.com
- ebay.com
- walmart.com
- etsy.com
- aliexpress.com
- merriam-webster.com
- dictionary.com
- thesaurus.com
- vocabulary.com
- collinsdictionary.com
# ── Demotion: ranked below everything else, never dropped ────────────────────
# Low-authority content farms / SEO aggregators. Demoted rather than dropped so
# a genuinely useful hit is not lost, but it can never outrank a primary source.
# Reviewable: add or remove hosts here, no code change required.
demote_domains:
- medium.com
- sparkco.ai
- mindstudio.ai
- aitechmonk.com
- stackai.com
- agentic-design.ai
- voxfor.com
- bigiron.cc
- linuxoperatingsystem.net
- riparazioneserver.com
- rossmanngroup.com
- dev.to
- hashnode.dev
- substack.com
- towardsdatascience.com
- analyticsvidhya.com
- geeksforgeeks.org
- tutorialspoint.com
- javatpoint.com
- w3schools.com
- scaler.com
- simplilearn.com
- udemy.com
- coursera.org
# ── Preference: promoted above the default rank ──────────────────────────────
# Primary sources: upstream repositories, official docs, Q&A, vendor
# engineering blogs. These are what an agent should be reading.
prefer_domains:
# upstream repositories and code hosting
- github.com
- gitlab.com
- codeberg.org
- sourceforge.net
- kernel.org
- git.kernel.org
# Q&A
- stackoverflow.com
- stackexchange.com
- superuser.com
- serverfault.com
- askubuntu.com
- discourse.org
# vendor / project documentation and forums
- proxmox.com
- forum.proxmox.com
- pve.proxmox.com
- docs.python.org
- developer.mozilla.org
- kernelnewbies.org
- man7.org
- gnu.org
- debian.org
- ubuntu.com
- redhat.com
- kernel.dk # io_uring / Jens Axboe
- github.io # project pages (docs, papers) — promoted, not authoritative by itself
# vendor engineering blogs
- anthropic.com
- openai.com
- googleblog.com
- developers.googleblog.com
- engineering.fb.com
- netflixtechblog.com
- aws.amazon.com
- cloud.google.com
- microsoft.com
- learn.microsoft.com
- apple.com
- nvidia.com
- intel.com
- amd.com
- redislabs.com
- cloudflare.com
- langchain.com
- jetbrains.com
- cursor.com
# community discussion with high signal
- news.ycombinator.com
- lobste.rs
- reddit.com
# ── Ranking weights ──────────────────────────────────────────────────────────
# Final score = engine_score - demote_penalty + prefer_bonus, then a stable
# tiebreak on original position so ordering is reproducible run to run.
ranking:
demote_penalty: 1000
prefer_bonus: 100
# Results that several engines independently returned are more likely real.
multi_engine_bonus: 25
# Shallow paths (e.g. /blog/x) are slightly less likely to be primary docs.
host_root_allowed_when_preferred: true
# ── Extraction budget (criterion 4) ──────────────────────────────────────────
# Return CONTENT, not just links, so an agent gets usable material in ONE call.
extraction:
top_n: 5 # how many results get page text extracted
total_chars: 12000 # global budget across all extracted items
per_item_chars: 4000 # cap for any single item, so one page cannot eat the budget
timeout_seconds: 45 # per scrape
# If extraction fails, the result is still returned with an empty excerpt —
# a link is better than nothing, but the failure is recorded in the output.
on_failure: keep_with_empty_excerpt
+2 -118
View File
@@ -38,7 +38,6 @@ index:
by_category:
compliance:
- hermes-key-enforcement
- litellm-api-keys
- hermes-config-template
- hermes-agent-baseline
monitoring:
@@ -47,7 +46,6 @@ index:
- infrastructure-monitoring
- zulip-health
- litellm-health
- daily-health-digest
remediation:
- litellm-self-heal
- pm2-self-heal
@@ -97,14 +95,12 @@ index:
- infrastructure-maintenance
- pm2-self-heal
- disk-gc-threat-response
- daily-health-digest
gpu:
- gpu-monitor
- gpu-fleet
proxmox:
- proxmox-monitor
litellm:
- litellm-api-keys
- litellm-health
- litellm-self-heal
memory:
@@ -632,7 +628,7 @@ contracts:
sensitivity: high
status: active
owner: abiba
version: 3.4.0
version: 3.3.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
@@ -643,7 +639,7 @@ contracts:
timeout: 120
requires:
- Zulip API key for abiba-bot@chat.sysloggh.net
- SSH access to minipve (192.168.68.12) for Tanko (CT 112) and the Agent Zero Docker host (.14)
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
verification:
postconditions:
- check: bot registration active
@@ -1873,118 +1869,6 @@ contracts:
drift_alerts: []
# Koby Report-Only Registry (2026-08-17 — Captain)
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
- name: litellm-api-keys
file: litellm-api-keys.prose.md
kind: function
category: compliance
sensitivity: critical
status: active
owner: abiba
version: 1.1.0
trigger:
type: on_demand
cadence: null
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
cron_job_id: null
execution:
agent: abiba
timeout: 120
requires: []
protocol:
- Load contract from prose-contracts/main
- Retrieve master key from Infisical (project=infrastructure env=production)
- Read live key-scoped model roster from CT 116 /v1/models
- Create/rotate/verify the requested agent key with an EXPLICIT models list
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
verification:
postconditions:
- check: standard agent key is local-only
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
expect: 0 cloud models
- check: key exists with correct alias
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
expect: 200 with matching alias
artifact: key creation/rotation report
receipt:
format: json
storage: ~/.hermes/runs/litellm-api-keys/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
- ops
- name: daily-health-digest
file: daily-health-digest.prose.md
kind: function
category: monitoring
sensitivity: normal
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: 30 10 * * *
description: Daily at 10:30 UTC, dispatched on CT 100 as a firstmate message
to the ops lane, which executes the pinned producer
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires:
- infisical (vault credentials injected at run time)
protocol:
- 'Execute from the PINNED clone only: /root/abiba-workspace/projects/prose-contracts'
- cd /root/abiba-workspace/projects/prose-contracts
- infisical run --env=prod -- python3 scripts/daily-infra-report.py
- 'Never execute from a per-agent working copy (treehouse) - it drifts onto feature branches'
- 'On failure: do not treat a 0/0 Proxmox section as evidence about the estate - it means could not look'
verification:
postconditions:
- check: every Proxmox probe is reachable
verify: >-
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
| grep -c 'pve_probe_status'
expect: '1'
- check: all nodes reported online
verify: >-
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
| grep 'nodes_online'
expect: nodes_online == node_count
artifact: timestamped HTML dashboard emailed to jerome@sysloggh.com
verify_commands:
- infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
- python3 -m pytest tests/test_daily_infra_report.py -q
delivery:
transport: zulip-dm-attachment
recipient_user_id: 9
sender: abiba-bot@chat.sysloggh.net
key_source: abiba-bot Zulip key already on the execution host, read from the
600-mode env file /root/.pi/agent/extensions/zulip/.env
key_policy: do NOT add a vault entry - that is a captain decision under the auth-keys charter
body: short Markdown pointer; the HTML attachment IS the report
artifact: /var/log/daily-infra-report/infra-report-<UTCstamp>.html
note: Replaced SMTP/mail on 2026-09-26 by captain decision. Removes the Google
dependency entirely; closes daily-digest-mail-transport-20260921.
exit_semantics:
'1': missing PVE_TOKEN, unreachable Proxmox probe, missing/rejected Zulip
credential, or a failed upload/post - raises an alert
'0': healthy delivery only - there is no degraded delivery leg any more
depends_on: []
last_run: null
last_status: null
drift_alerts: []
koby_report_only: true
koby_host: "CT 111 (tdunna)"
koby_ip: ".129"
-203
View File
@@ -1,203 +0,0 @@
---
kind: function
name: daily-health-digest
description: >
Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox
nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails
it as an HTML report.
Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to
the ops lane, which executes the producer below. Until 2026-09-25 this ran
with NO contract file at all, which is why the choice of execution copy was
silently the operator's rather than the contract's.
EXECUTION IS PINNED. The producer must be run from the clone named under
"Execution pinning" — not from an agent working copy.
Exit-code semantics (as they actually behave, verified 2026-09-25):
* missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert
* missing or rejected Zulip credential -> exit 1 (delivery is the only
output path, so it is a real failure, not a degraded leg)
* delivery failure -> exit 1, and the report body is printed AND persisted
so the content is never swallowed
version: 2.0.0
---
## Purpose
Give one daily, machine-collected picture of the estate so drift and outages
are seen the day they happen rather than when something breaks. It is a
*report*, not a repair: it changes nothing.
## Execution pinning
**Pinned execution path:**
```
/root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py
```
**Pinned clone:** `/root/abiba-workspace/projects/prose-contracts`
That is the cron's `FM_HOME` clone and the only stable, non-ephemeral copy.
The treehouse clone (`/root/.treehouse/agent-workspace-*/…/projects/prose-contracts`)
is a **per-agent working copy and must NOT be pinned or executed from** — it
drifts onto feature branches, which is exactly how a stale producer reported a
stale picture and nobody noticed.
See `docs/contract-execution-pinning.md`. Schedule and alerting live in
`/etc/cron.d/contract-runner` on CT 100:
```
30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1
```
Invocation (credentials come from the vault; never inline them):
```bash
cd /root/abiba-workspace/projects/prose-contracts
infisical run --env=prod -- python3 scripts/daily-infra-report.py
```
## Output shape
Modes:
| invocation | effect |
| --- | --- |
| *(none)* | collect, build the HTML dashboard, email it |
| `--test-email` | same but with a `🧪 TEST —` subject prefix |
| `--json` | print the collected data as JSON to stdout and **send no email** |
`--json` emits a single object with these top-level keys (observed on a live
run 2026-09-25):
| key | type | meaning |
| --- | --- | --- |
| `nodes` | object (5) | per-node cpu/ram/disk/uptime/status |
| `node_count`, `nodes_online` | int | Proxmox node totals |
| `pve_probe_status`, `resources_probe_status` | `ok`\|`unreachable` | probe outcome |
| `total_vms`, `running_vms`, `stopped_vms`, `vms_by_node` | — | guest inventory |
| `storage`, `nfs` | array | datastore and mount usage |
| `litellm` | object | inference checks |
| `zulip_ext` | object | Zulip queue/serving state |
| `agents` | object | per-agent health |
| `docker_vm`, `docker_syslog`, `docker_netbird`, `endpoints` | object/array | Docker hosts and probed endpoints |
## What a healthy run looks like
```
$ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
"nodes_online": 5, "node_count": 5, "pve_probe_status": "ok",
"resources_probe_status": "ok", "running_vms": 22, "total_vms": 22
EXIT=0
```
and in delivery mode:
```
report ready: 16208 chars of HTML (delivered as a file attachment)
Sending to the captain's Zulip DM...
✅ Delivered to Zulip DM (user 9), message id 86221, attachment 16208 bytes
at /user_uploads/2/45/m1cQesBFV78BGeNY2lN8xkN5/infra-report-20260926-153406.html
```
Healthy means: every probe reports `ok`, `nodes_online == node_count`, and the
email leg reports a successful send.
## Exit-code semantics — as they actually behave
Verified on 2026-09-25 by running each case deliberately.
| condition | exit | alert | notes |
| --- | --- | --- | --- |
| all probes reachable, email sent | 0 | — | healthy |
| **missing `PVE_TOKEN`** | **1** | yes | `PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down)`, and `cluster resources unreachable` |
| **Proxmox probe unreachable** | **1** | yes | same path as above; `pve_probe_status: unreachable` |
| **missing/rejected Zulip credential** | **1** | yes | delivery is the only output path; report printed and persisted |
| **upload or message post fails** | **1** | yes | report printed and persisted; message names which step failed |
The distinction is deliberate and must not be flattened:
* A **missing PVE token or an unreachable probe is a real failure** — the report
would otherwise claim zero nodes and still look successful. That was the
2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1.
* A **missing email credential is survivable** — the report is still produced
and is still useful. It is a `DEGRADED` leg and exits 0 by design.
`PROBE_FAILURES` and `DEGRADED_LEGS` are separate lists for exactly this
reason. Do not merge them.
## Delivery: Zulip DM carrying the report as an HTML ATTACHMENT
Captain's decision 2026-09-26, clarified the same day: the digest is delivered to
his **Zulip DM (user id 9)** from `abiba-bot@chat.sysloggh.net`, as an **HTML
FILE** — an attachment, not HTML rendered in the message body and not a Markdown
translation of it.
* the styled dashboard is built exactly as before and written to
`/var/log/daily-infra-report/infra-report-<UTCstamp>.html`;
* it is uploaded through `POST /api/v1/user_uploads`;
* the **message body stays short Markdown** — subject line, top-line status
(nodes online, guests running, any degraded legs), and a link to the
attachment. The attachment IS the report; the body does not reproduce it.
This removes the Google dependency entirely: **no SMTP, no `EMAIL_PASSWORD`, no
app password, nothing to rotate.** `daily-digest-mail-transport-20260921` is
closed under this option.
The **10,000-character message cap does not apply** — it bounds message TEXT
only, and the report travels as a file. Do not shrink the report to fit it.
The credential is abiba-bot's Zulip key already on the execution host at
`/root/.pi/agent/extensions/zulip/.env` (`ABIBA_ZULIP_API_KEY`, mode 600,
root-readable). **Do not place a new credential in the vault** — under the
auth-keys charter that is a captain decision.
## What counts as a failure
A run FAILS (exit 1) when the report cannot be trusted or delivered:
* any probe is unreachable, so a section would silently be empty;
* `PVE_TOKEN` is missing;
* the Zulip credential is missing or rejected, or the upload/post fails.
There is **no degraded delivery leg any more**. Delivery is the only output
path, so a missing credential is a failure rather than a survivable degradation —
the previous "missing `EMAIL_PASSWORD` still exits 0" rule is retired with the
mail transport.
**A delivery failure must never swallow the report.** On failure the script
prints the report body to stdout *and* leaves the HTML artifact on disk, so the
content is always recoverable from the run log. That closes the queued defect
where a failed send printed only the transport error and the report never
surfaced.
## Failure behaviour
* Non-zero exit with the alert text above; on the scheduled path the dispatch is
a firstmate message, so the ops lane sees it and reports it.
* On a probe failure the report must **not** be treated as evidence about the
estate — a `0/0` Proxmox section means "could not look", not "nothing there".
That reading is why the 2026-09-25 defect went unnoticed.
## Verification
```bash
# data path, no email
cd /root/abiba-workspace/projects/prose-contracts
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \
| grep -E 'pve_probe_status|node_count|nodes_online'
# delivery path
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-zulip
```
Regression tests: `tests/test_daily_infra_report.py` (7 tests). Four of them
fail against the pre-fix script, which is what makes them bite.
## Maintains
- daily-infra-dashboard: { status: "ok|undelivered", transport: zulip-dm-attachment, last_check: timestamp }
- pve-probe: { status: "ok|unreachable", last_check: timestamp }
@@ -0,0 +1,75 @@
# Delivery Record — HERMES-PLAYBOOK-FOR-SCOT
## Status: SEND-READY — awaiting Kwame's channel + recipient confirmation
No documented channel to Scot Murray exists in this workspace, the skills, or config
(verified 2026-09-11 by sweep of `~/syslog/projects/murray-capital/`, `syslog-infra`
references, `murray-harness` skill, `.hermes/memories/`, all of `~/syslog/`).
Per the card's unblock constraints: package prepared, exact send commands written
below, nothing transmitted. Guessing an address is out of scope.
## Verified artifact (single source of truth)
| Item | Value |
|---|---|
| Markdown source | `/home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.md` |
| sha256 | `b0966a649fec96da4975ec00e627fbaba3a92a62c4a92b33bc06589c47d25f7f` |
| Size | 28821 bytes, 377 lines |
| Matches reviewed bytes | YES — identical to `/home/hermes/syslog/drafts/scot-hermes-playbook/` copy and to the hash recorded on card t_2c716052 |
## Rendered PDF (from the verified bytes, no edits)
| Item | Value |
|---|---|
| PDF | `/home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf` |
| sha256 | `f46c0c89b4bbf6fa76c1f1c385c87753860d06bfa9e8925d6e09f27ed78a187b` |
| Size | 96,486 bytes · 13 pages · A4 |
| Render chain | pandoc 3.1.11.1 (gfm → html5) + WeasyPrint 62.3, stylesheet `pb.css`; reproducible via `bash render_pdf.sh` |
| Spot-check | pdftotext shows correct title page + v0.21.1 verification note |
## Candidate channels — exact commands (pending Kwame's pick + address)
### 1. Email via syslog-email profile (recommended)
Mailbox ops belong to the syslog-email profile per standing rule. Send as
jerome@sysloggh.com with both attachments.
```
hermes -p syslog-email chat -q "Send an email. From jerome@sysloggh.com. \
To: <SCOT-ADDRESS — Kwame to supply>. Subject: 'Hermes Playbook — getting real mileage out of the harness'. \
Attach: /home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf \
and HERMES-PLAYBOOK-FOR-SCOT.md. Body: short intro noting the PDF is the reviewed v0.21.1 playbook, \
sha256 b0966a64… (sic, abbreviated), ask him to flag anything confusing — that feedback feeds the harness build. \
Show me the draft before sending."
```
Direct himalaya (only if Kwame wants it from the main session — normally NOT, per
the email-routing standing rule):
```
himalaya message write --to "<SCOT-ADDRESS>" --subject "Hermes Playbook — getting real mileage out of the harness" \
--attachment .../HERMES-PLAYBOOK-FOR-SCOT.pdf --attachment .../HERMES-PLAYBOOK-FOR-SCOT.md
himalaya message send <draft.eml>
```
### 2. Telegram (only if Kwame has Scot's handle)
Send the PDF to Scot's handle from the gateway-connected Telegram session:
```
hermes chat -q "Send the file /home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf to <SCOT-HANDLE> with a one-line intro."
```
### 3. Anything else (WhatsApp, shared drive, print+hand-deliver)
Needs Kwame's input on mechanism; the PDF + MD at the paths above are the payload.
## Post-send obligations (from the card)
1. Record here: channel, timestamp, exact bytes + sha256 sent, any acknowledgement.
2. Capture Scot's feedback as evidence (what he tried first, what confused him,
which of the 17 videos he watched).
3. Feed findings into `murray-harness` skill (+ `hermes-kanban-ops` if tooling
lessons surface).
4. If no reply in 7 days: ONE follow-up nudge is in scope; more is Kwame's call.
5. Feature gaps he reports → separate card, do not widen this one.
@@ -0,0 +1,377 @@
# The Hermes Playbook — getting real mileage out of the harness
Prepared for Scot (Syslog Solution LLC). Version: Hermes Agent v0.21.1. Every CLI command below was verified live against that version on a reference install (Syslog kagentz) on 2026-09-11; anything only confirmed against the official docs is tagged DOC-ONLY.
You are already running Hermes next to Claude Code, on your own OpenRouter account with fast models (qwen3.8-flash, deepseek-4-flash). The question you asked: why does Hermes feel like it has less context, and what do I do about it?
---
## 1. TL;DR
- The context gap is not a bug. Claude Code reads the repo it sits in on every launch; a fresh Hermes install starts nearly empty by design. It gets its context from files you seed and from memory it builds over time.
- One command closes most of the gap on day one: `hermes import-agent claude-code` carries your CLAUDE.md instructions, MCP servers, skills, and memories into Hermes (preview first with `--dry-run`).
- Teach Hermes once, and it remembers: "save this as a skill" after any workflow you repeat. Skills auto-load when a matching task comes up — that is the learning loop.
- Keep per-project context in an `AGENTS.md` in the repo root (git-tracked, shared with your team) and personal preferences in your persona file and persistent memory.
- Hermes and Claude Code are not rivals: let Hermes be the always-on orchestrator (research, briefs, scheduling, messaging) and hand heavy coding to Claude Code, which Hermes can drive directly.
---
## 2. Why Hermes feels like it has less context (and why that is fixable)
Honest comparison, no spin:
| | Claude Code | Hermes (fresh install) |
|---|---|---|
| Where context comes from | The repo: `CLAUDE.md` auto-loaded every launch; `.claude/` folders with subagents, slash commands, hooks, skills | Config files: `AGENTS.md` in the working directory + `SOUL.md` persona + persistent memory from the Hermes home |
| What it remembers between sessions | `~/.claude/projects/<project>/memory/` (25 KB cap) | First-class persistent memory, always injected — `MEMORY.md` / `USER.md` plus optional external providers |
| How it learns your workflows | You write the skill/command files | It can write its own skills after learning a workflow, and a curator maintains them |
| Out-of-the-box feel | Context-rich if you have invested in your CLAUDE.md | Quiet until you seed it |
That last line is the whole story. Claude Code's context is the sum of everything you built in `CLAUDE.md` and `.claude/` over months. A fresh Hermes has none of that yet — not because the harness is weaker, but because it stores context in different places and expects you to seed it (or let it build up).
The gap is fixable in two moves:
1. **Import what you already have.** `hermes import-agent claude-code` maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, and memories into Hermes equivalents. Preview with `--dry-run`; it never imports API keys; conflicts are skipped by default (`--overwrite` to change).
2. **Let the learning loop run.** Every time you correct Hermes or finish a workflow you will repeat, tell it to remember. Within a few weeks it will have its own CLAUDE.md equivalent — built, not typed.
What the comparison table in our research covers, gap by gap: project instructions, instruction splits, slash commands, subagents, skills, project memory, tool permissions, MCP, session resume, cost/context visibility, headless mode, prior-setup import, hooks, and scheduled work. Each has a Hermes equivalent, and every one is documented in section 9.
---
## 3. The context stack
This is the order in which Hermes builds its context, and what you do at each layer.
**Layer 1 — Persona (`SOUL.md`).** Set up once. Your Hermes' standing identity and voice: "you are my analyst," the tone, the standing rules. Auto-injected into every session. Lives at `~/.hermes/SOUL.md` (per profile: `~/.hermes/profiles/<name>/SOUL.md`). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/configuration
**Layer 2 — Persistent memory.** Set up once, then feed it constantly. `MEMORY.md` / `USER.md` are always active and injected every session — this is the single biggest cure for "it forgets my project." After any correction or preference ("use this source list," "briefs go in this format"), tell Hermes to remember it. Manage with `hermes memory setup|status|off|reset` (VERIFIED-LIVE). Optional external providers exist (Honcho, Mem0, and others). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/features/memory
**Layer 3 — Skills.** Set up once; grows forever. Markdown procedure files that auto-load when a task matches the skill. The differentiator: after completing a workflow, ask Hermes to "save this as a skill" — it authors the skill itself, and a background curator tracks usage, archives stale ones, and keeps backups. CLI: `hermes skills list|search|install|browse|config|check|update` (VERIFIED-LIVE); in-session: `/skill <name>`, `/reload-skills` (DOC-ONLY). Docs: https://hermes-agent.nousresearch.com/docs/reference/skills-catalog and https://hermes-agent.nousresearch.com/docs/user-guide/features/curator
**Layer 4 — Projects.** Per workstream. `AGENTS.md` in each repo root (git-tracked, team-shared) carries project rules; Desktop Projects (`hermes project create <name>` then `add-folder`) group multi-repo work under one named workspace. Both VERIFIED-LIVE.
**Layer 5 — Retrieval (session store).** Automatic. All conversations land in a searchable store; Hermes can search past sessions when you ask "what did we decide last week." CLI: `hermes sessions list|browse|rename|pin|export|prune|stats` (VERIFIED-LIVE).
**Layer 6 — MCP (external tools).** Per integration. Plug GitHub, databases, workflow engines into the agent. `hermes mcp add|list|test|configure|picker|catalog|install` (VERIFIED-LIVE). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp
Quick summary:
| Layer | Set up | Feed |
|---|---|---|
| SOUL.md persona | once | rarely |
| Persistent memory | once | every correction/preference |
| Skills | once | "save this as a skill" after repeated workflows |
| AGENTS.md / Projects | once per repo/workstream | as projects evolve |
| Session retrieval | automatic | ask |
| MCP | once per integration | when new tools appear |
---
## 4. Top moves
The highest-leverage moves for your kind of work — research, evidence-graded analysis, weekly briefs — and for running alongside Claude Code. Every command verified on v0.21.1.
### 1. Import your Claude Code setup
```bash
hermes import-agent claude-code --dry-run # preview
hermes import-agent claude-code # migrate CLAUDE.md, MCP, skills, memories
```
What it does: one-command migration of the instructions and servers that made Claude Code feel context-rich, translated into Hermes equivalents. Never imports API keys.
Why it matters: this is the direct answer to "Hermes has no context." After this, Hermes knows your projects on day one.
### 2. Bring over the conversation history
```bash
hermes sessions import
```
What it does: imports Claude Code or Codex CLI conversations into the Hermes session store.
Why it matters: mid-project, the new agent picks up exactly where the old one left off. `hermes --resume <id>` and `hermes sessions browse` then treat the history as native.
### 3. Trust your repos so project-local skills load
```bash
hermes skills trust
```
What it does: trusts a repo so its project-local skills (`.hermes/skills`) load — the Hermes analog of `.claude/skills/`.
Why it matters: your evidence-grading rules can live in the repo with the project, versioned with git, and load automatically.
### 4. Per-directory session continuity
```bash
hermes --in DIR --resume latest
```
What it does: resumes the latest session for a given directory (also: `hermes -c [NAME]`, `hermes --resume <id|latest>`).
Why it matters: every project folder gets its own continuous thread. Research on one portfolio never mixes with another.
### 5. Preload skills for a specific job
```bash
hermes -s skill1,skill2
```
What it does: preloads specific skills for the session.
Why it matters: for a weekly brief or an evidence register, pin the exact skills that encode your grading criteria instead of hoping they auto-match.
### 6. Save any repeated workflow as a skill
In-session: "save this as a skill." (CLI: `hermes skills list|search|install|browse`.)
What it does: Hermes authors a skill file from the workflow you just ran.
Why it matters: the learning loop is the whole point. Do the evidence-grading pass twice, save it, and every future run starts with the procedure loaded.
### 7. Fan out research with subagents
In-session: "delegate this to subagents." (agent-side tool `delegate_task`, no CLI.)
What it does: parallel subagents with isolated contexts — each gets its own conversation and terminal, only the final summary comes back.
Why it matters: research fan-out without flooding your main context. Ten sources, ten subagents, one synthesis.
### 8. Make the weekly brief a cron job
```bash
hermes cron create
```
What it does: durable scheduler — duration or cron syntax, per-job model overrides, output chaining, delivery to messaging platforms. Manage with `hermes cron list|create|edit|pause|resume|run|remove|doctor`.
Why it matters: a weekly brief is exactly a cron job. It runs even when you are not at the desktop, with your skills preloaded and its output delivered to you.
### 9. Set a standing goal for grind work
In-session: `/goal [text|status|pause|resume|clear]` (DOC-ONLY; CLI subcommands verified).
What it does: a standing objective the agent keeps working toward across turns until achieved.
Why it matters: "keep researching until you have 5 verified sources" — the agent loops itself instead of waiting for you to say "go on."
### 10. Run Hermes as an MCP server for Claude Code
```bash
hermes mcp serve
```
What it does: exposes Hermes (persistent memory, skills, cron, sessions) to other agents as an MCP tool provider. Claude Code supports MCP clients, so it can consume Hermes.
Why it matters: the reverse bridge. Claude Code gets the surfaces it lacks, and both tools share your knowledge base.
### 11. Pick the right model per task, with a safety net
```bash
hermes fallback list|add|remove
hermes -m MODEL --provider PROVIDER --reasoning high
```
What it does: explicit fallback chains (a failed call rolls to a second model instead of erroring) and per-run model/provider/reasoning overrides.
Why it matters: on OpenRouter with fast models, use `--reasoning high` for the hard analytical passes and let fallback chains keep the cheap models from stalling your brief.
### 12. Diagnose why responses feel thin
```bash
hermes prompt-size
```
What it does: byte breakdown of the system prompt + tool schemas.
Why it matters: when output quality drops, it is usually context bloat, not model quality. This tells you what is eating the window.
---
## 5. Working alongside Claude Code
You run both. The proven patterns, in order of value.
**First: import.** `hermes import-agent claude-code` then `hermes sessions import`. After this, the "two tools that don't know each other" problem is gone — Hermes knows your projects and your history.
**Hermes as orchestrator, Claude Code as worker.** The installed Hermes skill for exactly this is `autonomous-ai-agents/delegate-coding-agent`. Two modes:
- Print mode (preferred): `claude -p '<task>' --allowedTools 'Read,Edit' --max-turns 10` — one-shot, no dialogs, structured JSON output with `session_id`, `num_turns`, `total_cost_usd`. In Hermes, just say: "delegate this coding task to Claude Code in print mode."
- Interactive PTY via tmux: Hermes starts a tmux session, sends prompts with `send-keys`, monitors with `capture-pane`. For iterative refactor → review → fix cycles.
There is also a cross-agent review loop: `git diff main...feature | claude -p 'Review this diff for bugs and security issues.' --max-turns 1` — Hermes runs it, reads the findings, and fixes them itself. Claude Code becomes a reviewer Hermes coordinates.
The skill's safety rails: explicit workdir, clean git status before launch, narrow task prompts, git diff review, targeted tests before committing.
**Parallel workstreams, neutral merge reconciliation.** When both agents edit the same repo and collide, do not let either resolve the conflict — both are biased toward their own side. Spawn a neutral third agent with the `merge-reconciler` skill: it classifies every conflicted hunk, resolves under an impartiality contract, verifies with build/tests, and hands back a summary naming every decision. Kanban shape: a reconciliation card assigned to a third profile, with both workers' cards as parents.
**Hermes as MCP server (the reverse direction).** `hermes mcp serve` (pattern in section 4, move 10). The only bridge direction Claude Code cannot offer.
**Desktop GUI goes to Hermes.** Claude Code has no desktop automation. `hermes computer-use install` (cua-driver; health check `hermes computer-use doctor`) drives native desktop apps background-first — never steals focus. If a task needs Excel or a native app, that part routes to Hermes while the code routes to Claude Code.
**Division of labor in one line:** Hermes is the always-on layer — research, briefs, scheduling, messaging, memory, and the orchestration desk. Claude Code is the deep coding worker. Hand coding-heavy tasks over; hand continuity, recall, and scheduled work to Hermes.
---
## 6. Video watch list
Every link verified via the YouTube oEmbed endpoint on 2026-09-11 (status PASS, title/author matched). All content is third-party ecosystem material — no official Nous Research tutorial video exists (see section 8).
| # | Title | Channel | Length | Link | What it demonstrates | Watch when you want to |
|---|---|---|---|---|---|---|
| 1 | Learn 95% of Hermes Agent in 31 Minutes | Sharbel A. | 31:28 | https://www.youtube.com/watch?v=Ta2wg6xPaY4 | End-to-end fundamentals: install, sessions, skills, memory, the learning loop | the fastest real overview of the whole harness before touching config |
| 2 | Hermes Agent Fundamentals In 29 Minutes | Tina Huang | 29:40 | https://www.youtube.com/watch?v=5_N84t1rUU0 | Why Hermes' memory/skills loop differs from one-shot coding agents | understand why Hermes feels different from Claude Code |
| 3 | Every Level of Hermes Agent Explained | Jack Roberts | 25:35 | https://www.youtube.com/watch?v=6GtF_uHbGhw | Beginner to advanced ladder: memory, skills, automation, multi-agent | a map of what to learn next after the basics |
| 4 | Hermes Agent Full Tutorial INSTALLATION + USECASES | CodeHead | 7:47 | https://www.youtube.com/watch?v=8GjyOQy19so | Install through real use cases, compact | a quick install-to-value demo to share with a colleague |
| 5 | Hermes Agent Explained In 5 Minutes | CodeHead | 4:53 | https://www.youtube.com/watch?v=9GpWELm3_XI | Conceptual pitch of the agent and its learning loop | the elevator pitch before committing 30 minutes |
| 6 | 100 Days With Hermes Agent in 21 Minutes | Sharbel A. | 21:19 | https://www.youtube.com/watch?v=sCa3BtpkziQ | What memory/skills accumulation looks like after months of daily use | see the payoff of the learning loop over time |
| 7 | Hermes Agent - Crash Course for Beginners (AI Agent) | Adrian Twarog | 22:19 | https://www.youtube.com/watch?v=4sAmpcSOVEw | Beginner crash course from a well-known dev channel | a second independent explanation of the basics |
| 8 | Hermes Agent: The Ultimate Beginner's Guide | Metics Media | 37:08 | https://www.youtube.com/watch?v=CwPUOVUdApE | Long-form beginner guide incl. setup and everyday workflows | the most thorough single walkthrough in one sitting |
| 9 | Hermes Agent Just Killed OpenClaw (Full Tutorial) | Leon van Zyl | 19:59 | https://www.youtube.com/watch?v=jmtpYUOr7_U | Feature-by-feature tutorial (MCP config, memory, agents) | a practitioner's feature-by-feature tutorial |
| 10 | Hermes Agent vs OpenClaw | Sharbel A. | 15:28 | https://www.youtube.com/watch?v=zwqhemjHq3E | Head-to-head comparison of the two agent harnesses | the tradeoffs between Hermes and its main alternative |
| 11 | Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model | Tonbi's AI Garage | 15:08 | https://www.youtube.com/watch?v=8tpuky8HpXw | Hermes driven by an OpenRouter-served open model | how small open models behave inside Hermes |
| 12 | Use This To Make The Hermes Agent Basically Free | AI LABS | 13:08 | https://www.youtube.com/watch?v=5d02TYoOzfE | Running Hermes on cheap/free model backends | cut inference costs on an OpenRouter account |
| 13 | Hermes Agent The 24/7 Self-Evolving AI Agent! | WorldofAI | 9:15 | https://www.youtube.com/watch?v=cu2fgknmemA | Always-on operation: gateway, cron, background automation | turn Hermes from a chat window into a 24/7 assistant |
| 14 | Hermes Co-Founder on Building an AI Agent That Improves Itself \| Karan Malhotra | Peter Yang | 46:45 | https://www.youtube.com/watch?v=UWjh5Z4s8jY | Interview on design philosophy (self-improving agents, skills as memory) | where the product is going |
| 15 | Hermes Agent: Agents that grow with you \| Episode #357 | Practical AI | 47:34 | https://www.youtube.com/watch?v=UTZhvPXnmwA | Podcast-depth technical discussion of the agent architecture | the engineering story behind the learning loop |
| 16 | Did Hermes Agent just kill OpenClaw? (full guide) | Alex Finn | 13:55 | https://www.youtube.com/watch?v=tP6yf22OJdI | Guide-style comparison/switch content | a switcher's guide perspective |
| 17 | Hermes Agent: Why Everyone's Ditching OpenClaw in 2026 | Luke Alexander AI | 18:03 | https://www.youtube.com/watch?v=1UgXUjT-QtI | Comparison content | more comparison context |
Suggested order: 1 or 5 first (whichever mood you are in), then 2, then 6 once you have a few weeks of use under your belt.
---
## 7. Your first 7 days
One action per day, each finishable in 15 minutes.
**Day 1 — Import.** `hermes import-agent claude-code --dry-run`, review the preview, then run it without the flag. Your CLAUDE.md context now lives in Hermes.
**Day 2 — Write your SOUL.md.** Open `~/.hermes/SOUL.md` and write who this agent is for you: its role, your tone, three standing rules (e.g., how to grade evidence, where briefs go, how to flag uncertainty). Ten lines is plenty.
**Day 3 — Per-directory sessions.** Pick your most active project folder. Work one task there via `hermes --in DIR --resume latest`. Notice the thread is separate from everything else.
**Day 4 — First skill.** Finish a small repeated workflow (a source-check pass, a brief section). At the end, say "save this as a skill." Next day, watch it load by itself.
**Day 5 — One cron job.** `hermes cron create` for a small daily check (inbox digest, a price or news watch, whatever you already do by hand). Deliver it somewhere you actually look.
**Day 6 — Hand a task to Claude Code.** In Hermes: "delegate this coding task to Claude Code in print mode." Read the JSON result. This is the bridge working.
**Day 7 — Recall test.** Ask Hermes "what did we decide last week about [your project]?" If it can answer from the session store, the stack is working. If not, `/compress` the bloat and try `hermes prompt-size` to see what is eating the window.
---
## 8. What NOT to expect
- **A bigger context window than you have.** Model choice does not change the window size. Fast models on OpenRouter (qwen3.8-flash, deepseek-4-flash) are cheap and quick, but they carry fewer bytes per turn than a frontier model. The harness compresses automatically near the limit — you will not watch a meter like Claude Code's `/context` — but compression is lossy. For the heaviest analytical passes, use `--reasoning high` and a larger model for that run.
- **Model choice as a silver bullet.** What a different model buys: better reasoning, better tool-calling, more reliable long-horizon work. What it does not buy: memory of your projects, your workflows, or last week's decisions. That lives in your context stack, not the model.
- **Desktop = everything.** The desktop app is a thin client over a local agent: config, memory, skills, sessions, cron, and kanban all live in the Hermes home, not in the window. Close the window and the work keeps living; that is a feature, not a bug.
- **Cron limits.** Cron jobs are durable, but they run on their own budgets: wall-clock caps, per-job model overrides, and delivery depends on configured platforms. A job is not an infinite second brain — design it as a bounded task with a bounded output.
- **It will still need to be told things twice.** If you did not save it as memory or a skill, the next session does not know. The learning loop only works if you trigger it. "Remember this" and "save this as a skill" are deliberate moves, not magic.
- **Official tutorial videos.** None exist from Nous Research; the watch list is verified third-party content. The docs (hermes-agent.nousresearch.com/docs) are the authoritative source, and `/help` inside a session lists the exact commands your version supports.
- **Slash commands behave like the CLI does.** The slash registry is version-dependent; anything tagged DOC-ONLY here was confirmed against the docs but not exercised live from a headless session. `/help` in your own session is the final word.
---
## 9. Appendix: command reference
Tags: **VERIFIED-LIVE** = confirmed against `hermes --help` / `hermes <cmd> --help` on v0.21.1 (2026.9.7), reference install (Syslog kagentz), 2026-09-11. **DOC-ONLY** = confirmed against the official docs (slash commands run inside a chat session and were not exercised from a headless research run; their CLI subcommands were verified live).
### Setup & health
| Command | What it does | Tag |
|---|---|---|
| `hermes setup` | Interactive setup wizard | VERIFIED-LIVE |
| `hermes doctor [--fix] [--live]` | Diagnose config/deps; `--fix` auto-repairs | VERIFIED-LIVE |
| `hermes status [--all] [--deep]` | Component status | VERIFIED-LIVE |
| `hermes config show/edit/get/set/unset/path/env-path/check/migrate` | View/edit config | VERIFIED-LIVE |
| `hermes update` | Update Hermes to latest | VERIFIED-LIVE |
### The Claude Code bridge (highest value for you)
| Command | What it does | Tag |
|---|---|---|
| `hermes import-agent claude-code [--dry-run] [--overwrite] [--yes]` | One-command import of a Claude Code setup: maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, memories into Hermes equivalents. Never imports API keys. | VERIFIED-LIVE |
| `hermes import-agent codex` | Same for Codex CLI setups | VERIFIED-LIVE |
| `hermes sessions import` | Import a Claude Code or Codex CLI session/conversation into Hermes | VERIFIED-LIVE |
| `hermes skills trust` | Trust a repo so its project-local skills (`.hermes/skills`) load — the Hermes analog of `.claude/skills/` | VERIFIED-LIVE |
### Daily driving
| Command | What it does | Tag |
|---|---|---|
| `hermes` / `hermes chat` | Interactive session | VERIFIED-LIVE |
| `hermes -c [NAME]` / `hermes --resume <id\|latest>` | Resume by name or ID | VERIFIED-LIVE |
| `hermes --in DIR --resume latest` | Resume the latest session for a directory | VERIFIED-LIVE |
| `hermes -z "PROMPT"` | One-shot: prints ONLY the final answer (scripting/CI); tools, memory, and AGENTS.md still load | VERIFIED-LIVE |
| `hermes chat -q "PROMPT"` | Single-query mode | VERIFIED-LIVE |
| `hermes -m MODEL --provider PROVIDER --reasoning LEVEL` | Per-run model/provider/reasoning overrides (`none…ultra`) | VERIFIED-LIVE |
| `hermes -s SKILL1,SKILL2` | Preload specific skills for the session | VERIFIED-LIVE |
| `hermes -t TOOLSETS` | Restrict toolsets for this run | VERIFIED-LIVE |
| `hermes -w` | Isolated git worktree session (parallel agents on one repo) | VERIFIED-LIVE |
| `hermes chat --checkpoints` | Enable filesystem checkpoints (`/rollback` to restore) | VERIFIED-LIVE |
| `hermes chat --max-turns N` / `--run-budget SECONDS` | Cap loop iterations / wall-clock budget | VERIFIED-LIVE |
| `hermes -yolo` | Bypass command approval prompts (use with care) | VERIFIED-LIVE |
| `hermes pause` / `hermes resume` | Emergency stop / lift (pauses cron, kanban dispatch, gateway turns) | VERIFIED-LIVE |
### Context & memory management
| Command | What it does | Tag |
|---|---|---|
| `hermes memory setup/status/off/reset` | External memory provider management (built-in MEMORY.md/USER.md always active) | VERIFIED-LIVE |
| `hermes sessions list/browse/rename/pin/export/prune/stats` | Session store management | VERIFIED-LIVE |
| `hermes skills list/search/install/inspect/browse/config/check/update` | Skill management | VERIFIED-LIVE |
| `hermes skills trust/untrust` | Repo-local skill trust | VERIFIED-LIVE |
| `hermes curator status/run/pause/pin/...` | Background skill maintenance (auto-archive, backups) | VERIFIED-LIVE |
| `hermes prompt-size` | Byte breakdown of system prompt + tool schemas (context-bloat diagnosis) | VERIFIED-LIVE |
| `hermes insights [--days N]` | Usage analytics | VERIFIED-LIVE |
### Tools, MCP, integrations
| Command | What it does | Tag |
|---|---|---|
| `hermes tools` (interactive) / `list/enable/disable` | Per-platform toolset toggles; MCP tools as `server:tool` | VERIFIED-LIVE |
| `hermes mcp add/remove/list/test/configure/picker/catalog/install` | MCP server management (incl. one-click catalog installs) | VERIFIED-LIVE |
| `hermes mcp serve` | Run Hermes AS an MCP server for other agents | VERIFIED-LIVE |
| `hermes computer-use install/status/doctor` | Desktop-control backend (cua-driver) | VERIFIED-LIVE |
| `hermes gateway run/install/start/status/setup` | Messaging gateway (Telegram, Discord, Slack, WhatsApp, …) | VERIFIED-LIVE |
| `hermes send` | Send a message to a configured platform (scripts/cron/CI) | VERIFIED-LIVE |
### Automation & multi-agent
| Command | What it does | Tag |
|---|---|---|
| `hermes cron list/create/edit/pause/resume/run/remove/doctor` | Scheduled jobs (durable, multi-platform delivery) | VERIFIED-LIVE |
| `hermes cron notepad` | Durable per-job key-value notepad across runs | VERIFIED-LIVE |
| `hermes kanban create/list/show/link/complete/swarm/...` | Durable multi-profile task board (40+ verbs) | VERIFIED-LIVE |
| `hermes kanban swarm` | Generate a parallel-workers → verifier → synthesizer task graph | VERIFIED-LIVE |
| `hermes project create/list/add-folder/bind-board` | Named multi-folder workspaces (desktop Projects) | VERIFIED-LIVE |
| `hermes profile list/create/use/alias/export/import` | Isolated Hermes instances | VERIFIED-LIVE |
| `hermes auth add/list/priority/reset` | Pooled credentials per provider (rotation) | VERIFIED-LIVE |
| `hermes fallback list/add/remove` | Fallback model chain (auto-rollover on failure) | VERIFIED-LIVE |
| `hermes model` | Interactive model/provider picker | VERIFIED-LIVE |
### In-session slash commands (DOC-ONLY)
Source: https://hermes-agent.nousresearch.com/docs/reference/slash-commands
| Command | What it does |
|---|---|
| `/help` | List all commands (authoritative in your version) |
| `/new` (`/reset`) | Fresh session |
| `/resume [name]` | Resume a named/recent session |
| `/branch` (`/fork`) | Branch the current session |
| `/compress` | Manually compress context (auto-compression also exists) |
| `/undo` | Remove last exchange |
| `/retry` | Resend last message |
| `/title [name]` | Name the session |
| `/save` | Save conversation to file |
| `/history` | Show conversation history |
| `/skill <name>` | Load a skill into the session |
| `/skills` | Search/install skills |
| `/reload-skills` | Re-scan skill directory |
| `/tools` / `/toolsets` | Manage tools |
| `/goal [text]` | Set a standing goal the agent works toward across turns (`/goal status/pause/clear` to manage) |
| `/background <prompt>` | Run a prompt in the background |
| `/queue <prompt>` | Queue a prompt for the next turn |
| `/steer <prompt>` | Inject a course-correction after the next tool call without interrupting |
| `/agents` | Show active agents and running tasks |
| `/cron` | Manage cron jobs in-session |
| `/kanban` | Multi-profile collaboration board in-session |
| `/model [name]` | Show/change model mid-session |
| `/reasoning [level]` | Set reasoning effort |
| `/voice [on\|off\|tts]` | Voice mode |
| `/rollback [N]` | Restore filesystem checkpoint (needs `--checkpoints`) |
| `/usage` | Token usage |
| `/insights [days]` | Usage analytics |
| `/platforms` | Gateway platform status |
Note: Hermes compresses automatically near the context limit; no manual threshold watch is needed the way Claude Code's `/context` grid is.
+100
View File
@@ -0,0 +1,100 @@
# Review Results: Scot Murray Hermes Playbook (t_fefdf30b)
**VERDICT: APPROVED-WITH-FIXES**
## Summary
The playbook is well-structured, factually accurate, and provides genuine value for a new Hermes user. All 17 video links verified live (17/17 PASS), all 10 source URLs resolved successfully, all 26+ CLI commands verified against v0.21.1, no client data leaks detected, and all 9 required sections present with substantive content. One minor documentation accuracy issue requires correction.
## Per-Check Results
### 1. COMMANDS: ✅ PASS
- 26 top-level commands and subcommands verified live on Hermes Agent v0.21.1 (2026.9.7)
- All VERIFIED-LIVE tags confirmed: `hermes import-agent claude-code --dry-run`, `hermes skills trust`, `hermes mcp serve`, `hermes prompt-size`, `hermes fallback`, `hermes curator`, etc.
- All DOC-ONLY commands (in-session slash commands) confirmed against official docs
- No fabricated or non-existent commands found
### 2. VIDEO LINKS: ✅ PASS
- All 17 YouTube URLs verified via oEmbed endpoint
- **17/17 PASS** - All titles and channels match the documentation claims
- Videos: https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=<ID>&format=json
- Example verified: Ta2wg6xPaY4 → "Learn 95% of Hermes Agent in 31 Minutes" | Sharbel A. ✅
### 3. SOURCES: ✅ PASS
- 10/10 URLs in sources.md resolved successfully (HTTP 200)
- No dead links or inaccessible URLs found
- All official docs and GitHub repo accessible
### 4. COMPLETENESS: ✅ PASS
- All 9 required sections present and substantive:
- 1. TL;DR ✅
- 2. Context gap explanation ✅
- 3. Context stack ✅
- 4. Top moves (12 items) ✅
- 5. Working alongside Claude Code ✅
- 6. Video watch list (17 videos) ✅
- 7. 7-day ramp ✅
- 8. What NOT to expect ✅
- 9. Appendix: command reference ✅
- Top moves count: 12/12 (within 12 limit) ✅
- 7-day ramp is actionable with specific commands ✅
### 5. CLIENT-DATA LEAK: ✅ PASS
- **No private financial data or personal data found**
- Grep patterns searched: murray, jds, portfolio, allocation, holding, ticker, position, dollar, 192.168.68.17, syslog solution llc
- Only mentions of "portfolio" are generic workflow descriptions, not specific financial data
- No Murray Capital/JDS portfolio details, positions, or dollar figures found
### 6. HONESTY/OVER-CLAIM: ✅ PASS (1 minor issue)
- **Minor issue found:** The "save this as a skill" workflow description is slightly misleading
- Book says: "after you complete a workflow twice, ask Hermes to 'save this as a skill'"
- Reality: The workflow works on a single workflow completion (not after two)
- The language "after you complete a workflow twice" suggests a minimum repetition requirement that doesn't exist
- **Recommendation:** Change to "after completing a workflow, ask Hermes to save this as a skill"
- No major over-claims about features that don't exist
- All model claims are accurate for OpenRouter-only setup
- Honest about "no official Nous Research tutorial videos exist" ✅
### 7. USEFULNESS: ✅ PASS
- **Strongest section:** Section 4 "Top moves" - provides 12 highly actionable, verified commands
- **Strongest section:** Section 6 "Video watch list" - all links verified, titles/channels accurate
- **Strongest section:** Section 7 "Your first 7 days" - practical, incremental onboarding plan
- **Weakest section:** Section 2 "Why Hermes feels like it has less context" - could benefit from more concrete examples
- Overall: Would genuinely help a new user close the context gap with actionable, verified steps
## Prioritized Fixes
### SHOULD-FIX
1. **Fix "save as skill" workflow description** - Change "after you complete a workflow twice" to "after completing a workflow" (Section 3, paragraph 4)
- This is the only minor issue found
- Doesn't affect functionality but could create false expectations about repetition requirements
### NIT
- None identified - all content is accurate and well-organized
## Edits Applied
Applied the SHOULD-FIX correction directly to the playbook:
- **Section 3, Layer 3 (Skills):** Changed "after you complete a workflow twice" → "after completing a workflow"
---
## FINAL SUMMARY
**Verdict: APPROVED-WITH-FIXES**
**Per-check results:**
- Check 1 (COMMANDS): ✅ PASS - 26+ commands verified live
- Check 2 (VIDEO LINKS): ✅ PASS - 17/17 valid with matching titles/channels
- Check 3 (SOURCES): ✅ PASS - 10/10 URLs resolved
- Check 4 (COMPLETENESS): ✅ PASS - 9/9 sections, 12 top moves, actionable 7-day ramp
- Check 5 (CLIENT-DATA LEAK): ✅ PASS - No private data found
- Check 6 (HONESTY/OVER-CLAIM): ✅ PASS - 1 minor issue identified and fixed
- Check 7 (USEFULNESS): ✅ PASS - Strong actionable content
**Video link pass/fail count:** 17/17 pass, 0 fail
**Fabricated/non-existent commands:** None found
**Dead links:** None found
**Path to REVIEW.md:** /home/hermes/syslog/drafts/scot-hermes-playbook/REVIEW.md
+12
View File
@@ -0,0 +1,12 @@
@page { size: A4; margin: 2cm 1.8cm; @bottom-center { content: counter(page); font-size: 9pt; color: #666; } }
body { font-family: 'DejaVu Sans', sans-serif; font-size: 10pt; line-height: 1.5; color: #1a1a1a; }
h1 { font-size: 20pt; border-bottom: 2px solid #222; padding-bottom: 6px; }
h2 { font-size: 14pt; border-bottom: 1px solid #bbb; padding-bottom: 3px; margin-top: 1.4em; }
h3 { font-size: 11.5pt; margin-top: 1.2em; }
code { font-family: 'DejaVu Sans Mono', monospace; font-size: 8.5pt; background: #f2f2f2; padding: 1px 3px; border-radius: 3px; }
pre { background: #f6f6f6; border: 1px solid #ddd; padding: 8px 10px; border-radius: 4px; white-space: pre-wrap; }
pre code { background: none; padding: 0; }
table { border-collapse: collapse; width: 100%; margin: 0.8em 0; font-size: 9pt; }
th, td { border: 1px solid #999; padding: 4px 6px; text-align: left; vertical-align: top; }
th { background: #eee; }
blockquote { border-left: 3px solid #888; margin-left: 0; padding-left: 12px; color: #444; }
@@ -0,0 +1,18 @@
#!/usr/bin/env bash
# Render HERMES-PLAYBOOK-FOR-SCOT.md -> PDF (send-ready package for t_2c716052).
# Toolchain: pandoc (md->html) + system weasyprint (html->pdf), both from Debian repo.
set -euo pipefail
DIR="$(cd "$(dirname "$0")" && pwd)"
SRC="$DIR/HERMES-PLAYBOOK-FOR-SCOT.md"
OUT="$DIR/HERMES-PLAYBOOK-FOR-SCOT.pdf"
echo "source sha256 : $(sha256sum "$SRC" | awk '{print $1}')"
echo "source size : $(wc -c < "$SRC") bytes"
pandoc "$SRC" -f gfm -t html5 -s --metadata title="Hermes Playbook for Scot" \
-c pb.css -o /tmp/pb.html
weasyprint -u "$DIR/" /tmp/pb.html "$OUT"
echo "pdf path : $OUT"
echo "pdf size : $(wc -c < "$OUT") bytes"
echo "pdf sha256 : $(sha256sum "$OUT" | awk '{print $1}')"
@@ -0,0 +1,50 @@
# 01 — The Context Gap: Claude Code vs a Fresh Hermes Install
**Audience:** internal research for the Scot Murray playbook (writer takes over from here).
**Prepared:** 2026-09-11. Primary sources: official docs (hermes-agent.nousresearch.com/docs) and live CLI verification on the reference install (Syslog kagentz). Every "exact command/file" was checked against `hermes --help` / `hermes <cmd> --help` output on v0.21.1 unless marked DOC-ONLY.
## Why the gap exists (30-second framing)
Claude Code discovers context from the repo it sits in: a `CLAUDE.md` it reads on every
launch, `.claude/` folders that ship subagents, slash commands, hooks, and skills. A fresh
Hermes install starts nearly empty by design — its philosophy is that the agent *builds* its
own context over time (memory, skills) and that context comes from config files, not the
repo. The "lack of context" Scot noticed is just Hermes waiting to be seeded. Below: every
gap and the Hermes mechanism that closes it.
## Gap table
| # | Gap | Claude Code behaviour (out of the box) | Hermes equivalent | Exact command / file |
|---|-----|----------------------------------------|-------------------|----------------------|
| 1 | Project instructions | Auto-loads `CLAUDE.md` from project root; `#` prefix adds memory live; `claude /init` scaffolds it | Auto-injects `AGENTS.md` (and `.cursorrules`) from the working directory + `SOUL.md` persona + persistent memory from the Hermes home. `hermes import-agent claude-code` migrates existing CLAUDE.md content in one shot | File: `AGENTS.md` in the project root (git-tracked). Command: `hermes import-agent claude-code [--dry-run]` — VERIFIED-LIVE |
| 2 | Team/personal instruction split | `.claude/rules/*.md` (project) + `~/.claude/rules/*.md` (personal) | Rules via `AGENTS.md` in the repo (team) vs `SOUL.md` + memory in `~/.hermes/` (personal). Config for everything else: `hermes config edit` | Files: `AGENTS.md` (repo), `SOUL.md` (`~/.hermes/`). VERIFIED-LIVE (documented in `--ignore-rules` help text, which names exactly what gets injected) |
| 3 | Slash commands | Ships dozens built-in; custom ones in `.claude/commands/<name>.md` | Rich built-in registry (`/help` to list); custom automation goes into skills instead of command files | In-session: `/help`, `/skills`. Doc: https://hermes-agent.nousresearch.com/docs/reference/slash-commands — VERIFIED-LIVE (registry derived from `hermes_cli/commands.py`) |
| 4 | Subagents / delegation | `.claude/agents/*.md`, `@agent` mentions, Task tool | Built-in `delegate_task` tool (isolated subagent contexts, parallel batches) plus full-process spawns (`hermes chat -q`, tmux) and the durable Kanban board for multi-profile work | In-session: ask Hermes to delegate; `hermes kanban create ...` for durable tasks. Doc: /docs/user-guide/features/kanban. VERIFIED-LIVE (`hermes kanban --help` shows 40+ verbs incl. `swarm`) |
| 5 | Skills (auto-invoked expertise) | `.claude/skills/*.md` markdown guides invoked by natural language match | Same concept, more infrastructure: skills auto-load by task match, can be authored BY the agent itself (`skill_manage`), installed from registries, maintained by the curator | CLI: `hermes skills list/search/install/config`; in-session: `/skill <name>`, `/reload-skills`. VERIFIED-LIVE. Hub: `hermes skills browse` |
| 6 | Project memory / auto-memory | `~/.claude/projects/<project>/memory/`, 25 KB cap | Persistent memory is first-class: built-in `MEMORY.md`/`USER.md` always active, pluggable providers (Honcho, Mem0, …) | CLI: `hermes memory setup/status/off`. VERIFIED-LIVE. Doc: /docs/user-guide/features/memory |
| 7 | Tool permissions | `/permissions`, `settings.json` allowlists | Per-platform toolset toggles + MCP tool allowlists (`server:tool` notation) | CLI: `hermes tools` (interactive UI), `hermes tools list/enable/disable`. VERIFIED-LIVE |
| 8 | MCP servers | `claude mcp add/list/remove`, scopes user/local/project | `hermes mcp add/list/test/configure`, one-click catalog installs, plus `hermes mcp serve` (Hermes AS an MCP server — Claude Code cannot do this) | VERIFIED-LIVE. Doc: /docs/user-guide/features/mcp |
| 9 | Session resume / history | `claude -c`, `claude -r <id>`, `/resume` | `hermes -c`, `hermes --resume <id|latest|title>`, named sessions, plus a durable SQLite store with search/export/pin | CLI: `hermes sessions list/browse/rename/pin/export`. VERIFIED-LIVE |
| 10 | Cost & context visibility | `/cost`, `/context` grid | `/usage`, `/insights [days]`, `/compress` (auto-compression built in), `/prompt-size` byte breakdown | VERIFIED-LIVE (`insights`, `logs` subcommands confirmed in `hermes --help`) |
| 11 | Headless/CI mode | `claude -p` print mode | `-z/--oneshot` flag (prints only final response) + `hermes chat -q` | VERIFIED-LIVE |
| 12 | Import of prior setup | n/a (it IS the incumbent) | **The single most important one for Scot:** `hermes import-agent claude-code` maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, and memories into Hermes equivalents (never API keys) | `hermes import-agent claude-code --dry-run` then without `--dry-run`. VERIFIED-LIVE. Also `hermes sessions import` for old Claude Code conversations — VERIFIED-LIVE |
| 13 | Hooks on tool events | 8 hook types in `settings.json` (PreToolUse, PostToolUse, …) | Shell-script hooks managed via `hermes hooks` | CLI: `hermes hooks`. VERIFIED-LIVE (in top-level command list) |
| 14 | Scheduled / recurring work | `claude /loop` (in-session only) | Durable cron scheduler with multi-platform delivery, chained outputs (`context_from`), per-job model overrides | CLI: `hermes cron list/create/edit/pause/resume/run/remove/doctor`. VERIFIED-LIVE. Doc: /docs/user-guide/features/cron |
## The one-command bridge (lead with this in the playbook)
```bash
hermes import-agent claude-code --dry-run # preview
hermes import-agent claude-code # migrate CLAUDE.md → AGENTS.md, MCP, skills, memories
```
This is the fastest way to eliminate the "Hermes has no context" feeling for someone who
already has a working Claude Code setup: it carries over the exact instructions and servers
that made Claude Code feel context-rich. Preview-only mode exists (`--dry-run`), it never
imports credentials, and conflicts are skipped by default (`--overwrite` to change).
## Sources
- Live CLI: `hermes --help`, `hermes chat --help`, `hermes import-agent --help`, `hermes kanban --help`, `hermes skills --help`, `hermes sessions --help`, `hermes mcp --help`, `hermes tools --help`, `hermes memory --help`, `hermes project --help`, `hermes cron --help`, `hermes config --help`, `hermes profile --help`, `hermes computer-use --help` on v0.21.1, reference install (Syslog kagentz), 2026-09-11. Raw dump: `cli-help-dump.txt` next to this file.
- Docs: https://hermes-agent.nousresearch.com/docs/ (index) — all URLs in sources.md
- Claude Code side: installed skill `delegate-coding-agent/references/claude-code.md` (Hermes Agent + Teknium, v2.2.1), `/home/hermes/.hermes/skills/autonomous-ai-agents/`
@@ -0,0 +1,101 @@
# 02 — High-Leverage Hermes Surfaces (the "harness power" inventory)
**Prepared:** 2026-09-11. Each surface: what it does, when to use it, exact command/file, doc URL. Verification: V-LIVE = confirmed against live CLI v0.21.1 on the reference install (Syslog kagentz); V-DOC = confirmed against official docs page (URL resolved HTTP 200); V-FILE = present on this machine's installed skills.
---
### 1. Persona / SOUL file
- **What:** `SOUL.md` is Hermes' personality + standing-identity file, auto-injected into the system prompt alongside `AGENTS.md` rules and memory (confirmed by `--ignore-rules` help text which lists exactly what gets injected).
- **When:** client wants the agent to have a consistent voice/role (e.g., "you are my analyst").
- **Where:** `~/.hermes/SOUL.md` (per-profile: `~/.hermes/profiles/<name>/SOUL.md`).
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/configuration [V-DOC]
### 2. Persistent memory (built-in + providers)
- **What:** Built-in `MEMORY.md` / `USER.md` always active; optional external providers (honcho, mem0, hindsight, byterover, …). Memory is injected every session — this is the single biggest cure for "it forgets my project."
- **When:** after any correction or preference the user states ("use bun, not npm") — tell Hermes to remember it and it persists.
- **Command:** `hermes memory setup|status|off|reset` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/memory [V-DOC]
### 3. Skills + skill authoring (the learning loop)
- **What:** Markdown procedure files that auto-load when a task matches. The differentiator: Hermes can WRITE its own skills after learning a workflow (self-improving), and the curator maintains them (usage tracking, archiving, backups).
- **When:** any workflow done twice — say "save this as a skill."
- **Commands:** `hermes skills list|search|install|browse|config|check|update` [V-LIVE]; in-session `/skill <name>`, `/reload-skills` [V-DOC]; authoring tool in-session is `skill_manage` (agent-side; writer should describe it as "ask your Hermes to save the procedure as a skill").
- **Docs:** https://hermes-agent.nousresearch.com/docs/reference/skills-catalog [V-DOC]; curator: https://hermes-agent.nousresearch.com/docs/user-guide/features/curator [V-DOC]
### 4. Desktop Projects
- **What:** Human-named workspaces spanning multiple folders/repos; anchor desktop session grouping; bindable to a Kanban board for deterministic worktree/branch conventions.
- **When:** Scot's multi-repo workflows (portfolio ops). `hermes project create <name>` then `add-folder`.
- **Command:** `hermes project create|list|show|add-folder|set-primary|use|bind-board` [V-LIVE]
### 5. MCP servers
- **What:** Plug external tools into the agent (GitHub, Postgres, n8n, …) via the Model Context Protocol. Also runs in reverse: `hermes mcp serve` exposes Hermes conversations to other agents.
- **Command:** `hermes mcp add|list|test|configure|picker|catalog|install|serve` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp [V-DOC]
### 6. Toolsets & deferred tool discovery
- **What:** ~30 built-in toolsets (web, browser, terminal, memory, kanban, tts, …) toggled per platform via `hermes tools`; the agent can also defer-load more tools at runtime via `tool_search` instead of carrying every schema in context.
- **When:** trim toolsets for focus/cost, or enable `browser` for web work.
- **Command:** `hermes tools` (interactive), `hermes tools list|enable|disable` [V-LIVE]; docs: https://hermes-agent.nousresearch.com/docs/reference/tools-reference [V-DOC]
### 7. Subagent delegation (delegate_task)
- **What:** In-session parallel subagents with isolated context + terminal sessions; leaf vs orchestrator roles; batched parallel spawns.
- **When:** research fan-out, parallel code review, anything that would flood the main context.
- **Command:** agent-side tool (no CLI). In-session: ask Hermes to "delegate X to subagents." Docs: /docs/user-guide/features (delegation section) [V-DOC]
### 8. Kanban (durable multi-agent board)
- **What:** SQLite board shared across profiles; tasks with dependencies, atomic claims, isolated workspaces, dispatcher; `swarm` verb builds parallel-worker → verifier → synthesizer graphs.
- **When:** recurring multi-step operations, handoffs between specialist profiles, long-running campaigns that must survive restarts.
- **Command:** `hermes kanban create|list|show|swarm|link|complete|watch|stats|dispatch` (40+ verbs) [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/kanban [V-DOC]
### 9. Cron jobs
- **What:** Durable scheduler: duration or cron syntax, per-job model/skills overrides, output chaining (`context_from`), multi-platform delivery.
- **When:** daily reports, monitoring with alerts, weekly reviews.
- **Command:** `hermes cron list|create|edit|pause|resume|run|remove|doctor|status` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/cron [V-DOC]
### 10. Session store + session_search
- **What:** All conversations in a searchable SQLite store: resume by ID/name/`latest`, pin, export to JSONL/Markdown, prune, stats.
- **When:** "what did we decide last week" — the agent can search past sessions; user can browse them.
- **Command:** `hermes sessions list|browse|rename|pin|export|prune|stats` [V-LIVE]; in-session `/resume`, `/branch` [V-DOC]
### 11. Browser + computer use
- **What:** Two surfaces: headless browser automation (browser toolset: navigate/click/snapshot) and full desktop control via `computer_use` (cua-driver, macOS/Windows/Linux, background-first input that never steals focus).
- **When:** web research → headless browser; native apps (Excel, Figma, native chat) → computer use.
- **Command:** `hermes computer-use install|status|doctor` [V-LIVE]; enable via `hermes tools` [V-LIVE]
### 12. Model/provider routing, credential pools, fallbacks
- **What:** Per-invocation model/provider overrides; interactive model picker; pooled credentials with rotation; explicit fallback chains; per-task model overrides on Kanban.
- **Command:** `hermes model` [V-LIVE], `hermes fallback list|add|remove` [V-LIVE], `hermes auth add|list|priority|reset` [V-LIVE]; per-run flags `-m`, `--provider`, `--reasoning` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/integrations/providers [V-DOC]
- **Note for Scot (OpenRouter + small fast models):** `--reasoning high` on hard tasks; `hermes fallback add` so a failed call rolls to a second model instead of erroring.
### 13. Profiles (isolated instances)
- **What:** Completely independent Hermes instances (config, memory, skills, sessions) with wrapper aliases; export/import for distribution.
- **When:** separate work/persona contexts, or one profile per client.
- **Command:** `hermes profile list|create|use|alias|export|import` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/profiles [V-DOC]
### 14. Goal loops
- **What:** `/goal <text>` sets a standing objective the agent keeps working toward across turns until achieved (judge-checked continuations).
- **When:** "keep the CI green until it passes," "keep researching until you have 5 verified sources."
- **Command:** in-session `/goal [text|status|pause|resume|clear]` [V-DOC: /docs/reference/slash-commands]
### 15. Gateway (messaging platform front-end)
- **What:** The same agent reachable from Telegram, Discord, Slack, WhatsApp, Signal, Email, and 10+ platforms with full tool access; runs as a background service.
- **Command:** `hermes gateway run|install|start|status|setup` [V-LIVE]
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ [V-DOC]
### 16. Checkpoints & rollback
- **What:** Filesystem snapshots before destructive file operations; `/rollback [N]` restores.
- **When:** letting the agent loose on important files.
- **Command:** `hermes chat --checkpoints` / `hermes checkpoints` [V-LIVE]; in-session `/rollback`, `/snapshot` [V-DOC]
### 17. Projects↔Kanban binding + worktree mode
- **What:** `hermes project bind-board` ties a board to a project (deterministic worktree + branch per task); `-w/--worktree` runs any session in an isolated git worktree.
- **When:** parallel coding agents that must not collide.
- **Command:** `hermes project bind-board` [V-LIVE]; `hermes -w` [V-LIVE]
### 18. Prompt-size introspection
- **What:** Byte breakdown of system prompt + tool schemas — diagnose why responses feel "dumb" (usually context bloat).
- **Command:** `hermes prompt-size` [V-LIVE]
@@ -0,0 +1,121 @@
# 03 — Command Cheatsheet (every entry verified)
**Verification method:** each VERIFIED-LIVE entry was confirmed against `hermes --help` or `hermes <cmd> --help` on Hermes Agent v0.21.1 (2026.9.7), reference install (Syslog kagentz), 2026-09-11. Raw output: `cli-help-dump.txt`. DOC-ONLY entries come from the official docs (URL given). Nothing is invented.
## (a) CLI — `hermes ...`
### Setup & health
| Command | What it does | Tag |
|---|---|---|
| `hermes setup` | Interactive setup wizard | VERIFIED-LIVE |
| `hermes doctor [--fix] [--live]` | Diagnose config/deps; `--fix` auto-repairs | VERIFIED-LIVE |
| `hermes status [--all] [--deep]` | Component status | VERIFIED-LIVE |
| `hermes config show/edit/get/set/unset/path/env-path/check/migrate` | View/edit config | VERIFIED-LIVE |
| `hermes update` | Update Hermes to latest | VERIFIED-LIVE |
### The Claude Code bridge (highest value for Scot)
| Command | What it does | Tag |
|---|---|---|
| `hermes import-agent claude-code [--dry-run] [--overwrite] [--yes]` | One-command import of a Claude Code setup: maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, memories into Hermes equivalents. Never imports API keys. | VERIFIED-LIVE |
| `hermes import-agent codex` | Same for Codex CLI setups | VERIFIED-LIVE |
| `hermes sessions import` | Import a Claude Code or Codex CLI **session/conversation** into Hermes | VERIFIED-LIVE (subcommand listed in `hermes sessions --help`) |
| `hermes skills trust` | Trust a repo so its project-local skills (`./.hermes/skills`) load — the Hermes analog of `.claude/skills/` | VERIFIED-LIVE |
### Daily driving
| Command | What it does | Tag |
|---|---|---|
| `hermes` / `hermes chat` | Interactive session | VERIFIED-LIVE |
| `hermes -c [NAME]` / `hermes --resume <id\|latest>` | Resume by name or ID | VERIFIED-LIVE |
| `hermes --in DIR --resume latest` | Resume the latest session for a directory | VERIFIED-LIVE |
| `hermes -z "PROMPT"` | One-shot: prints ONLY the final answer (scripting/CI); tools, memory, and AGENTS.md still load | VERIFIED-LIVE |
| `hermes chat -q "PROMPT"` | Single-query mode | VERIFIED-LIVE |
| `hermes -m MODEL --provider PROVIDER --reasoning LEVEL` | Per-run model/provider/reasoning overrides (`none…ultra`) | VERIFIED-LIVE |
| `hermes -s SKILL1,SKILL2` | Preload specific skills for the session | VERIFIED-LIVE |
| `hermes -t TOOLSETS` | Restrict toolsets for this run | VERIFIED-LIVE |
| `hermes -w` | Isolated git worktree session (parallel agents on one repo) | VERIFIED-LIVE |
| `hermes chat --checkpoints` | Enable filesystem checkpoints (`/rollback` to restore) | VERIFIED-LIVE |
| `hermes chat --max-turns N` / `--run-budget SECONDS` | Cap loop iterations / wall-clock budget | VERIFIED-LIVE |
### Context & memory management
| Command | What it does | Tag |
|---|---|---|
| `hermes memory setup/status/off/reset` | External memory provider management (built-in MEMORY.md/USER.md always active) | VERIFIED-LIVE |
| `hermes sessions list/browse/rename/pin/export/prune/stats` | Session store management | VERIFIED-LIVE |
| `hermes skills list/search/install/inspect/browse/config/check/update` | Skill management | VERIFIED-LIVE |
| `hermes skills trust/untrust` | Repo-local skill trust | VERIFIED-LIVE |
| `hermes curator status/run/pause/pin/...` | Background skill maintenance (auto-archive, backups) | VERIFIED-LIVE |
| `hermes prompt-size` | Byte breakdown of system prompt + tool schemas (context-bloat diagnosis) | VERIFIED-LIVE |
| `hermes insights [--days N]` | Usage analytics | VERIFIED-LIVE |
### Tools, MCP, integrations
| Command | What it does | Tag |
|---|---|---|
| `hermes tools` (interactive) / `list/enable/disable` | Per-platform toolset toggles; MCP tools as `server:tool` | VERIFIED-LIVE |
| `hermes mcp add/remove/list/test/configure/picker/catalog/install` | MCP server management (incl. one-click catalog installs) | VERIFIED-LIVE |
| `hermes mcp serve` | Run Hermes AS an MCP server for other agents | VERIFIED-LIVE |
| `hermes computer-use install/status/doctor` | Desktop-control backend (cua-driver) | VERIFIED-LIVE |
| `hermes gateway run/install/start/status/setup` | Messaging gateway (Telegram, Discord, Slack, WhatsApp, …) | VERIFIED-LIVE |
| `hermes send` | Send a message to a configured platform (scripts/cron/CI) | VERIFIED-LIVE |
### Automation & multi-agent
| Command | What it does | Tag |
|---|---|---|
| `hermes cron list/create/edit/pause/resume/run/remove/doctor` | Scheduled jobs (durable, multi-platform delivery) | VERIFIED-LIVE |
| `hermes cron notepad` | Durable per-job key-value notepad across runs | VERIFIED-LIVE |
| `hermes kanban create/list/show/link/complete/swarm/...` | Durable multi-profile task board (40+ verbs) | VERIFIED-LIVE |
| `hermes kanban swarm` | Generate a parallel-workers → verifier → synthesizer task graph | VERIFIED-LIVE |
| `hermes project create/list/add-folder/bind-board` | Named multi-folder workspaces (desktop Projects) | VERIFIED-LIVE |
| `hermes profile list/create/use/alias/export/import` | Isolated Hermes instances | VERIFIED-LIVE |
| `hermes auth add/list/priority/reset` | Pooled credentials per provider (rotation) | VERIFIED-LIVE |
| `hermes fallback list/add/remove` | Fallback model chain (auto-rollover on failure) | VERIFIED-LIVE |
| `hermes model` | Interactive model/provider picker | VERIFIED-LIVE |
| `hermes -yolo` | Bypass command approval prompts (use with care) | VERIFIED-LIVE |
| `hermes pause` / `hermes resume` | Emergency stop / lift (pauses cron, kanban dispatch, gateway turns) | VERIFIED-LIVE |
## (b) In-session slash commands
Source: official slash-commands reference https://hermes-agent.nousresearch.com/docs/reference/slash-commands (DOC-ONLY — slash commands run inside a chat session and were not exercised from this headless research run; the CLI subcommands they map to were verified live). DOC-ONLY.
### Context & session
| Command | What it does |
|---|---|
| `/help` | List all commands (authoritative in your version) |
| `/new` (`/reset`) | Fresh session |
| `/resume [name]` | Resume a named/recent session |
| `/branch` (`/fork`) | Branch the current session |
| `/compress` | Manually compress context (auto-compression also exists) |
| `/undo` | Remove last exchange |
| `/retry` | Resend last message |
| `/title [name]` | Name the session |
| `/save` | Save conversation to file |
| `/history` | Show conversation history |
### Power surfaces
| Command | What it does |
|---|---|
| `/skill <name>` | Load a skill into the session |
| `/skills` | Search/install skills |
| `/reload-skills` | Re-scan skill directory |
| `/tools` / `/toolsets` | Manage tools |
| `/goal [text]` | Set a standing goal the agent works toward across turns (`/goal status/pause/clear` to manage) |
| `/background <prompt>` | Run a prompt in the background |
| `/queue <prompt>` | Queue a prompt for the next turn |
| `/steer <prompt>` | Inject a course-correction after the next tool call without interrupting |
| `/agents` | Show active agents and running tasks |
| `/cron` | Manage cron jobs in-session |
| `/kanban` | Multi-profile collaboration board in-session |
| `/model [name]` | Show/change model mid-session |
| `/reasoning [level]` | Set reasoning effort |
| `/voice [on\|off\|tts]` | Voice mode |
| `/rollback [N]` | Restore filesystem checkpoint (needs `--checkpoints`) |
| `/usage` | Token usage |
| `/insights [days]` | Usage analytics |
| `/platforms` | Gateway platform status |
| `/compact`-equivalent note | Hermes compresses automatically near the context limit; no manual threshold watch needed like Claude Code's `/context` |
### The "work alongside Claude Code" shortlist
1. `hermes import-agent claude-code --dry-run` → migrate the setup (VERIFIED-LIVE)
2. `hermes sessions import` → bring the conversation history over (VERIFIED-LIVE)
3. `hermes skills trust` → load repo-local skills like `.claude/skills/` (VERIFIED-LIVE)
4. `hermes -c` / `hermes --in <repo> --resume latest` → per-directory session continuity (VERIFIED-LIVE)
5. `hermes mcp serve` → expose Hermes to Claude Code as an MCP server (VERIFIED-LIVE) — the reverse direction Claude Code can't do
@@ -0,0 +1,76 @@
# 04 — The Claude Code Bridge: Running Hermes WITH Claude Code
**Prepared:** 2026-09-11. Scot already runs both tools. This file documents the proven integration patterns, citing the installed skills on this host (paths under `/home/hermes/.hermes/skills/`) and official docs.
## Pattern 0 — Import (do this first)
`hermes import-agent claude-code` [VERIFIED-LIVE] maps CLAUDE.md/AGENTS.md instructions,
permission allowlists, MCP servers, skills, and memories into Hermes equivalents. It always
shows a preview, never imports credentials. `hermes sessions import` [VERIFIED-LIVE] pulls
in old Claude Code conversations. After import, Hermes "knows" the projects — the context
gap disappears on day one.
## Pattern 1 — Hermes as orchestrator, Claude Code as worker
Source: installed skill **`autonomous-ai-agents/delegate-coding-agent`** (v1.0.0) + its
reference `references/claude-code.md` (v2.2.1) [V-FILE]. The skill is an official Hermes
skill authored for exactly this.
Two orchestration modes (verbatim from the skill):
- **Print mode (preferred):** `claude -p '<task>' --allowedTools 'Read,Edit' --max-turns 10` —
one-shot, no dialogs, structured JSON output with `session_id`, `num_turns`,
`total_cost_usd`. Ask Hermes: *"delegate this coding task to Claude Code in print mode."*
- **Interactive PTY via tmux:** multi-turn sessions — Hermes starts `tmux new-session`,
sends prompts with `send-keys`, monitors with `capture-pane`. For iterative
refactor → review → fix cycles.
Cross-agent review loop (also from the skill):
```
git diff main...feature | claude -p 'Review this diff for bugs and security issues.' --max-turns 1
```
Hermes runs this, reads the output, and fixes findings itself — Claude Code becomes a
reviewer Hermes coordinates.
Safety rails the skill prescribes: explicit `workdir`, clean git status before launch,
narrow task prompts, `git diff` review, targeted tests before committing.
## Pattern 2 — Parallel workstreams + neutral merge reconciliation
Source: installed skill **`autonomous-ai-agents/merge-reconciler`** [V-FILE].
When Hermes and Claude Code (or two Hermes workers) both edit the same repo and collide:
- Do NOT let either agent resolve the conflict — both are biased toward their own side.
- Spawn a **neutral third agent** with the merge-reconciler skill; it classifies every
conflicted hunk (disjoint-intent / same-question-different-answer / superseded), resolves
under an impartiality contract (touch only conflict markers, surface every design call),
verifies with build/tests, and hands back a summary naming every hunk decision.
- Kanban-native shape: a reconciliation card assigned to a **third profile** with both
workers' cards as parents — parent links carry both sides' completion summaries into the
reconciler's context automatically.
## Pattern 3 — Hermes as MCP server (Claude Code gets Hermes tools)
`hermes mcp serve` [VERIFIED-LIVE] runs Hermes as an MCP server exposing its conversations
and capabilities. Claude Code supports MCP clients (`claude mcp add`), so Claude Code can
consume Hermes as a tool provider — persistent memory, skills, cron — the surfaces Claude
Code lacks. This is the reverse-bridge only Hermes can offer. Docs:
https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp
## Pattern 4 — Import legacy sessions for continuity
`hermes sessions import` [VERIFIED-LIVE] imports a Claude Code session into the Hermes
store; from then on `hermes --resume <id>` / `hermes sessions browse` treat it as native
history. Use when mid-project: the new agent picks up exactly where Claude Code left off.
## Pattern 5 — Desktop GUI automation either side can use
Source: installed skill **`autonomous-ai-agents/computer-use`** (v2.0.0) [V-FILE].
`hermes computer-use install` sets up cua-driver; the `computer_use` toolset drives native
desktop apps background-first (never steals focus/cursor), any-model, cross-platform.
Relevant to the bridge because Claude Code has no desktop automation — if a task needs
Figma/Excel/native apps, that part routes to Hermes while the code routes to Claude Code.
Cmd: `hermes computer-use doctor` for health checks.
## Pattern 6 — The import-agent philosophy in one line
Claude Code holds repo context in `CLAUDE.md`; Hermes holds it in `AGENTS.md` + memory +
skills. `hermes import-agent claude-code` translates the first; the learning loop
("save this as a skill") rebuilds the rest automatically the more Scot uses Hermes.
## Reference paths (for the writer)
- `/home/hermes/.hermes/skills/autonomous-ai-agents/delegate-coding-agent/SKILL.md` and `references/claude-code.md`
- `/home/hermes/.hermes/skills/autonomous-ai-agents/merge-reconciler/SKILL.md`
- `/home/hermes/.hermes/skills/autonomous-ai-agents/computer-use/SKILL.md`
- `hermes import-agent --help` raw output in `cli-help-dump.txt` (lines 554-577)
@@ -0,0 +1,119 @@
# 05 — The Video Watch List (every URL verified 2026-09-11)
**Verification method:** each video was found via YouTube search-results scrape (`videoRenderer` metadata), then confirmed with the YouTube oEmbed endpoint (`curl -s "https://www.youtube.com/oembed?url=<URL>&format=json"`) — every entry below returned HTTP 200 with matching title/author (status PASS). Publish dates, durations, and view counts were read from each watch page's metadata. Raw evidence for all 21 entries: `video-verification.json` in this directory. **21/21 PASS, 0 FAIL.**
## Tier 1 — Hermes-specific, start here
### 1. Learn 95% of Hermes Agent in 31 Minutes
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=Ta2wg6xPaY4 | **Duration:** 31:28 | **Published:** 2026-08-09 | **Views:** ~129k
- **oEmbed:** PASS (title/author match)
- **What it demonstrates:** end-to-end Hermes fundamentals — install, sessions, skills, memory, the learning loop. The most complete single-video orientation found.
- **Watch this when you want** the fastest real overview of the whole harness before touching config.
### 2. Hermes Agent Fundamentals In 29 Minutes
- **Channel:** Tina Huang | **URL:** https://www.youtube.com/watch?v=5_N84t1rUU0 | **Duration:** 29:40 | **Published:** 2026-07-20 | **Views:** ~463k (highest-reach Hermes video found)
- **oEmbed:** PASS
- **What it demonstrates:** conceptual grounding — why Hermes' memory/skills loop differs from one-shot coding agents; practical walkthrough.
- **Watch this when you want** to understand *why* Hermes feels different from Claude Code, not just which buttons to press.
### 3. Every Level of Hermes Agent Explained
- **Channel:** Jack Roberts | **URL:** https://www.youtube.com/watch?v=6GtF_uHbGhw | **Duration:** 25:35 | **Published:** 2026-06-17 | **Views:** ~163k
- **oEmbed:** PASS
- **What it demonstrates:** beginner → advanced ladder of features (memory, skills, automation, multi-agent).
- **Watch this when you want** a map of what to learn next after the basics.
### 4. Hermes Agent Full Tutorial INSTALLATION + USECASES
- **Channel:** CodeHead | **URL:** https://www.youtube.com/watch?v=8GjyOQy19so | **Duration:** 7:47 | **Published:** 2026-05-14 | **Views:** ~64k
- **oEmbed:** PASS
- **What it demonstrates:** install through real use-cases, compact.
- **Watch this when you want** a quick install-to-value demo to share with a colleague.
### 5. Hermes Agent Explained In 5 Minutes
- **Channel:** CodeHead | **URL:** https://www.youtube.com/watch?v=9GpWELm3_XI | **Duration:** 4:53 | **Published:** 2026-05-23 | **Views:** ~251k
- **oEmbed:** PASS
- **What it demonstrates:** 5-minute conceptual pitch of the agent and its learning loop.
- **Watch this when you want** the elevator pitch before committing 30 minutes.
## Tier 2 — Hermes-specific deep dives
### 6. 100 Days With Hermes Agent in 21 Minutes
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=sCa3BtpkziQ | **Duration:** 21:19 | **Published:** 2026-06-17 | **Views:** ~58k
- **oEmbed:** PASS
- **What it demonstrates:** long-horizon usage — what memory/skills accumulation actually looks like after months of daily use.
- **Watch this when you want** to see the payoff of the learning loop over time.
### 7. Hermes Agent - Crash Course for Beginners (AI Agent)
- **Channel:** Adrian Twarog | **URL:** https://www.youtube.com/watch?v=4sAmpcSOVEw | **Duration:** 22:19 | **Published:** 2026-07-21 | **Views:** ~46k
- **oEmbed:** PASS
- **What it demonstrates:** beginner crash course from a well-known dev-YouTube creator.
- **Watch this when you want** a second independent explanation of the basics.
### 8. Hermes Agent: The Ultimate Beginner's Guide
- **Channel:** Metics Media | **URL:** https://www.youtube.com/watch?v=CwPUOVUdApE | **Duration:** 37:08 | **Published:** 2026-04-24 | **Views:** ~119k
- **oEmbed:** PASS
- **What it demonstrates:** long-form beginner guide incl. setup and everyday workflows.
- **Watch this when you want** the most thorough single walkthrough in one sitting.
### 9. Hermes Agent Just Killed OpenClaw (Full Tutorial)
- **Channel:** Leon van Zyl | **URL:** https://www.youtube.com/watch?v=jmtpYUOr7_U | **Duration:** 19:59 | **Published:** 2026-04-28 | **Views:** ~16k
- **oEmbed:** PASS
- **What it demonstrates:** full tutorial framing Hermes against the OpenClaw workflow (MCP config, memory, agents).
- **Watch this when you want** a practitioner's feature-by-feature tutorial.
### 10. Hermes Agent vs OpenClaw
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=zwqhemjHq3E | **Duration:** 15:28 | **Published:** 2026-04-20 | **Views:** ~35k
- **oEmbed:** PASS
- **What it demonstrates:** head-to-head comparison of the two agent harnesses.
- **Watch this when you want** the tradeoffs between Hermes and its main alternative.
### 11. Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model
- **Channel:** Tonbi's AI Garage | **URL:** https://www.youtube.com/watch?v=8tpuky8HpXw | **Duration:** 15:08 | **Published:** 2026-03-11 | **Views:** ~18k
- **oEmbed:** PASS
- **What it demonstrates:** Hermes driven by an OpenRouter-served open model (Qwen 3) — directly relevant to an OpenRouter-connected install.
- **Watch this when you want** to see how small open models behave inside Hermes.
### 12. Use This To Make The Hermes Agent Basically Free
- **Channel:** AI LABS | **URL:** https://www.youtube.com/watch?v=5d02TYoOzfE | **Duration:** 13:08 | **Published:** 2026-07-01 | **Views:** ~55k
- **oEmbed:** PASS
- **What it demonstrates:** running Hermes on cheap/free model backends.
- **Watch this when you want** to cut inference costs on an OpenRouter account.
### 13. Hermes Agent The 24/7 Self-Evolving AI Agent!
- **Channel:** WorldofAI | **URL:** https://www.youtube.com/watch?v=cu2fgknmemA | **Duration:** 9:15 | **Published:** 2026-04-07 | **Views:** ~47k
- **oEmbed:** PASS
- **What it demonstrates:** always-on operation: gateway, cron, background automation.
- **Watch this when you want** to turn Hermes from a chat window into a 24/7 assistant.
## Tier 3 — Adjacent (origin/philosophy; not tutorials)
### 14. Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra
- **Channel:** Peter Yang | **URL:** https://www.youtube.com/watch?v=UWjh5Z4s8jY | **Duration:** 46:45 | **Published:** 2026-08-02 | **Views:** ~37k
- **oEmbed:** PASS
- **What it demonstrates:** interview with Hermes' co-founder on the design philosophy (self-improving agents, skills as memory).
- **Watch this when you want** to understand where the product is going.
### 15. Hermes Agent: Agents that grow with you | Episode #357
- **Channel:** Practical AI | **URL:** https://www.youtube.com/watch?v=UTZhvPXnmwA | **Duration:** 47:34 | **Published:** 2026-05-20 | **Views:** ~1.9k
- **oEmbed:** PASS
- **What it demonstrates:** podcast-depth technical discussion of the agent architecture.
- **Watch this when you want** the engineering story behind the learning loop.
### 16. Did Hermes Agent just kill OpenClaw? (full guide)
- **Channel:** Alex Finn | **URL:** https://www.youtube.com/watch?v=tP6yf22OJdI | **Duration:** 13:55 | **Published:** 2026-03-31 | **Views:** ~132k
- **oEmbed:** PASS
- **What it demonstrates:** guide-style comparison/switch content.
- **Watch this when you want** a switcher's guide perspective.
### 17. Hermes Agent: Why Everyone's Ditching OpenClaw in 2026
- **Channel:** Luke Alexander AI | **URL:** https://www.youtube.com/watch?v=1UgXUjT-QtI | **Duration:** 18:03 | **Published:** 2026-03-26 | **Views:** ~14k
- **oEmbed:** PASS — adjacent, comparison content.
- **Watch this when you want** more comparison context.
## Honesty note (required by task spec)
At least 16 of the 17 entries above are directly Hermes-specific (not merely adjacent); the
"fewer than 5 exist" fallback clause was NOT needed — no padding was necessary. Entries
found in search but excluded as thin/low-signal: `iqN6MVzpJTk` (3.1k views, news-style),
`83nWNRKZTCE` (465 views), `6M2tItdARew` (1.2k views), `P2LIFtrRr2U` (promo-style) — all
also verified PASS and kept in `video-verification.json` as spares. No official Nous
Research YouTube tutorial channel was found in searches; the strongest signal of
Hermes-specific video content is the third-party ecosystem above.
@@ -0,0 +1,577 @@
===== hermes chat --help =====
usage: hermes chat [-h] [-q QUERY | --query-file PATH] [--oneshot]
[--image IMAGE] [-m MODEL] [-t TOOLSETS]
[--reasoning LEVEL] [-s SKILLS] [--provider PROVIDER] [-v]
[-Q] [--resume SESSION_ID] [--no-restore-cwd] [--in DIR]
[--continue [SESSION_NAME]] [--create-if-missing]
[--worktree] [--accept-hooks] [--checkpoints]
[--max-turns N] [--run-budget SECONDS] [--yolo]
[--pass-session-id] [--ignore-user-config] [--ignore-rules]
[--safe-mode] [--source SOURCE] [--tui] [--cli] [--dev]
Start an interactive chat session with Hermes Agent
options:
-h, --help show this help message and exit
-q, --query QUERY Query to run. On a real TTY the prompt seeds an
interactive session (submitted literally as the first
turn); combined with --oneshot or -Q, or on a non-TTY,
it answers and exits.
--query-file PATH Read the single query from a file instead of the
command line ('-' reads stdin). Safe for arbitrary
text: nothing is shell-interpreted, so quotes, $(...),
and backticks are preserved verbatim. Mutually
exclusive with -q.
--oneshot With -q/--query-file: answer the query and exit
(legacy single-query behavior) instead of seeding an
interactive session. Implied on non-TTY stdio and by
-Q/--quiet.
--image IMAGE Optional local image path to attach to a single query
-m, --model MODEL Model to use (e.g., anthropic/claude-sonnet-4)
-t, --toolsets TOOLSETS
Comma-separated toolsets to enable
--reasoning LEVEL Reasoning effort for this session: none, minimal, low,
medium, high, xhigh, max, or ultra. Overrides
agent.reasoning_effort for this run only (same levels
as the /reasoning slash command).
-s, --skills SKILLS Preload one or more skills for the session (repeat
flag or comma-separate)
--provider PROVIDER Inference provider (default: auto). Built-in or a
user-defined name from `providers:` in config.yaml.
-v, --verbose Verbose output
-Q, --quiet Quiet mode for programmatic use: suppress banner,
spinner, and tool previews. Only output the final
response and session info.
--resume, -r SESSION_ID
Resume a previous session by ID (shown on exit), or
'latest' for the most recent session
--no-restore-cwd Don't cd into a resumed session's recorded working
directory.
--in DIR Change into DIR before starting or resuming (scopes '
--resume latest' / -c lookups to DIR's workspace).
--continue, -c [SESSION_NAME]
Resume a session by name, or the most recent if no
name given
--create-if-missing With -c/--continue <name>: if no session matches the
name, create a new session with that title and proceed
(instead of failing with a not-found error).
Programmatic callers that want 'send to this named
thread, making it if needed'.
--worktree, -w Run in an isolated git worktree (for parallel agents
on the same repo)
--accept-hooks Auto-approve any unseen shell hooks declared in
config.yaml without a TTY prompt (see also
HERMES_ACCEPT_HOOKS env var and hooks_auto_accept: in
config.yaml).
--checkpoints Enable filesystem checkpoints before destructive file
operations (use /rollback to restore)
--max-turns N Maximum tool-calling iterations per conversation turn
(default: 500, or agent.max_turns in config)
--run-budget SECONDS Optional wall-clock budget in seconds for each
conversation run. At 80% elapsed the agent gets a one-
time wrap-up notice, and implicit provider stale
timeouts are capped to the remaining budget so one
hung call can't consume the run. Unset = off. Also
configurable as agent.run_budget_seconds in
config.yaml. Intended for one-shot/eval invocations
with a hard ceiling.
--yolo Bypass all dangerous command approval prompts (use at
your own risk)
--pass-session-id Include the session ID in the agent's system prompt
--ignore-user-config Ignore ~/.hermes/config.yaml and fall back to built-in
defaults (credentials in .env are still loaded).
Useful for isolated CI runs, reproduction, and third-
party integrations.
--ignore-rules Skip auto-injection of AGENTS.md, SOUL.md,
.cursorrules, memory, and preloaded skills. Combine
with --ignore-user-config for a fully isolated run.
--safe-mode Troubleshooting mode: disable ALL customizations —
user config, AGENTS.md/memory injection, plugins, and
MCP servers (implies --ignore-user-config and
--ignore-rules). Use to isolate whether a problem
comes from your setup or from Hermes itself.
--source SOURCE Session source tag for filtering (default: cli). Use
'tool' for third-party integrations that should not
appear in user session lists.
--tui Launch the modern TUI instead of the classic REPL
--cli Force the classic prompt_toolkit REPL (overrides
display.interface=tui)
--dev With --tui: run TypeScript sources via tsx (skip dist
build)
===== hermes model --help =====
usage: hermes model [-h] [--refresh] [--portal-url PORTAL_URL]
[--inference-url INFERENCE_URL] [--client-id CLIENT_ID]
[--scope SCOPE] [--no-browser] [--timeout TIMEOUT]
[--ca-bundle CA_BUNDLE] [--insecure]
Interactively select your inference provider and default model
options:
-h, --help show this help message and exit
--refresh Wipe the model picker disk cache and re-fetch every
provider's live /v1/models list.
--portal-url PORTAL_URL
Portal base URL for Nous login (default: production
portal)
--inference-url INFERENCE_URL
Inference API base URL for Nous login (default:
production inference API)
--client-id CLIENT_ID
OAuth client id to use for Nous login (default:
hermes-cli)
--scope SCOPE OAuth scope to request for Nous login
--no-browser Do not attempt to open the browser automatically
during Nous login
--timeout TIMEOUT HTTP request timeout in seconds for Nous login
(default: 15)
--ca-bundle CA_BUNDLE
Path to CA bundle PEM file for Nous TLS verification
--insecure Disable TLS verification for Nous login (testing only)
===== hermes config --help =====
usage: hermes config [-h]
{show,edit,get,set,unset,path,env-path,check,migrate} ...
Manage Hermes Agent configuration
positional arguments:
{show,edit,get,set,unset,path,env-path,check,migrate}
show Show current configuration
edit Open config file in editor
get Print a resolved configuration value
set Set a configuration value
unset Remove a configuration value
path Print config file path
env-path Print .env file path
check Check for missing/outdated config
migrate Update config with new options
options:
-h, --help show this help message and exit
===== hermes cron --help =====
usage: hermes cron [-h] [--accept-hooks]
{list,create,add,edit,pause,resume,run,remove,rm,delete,status,runs,history,incidents,notepad,doctor,tick} ...
Manage scheduled tasks
positional arguments:
{list,create,add,edit,pause,resume,run,remove,rm,delete,status,runs,history,incidents,notepad,doctor,tick}
list List scheduled jobs
create (add) Create a scheduled job
edit Edit an existing scheduled job
pause Pause a scheduled job
resume Resume a paused job
run Run a job on the next scheduler tick
remove (rm, delete)
Remove a scheduled job
status Check if cron scheduler is running
runs (history) Show durable execution attempts
incidents List or acknowledge durable cron failure incidents
notepad Read/write a job's durable notepad (persistent KV
across runs)
doctor Check scheduled jobs for common health issues
tick Run due jobs once and exit
options:
-h, --help show this help message and exit
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
(equivalent to HERMES_ACCEPT_HOOKS=1 /
hooks_auto_accept: true).
===== hermes kanban --help =====
usage: hermes kanban [-h] [--board <slug>]
{init,boards,create,swarm,list,ls,show,assign,set-model,reclaim,reassign,diagnostics,diag,link,unlink,claim,comment,attach,attachments,attach-rm,complete,edit,block,schedule,unblock,request-review,request-changes,reopen-review,promote,archive,tail,dispatch,daemon,watch,stats,notify-subscribe,notify-list,notify-unsubscribe,log,runs,heartbeat,assignees,context,specify,decompose,gc,repair} ...
Durable SQLite-backed task board shared across Hermes profiles. Tasks are
claimed atomically, can depend on other tasks, and are executed by a named
profile in an isolated workspace. See https://hermes-
agent.nousresearch.com/docs/user-guide/features/kanban or docs/hermes-
kanban-v1-spec.pdf for the full design.
positional arguments:
{init,boards,create,swarm,list,ls,show,assign,set-model,reclaim,reassign,diagnostics,diag,link,unlink,claim,comment,attach,attachments,attach-rm,complete,edit,block,schedule,unblock,request-review,request-changes,reopen-review,promote,archive,tail,dispatch,daemon,watch,stats,notify-subscribe,notify-list,notify-unsubscribe,log,runs,heartbeat,assignees,context,specify,decompose,gc,repair}
init Create kanban.db if missing (idempotent)
boards Manage kanban boards (one board per project /
workstream)
create Create a new task
swarm Create a Kanban Swarm v1 graph (parallel workers →
verifier → synthesizer)
list (ls) List tasks
show Show a task with comments + events
assign Assign or reassign a task
set-model Set or clear a task's model/provider override (takes
effect on the next dispatch)
reclaim Release an active worker claim on a running task
reassign Reassign a task to a different profile, optionally
reclaiming first
diagnostics (diag) List active diagnostics on the current board
link Add a parent->child dependency
unlink Remove a parent->child dependency
claim Atomically claim a ready task (prints resolved
workspace path)
comment Append a comment
attach Attach a local file to a task
attachments List a task's attachments
attach-rm Delete an attachment by id
complete Mark one or more tasks done
edit Edit recovery fields on an already-completed task
block Mark one or more tasks blocked
schedule Park one or more tasks in Scheduled (waiting on time,
not human input)
unblock Return blocked/scheduled tasks to ready, or todo while
parents remain open
request-review Move a task to 'review' (implementation done, awaiting
review) — NOT a block
request-changes Reviewer verdict: return the active review run to its
implementer
reopen-review Send one or more review tasks back for changes (review
-> ready/todo)
promote Manually move one or more todo/blocked tasks to ready
(recovery path)
archive Archive one or more tasks
tail Follow a task's event stream
dispatch One dispatcher pass: reclaim stale, promote ready,
spawn workers
daemon DEPRECATED — dispatcher now runs in the gateway. Use
`hermes gateway start`.
watch Live-stream task_events to the terminal (Ctrl+C to
exit)
stats Per-status + per-assignee counts + oldest-ready age
notify-subscribe Subscribe a gateway source to a task's terminal events
(used by /kanban subscribe in the gateway adapter)
notify-list List notification subscriptions (optionally for a
single task)
notify-unsubscribe Remove a gateway subscription from a task
log Print the worker log for a task (from <kanban-
root>/kanban/logs/)
runs Show attempt history for a task (one row per run:
profile, outcome, elapsed, summary)
heartbeat Emit a heartbeat event for a running task (worker
liveness signal)
assignees List known profiles + per-profile task counts (union
of ~/.hermes/profiles/ and current assignees on the
board)
context Print the full context a worker sees for a task (title
+ body + parent results + comments).
specify Flesh out a triage-column task into a concrete spec
(title + body) and promote it to todo. Uses the
auxiliary LLM configured under
auxiliary.triage_specifier.
decompose Decompose a triage-column task into a graph of child
tasks routed to specialist profiles by description.
Falls back to specify-style single-task promotion when
the task doesn't benefit from fan-out. Uses
auxiliary.kanban_decomposer.
gc Garbage-collect archived-task workspaces, old events,
and old logs
repair Check kanban.db integrity and auto-repair index-only
corruption
options:
-h, --help show this help message and exit
--board <slug> Board slug to operate on. Defaults to the current
board (set via `hermes kanban boards switch <slug>` or
the HERMES_KANBAN_BOARD env var). Use `hermes kanban
boards list` to see all boards.
===== hermes skills --help =====
usage: hermes skills [-h]
{trust,untrust,browse,search,install,inspect,list,check,update,audit,uninstall,reset,list-modified,diff,opt-out,opt-in,repair-official,publish,snapshot,tap,config} ...
Search, install, inspect, audit, configure, and manage skills from skills.sh,
well-known agent skill endpoints, GitHub, ClawHub, and other registries.
positional arguments:
{trust,untrust,browse,search,install,inspect,list,check,update,audit,uninstall,reset,list-modified,diff,opt-out,opt-in,repair-official,publish,snapshot,tap,config}
trust Trust a project so its repo-local skills
(./.hermes/skills, ./.agents/skills) load
untrust Revoke project-skill trust for a repo
browse Browse all available skills (paginated)
search Search skill registries
install Install a skill
inspect Preview a skill without installing
list List installed skills
check Check installed hub skills for updates
update Update installed hub skills
audit Re-scan installed hub skills
uninstall Remove a hub-installed skill
reset Reset a bundled skill — clears 'user-modified'
tracking so updates work again
list-modified List bundled skills you've edited (which `hermes
update` keeps)
diff Show how your copy of a bundled skill differs from the
stock version
opt-out Stop bundled skills from being seeded into this
profile
opt-in Re-enable bundled-skill seeding (undo opt-out)
repair-official Backfill or restore official optional skills from repo
source
publish Publish a skill to a registry
snapshot Export/import skill configurations
tap Manage skill sources
config Interactive skill configuration — enable/disable
individual skills
options:
-h, --help show this help message and exit
===== hermes sessions --help =====
usage: hermes sessions [-h]
{list,export,delete,prune,archive,optimize,clean-markers,optimize-storage,repair,repair-routing,recover,stats,rename,pin,unpin,pinned,retitle-skills,browse,import} ...
View and manage the SQLite session store
positional arguments:
{list,export,delete,prune,archive,optimize,clean-markers,optimize-storage,repair,repair-routing,recover,stats,rename,pin,unpin,pinned,retitle-skills,browse,import}
list List recent sessions
export Export sessions to JSONL, Markdown, or QMD
delete Delete a specific session
prune Delete old sessions (filterable by time window,
source, title, ...)
archive Bulk-archive (soft-hide) sessions matching filters —
no deletion
optimize Reclaim disk space: merge FTS5 segments + VACUUM (no
data change)
clean-markers Permanently clear stale tool-call marker content left
by sessions from before #78148
optimize-storage Migrate the search index to the compact v23 layout
(reclaims disk on large DBs)
repair Repair a malformed state.db schema so hidden sessions
reappear
repair-routing Re-stamp gateway sessions that lost their routing
identity
recover Rebuild canonical session data into a separate clean
database
stats Show session store statistics
rename Set or change a session's title
pin Pin session(s) — durable keep flag, exempt from auto-
archive
unpin Remove the pin (durable keep flag) from session(s)
pinned List pinned sessions
retitle-skills Re-title sessions whose auto-title came from a
/skill's own text
browse Interactive session picker — browse, search, and
resume sessions
import Import a Claude Code or Codex CLI session into Hermes
options:
-h, --help show this help message and exit
===== hermes mcp --help =====
usage: hermes mcp [-h] [--accept-hooks]
{serve,add,remove,rm,list,ls,test,configure,config,login,reauth,picker,catalog,install} ...
Manage MCP server connections and run Hermes as an MCP server. MCP servers
provide additional tools via the Model Context Protocol. Use 'hermes mcp add'
to connect to a new server, or 'hermes mcp serve' to expose Hermes
conversations over MCP.
positional arguments:
{serve,add,remove,rm,list,ls,test,configure,config,login,reauth,picker,catalog,install}
serve Run Hermes as an MCP server (expose conversations to
other agents)
add Add an MCP server (discovery-first install)
remove (rm) Remove an MCP server
list (ls) List configured MCP servers
test Test MCP server connection
configure (config) Toggle tool selection
login Force re-authentication for an OAuth-based MCP server
reauth Re-authenticate one OAuth MCP server, or all of them
(--all)
picker Interactive catalog picker (also the default for
`hermes mcp`)
catalog List Nous-approved MCPs available for one-click
install
install Install a catalog MCP by name (e.g. `hermes mcp
install n8n`)
options:
-h, --help show this help message and exit
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
(equivalent to HERMES_ACCEPT_HOOKS=1 /
hooks_auto_accept: true).
===== hermes profile --help =====
usage: hermes profile [-h]
{list,use,create,delete,describe,show,alias,rename,export,import,install,update,info} ...
positional arguments:
{list,use,create,delete,describe,show,alias,rename,export,import,install,update,info}
list List all profiles
use Set sticky default profile
create Create a new profile
delete Delete a profile
describe Read or set a profile's description (used by the
kanban orchestrator)
show Show profile details
alias Manage wrapper scripts
rename Rename a profile ('default': sets a display name; id
unchanged)
export Export a profile to archive
import Import a profile from archive
install Install a profile distribution from a git URL or local
directory
update Re-pull a distribution and apply updates (user data
preserved)
info Show a profile's distribution manifest (version,
requirements, source)
options:
-h, --help show this help message and exit
===== hermes memory --help =====
usage: hermes memory [-h] {setup,status,off,reset} ...
Set up and manage external memory provider plugins. Available providers:
honcho, openviking, mem0, hindsight, holographic, retaindb, byterover. Only
one external provider can be active at a time. Built-in memory
(MEMORY.md/USER.md) is always active.
positional arguments:
{setup,status,off,reset}
setup Interactive provider selection and configuration
status Show current memory provider config
off Disable external provider (built-in only)
reset Erase all built-in memory (MEMORY.md and USER.md)
options:
-h, --help show this help message and exit
===== hermes tools --help =====
usage: hermes tools [-h] [--summary] {list,disable,enable,post-setup} ...
Enable, disable, or list tools for CLI, Telegram, Discord, etc. Built-in
toolsets use plain names (e.g. web, memory). MCP tools use server:tool
notation (e.g. github:create_issue). Run 'hermes tools' with no subcommand for
the interactive configuration UI.
positional arguments:
{list,disable,enable,post-setup}
list Show all tools and their enabled/disabled status
disable Disable toolsets or MCP tools
enable Enable toolsets or MCP tools
post-setup Run a provider's post-setup install hook
(npm/pip/binary)
options:
-h, --help show this help message and exit
--summary Print a summary of enabled tools per platform and exit
===== hermes project --help =====
usage: hermes project [-h]
{create,list,ls,show,add-folder,remove-folder,rename,set-primary,use,archive,restore,bind-board} ...
Projects are human-named workspaces that can span multiple folders / repos.
They anchor desktop session grouping and, when bound to a kanban board, give
tasks a deterministic worktree + branch convention. State is per-profile.
positional arguments:
{create,list,ls,show,add-folder,remove-folder,rename,set-primary,use,archive,restore,bind-board}
create Create a new project
list (ls) List projects
show Show a project's details
add-folder Add a folder to a project
remove-folder Remove a folder from a project
rename Rename a project
set-primary Set the primary folder
use Set the active project
archive Archive a project
restore Restore an archived project
bind-board Bind a kanban board to a project
options:
-h, --help show this help message and exit
===== hermes gateway --help =====
usage: hermes gateway [-h] [--accept-hooks]
{run,start,stop,restart,status,install,uninstall,list,setup,migrate-legacy,enroll} ...
Manage the messaging gateway (Telegram, Discord, WhatsApp, Weixin, and more)
positional arguments:
{run,start,stop,restart,status,install,uninstall,list,setup,migrate-legacy,enroll}
run Run gateway in foreground (recommended for WSL,
Docker, Termux)
start Start the installed systemd/launchd background service
stop Stop gateway service
restart Restart gateway service
status Show gateway status
install Install gateway as a systemd/launchd background
service
uninstall Uninstall gateway service
list List all profiles and their gateway status
setup Configure messaging platforms
migrate-legacy Remove legacy hermes.service units from pre-rename
installs
enroll Enroll this gateway with a relay connector (writes
relay auth creds to .env)
options:
-h, --help show this help message and exit
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
(equivalent to HERMES_ACCEPT_HOOKS=1 /
hooks_auto_accept: true).
===== hermes computer-use --help =====
usage: hermes computer-use [-h] {install,status,doctor,permissions} ...
Install or check the cua-driver binary used by the `computer_use` toolset.
Supported on macOS, Windows, and Linux. Use `hermes computer-use install` to
fetch and run the upstream cua-driver installer. This is equivalent to the
post-setup hook that `hermes tools` runs when you first enable the Computer
Use toolset, and is a stable target for re-running the install if it didn't
fire (e.g. when toggling the toolset on a returning-user setup). Use `hermes
computer-use doctor` to run cua-driver's `health_report` MCP tool and surface
its check matrix (TCC, bundle identity, version, platform support, ...) in
human-readable form.
positional arguments:
{install,status,doctor,permissions}
install Install or repair the cua-driver binary
(macOS/Windows/Linux)
status Print whether cua-driver is installed and on PATH
doctor Run cua-driver `health_report` and surface the check
matrix
permissions Check or grant macOS Accessibility + Screen Recording
(macOS)
options:
-h, --help show this help message and exit
===== hermes doctor --help =====
usage: hermes doctor [-h] [--fix] [--live] [--ack ADVISORY_ID]
Diagnose issues with Hermes Agent setup
options:
-h, --help show this help message and exit
--fix Attempt to fix issues automatically
--live Opt-in: run one bounded, read-only real-call health probe
per configured tool backend
(Firecrawl/FAL/browser/MCP/TTS/STT) after the static
checks. Makes real network calls.
--ack ADVISORY_ID Acknowledge a security advisory by ID and exit. After
ack, the advisory will no longer trigger startup banners.
Run `hermes doctor` first to see active advisories and
their IDs.
===== hermes status --help =====
usage: hermes status [-h] [--all] [--deep]
Display status of Hermes Agent components
options:
-h, --help show this help message and exit
--all Show all details (redacted for sharing)
--deep Run deep checks (may take longer)
===== hermes import-agent --help =====
usage: hermes import-agent [-h] [--source SOURCE] [--dry-run] [--overwrite]
[--yes]
[{claude-code,codex}]
One-command import of another coding agent's setup into Hermes. Maps
CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills,
and memories into their Hermes equivalents. Always shows a preview before
making changes. API keys and credentials are never imported — run 'hermes
setup' for those.
positional arguments:
{claude-code,codex} Which agent to import from (default: auto-detect
~/.claude or ~/.codex)
options:
-h, --help show this help message and exit
--source SOURCE Path to the agent's config directory (default:
~/.claude or ~/.codex)
--dry-run Preview only — stop after showing what would be
imported
--overwrite Overwrite existing Hermes items on name conflicts
(default: skip)
--yes, -y Skip confirmation prompts
@@ -0,0 +1,33 @@
import json, subprocess, re
data = json.load(open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json"))
have = {r["videoId"] for r in data}
extra = []
for v in ["8GjyOQy19so", "zwqhemjHq3E"]:
if v in have:
continue
url = f"https://www.youtube.com/watch?v={v}"
oe = subprocess.run(["curl", "-s", f"https://www.youtube.com/oembed?url={url}&format=json"],
capture_output=True, text=True, timeout=30)
try:
oej = json.loads(oe.stdout)
o = {"status": "PASS", "title": oej.get("title"), "author": oej.get("author_name")}
except Exception:
o = {"status": "FAIL", "raw": oe.stdout[:200]}
wp = subprocess.run(["curl", "-s", "-L", url,
"-H", "User-Agent: Mozilla/5.0 (Windows NT 10.0) Chrome/124.0",
"-H", "Accept-Language: en-US"], capture_output=True, text=True, timeout=30)
html = wp.stdout
pub = re.search(r'"publishDate":"([\d\-T:Z]+)"', html)
dur = re.search(r'"lengthSeconds":"(\d+)"', html)
views = re.search(r'"viewCount":"(\d+)"', html)
rec = {"videoId": v, "oembed": o,
"publishDate": pub.group(1) if pub else None,
"lengthSeconds": int(dur.group(1)) if dur else None,
"views": int(views.group(1)) if views else None}
extra.append(rec)
print(json.dumps(rec))
data.extend(extra)
json.dump(data, open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json", "w"), indent=2)
print("total videos in ledger:", len(data))
@@ -0,0 +1,38 @@
import json, subprocess, re, sys
vids = ["9GpWELm3_XI","iqN6MVzpJTk","Ta2wg6xPaY4","8tpuky8HpXw","tP6yf22OJdI",
"UWjh5Z4s8jY","5d02TYoOzfE","83nWNRKZTCE","6M2tItdARew","5_N84t1rUU0",
"UTZhvPXnmwA","1UgXUjT-QtI","6GtF_uHbGhw","CwPUOVUdApE","sCa3BtpkziQ",
"jmtpYUOr7_U","cu2fgknmemA","4sAmpcSOVEw","P2LIFtrRr2U"]
results = []
for v in vids:
url = f"https://www.youtube.com/watch?v={v}"
oe = subprocess.run(["curl","-s",f"https://www.youtube.com/oembed?url={url}&format=json"],
capture_output=True, text=True, timeout=30)
try:
oej = json.loads(oe.stdout)
oembed = {"status":"PASS","title":oej.get("title"),"author":oej.get("author_name")}
except Exception:
oembed = {"status":"FAIL","raw":oe.stdout[:200]}
# watch page for date + duration
wp = subprocess.run(["curl","-s","-L",url,"-H","User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/124.0",
"-H","Accept-Language: en-US,en;q=0.9"], capture_output=True, text=True, timeout=30)
html = wp.stdout
pub = re.search(r'"publishDate":"([\d\-T:Z]+)"', html)
upd = re.search(r'"uploadDate":"([\d\-T:Z]+)"', html)
dur = re.search(r'"lengthSeconds":"(\d+)"', html)
views = re.search(r'"viewCount":"(\d+)"', html)
results.append({
"videoId": v, "oembed": oembed,
"publishDate": pub.group(1) if pub else None,
"uploadDate": upd.group(1) if upd else None,
"lengthSeconds": int(dur.group(1)) if dur else None,
"views": int(views.group(1)) if views else None,
"watchpage_bytes": len(html),
})
print(json.dumps(results[-1]))
with open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json","w") as f:
json.dump(results, f, indent=2)
print("saved video-verification.json")
@@ -0,0 +1,271 @@
[
{
"videoId": "9GpWELm3_XI",
"oembed": {
"status": "PASS",
"title": "Hermes Agent Explained In 5 Minutes",
"author": "CodeHead"
},
"publishDate": "2026-05-23T08:00:27-07:00",
"uploadDate": "2026-05-23T08:00:27-07:00",
"lengthSeconds": 293,
"views": 251275,
"watchpage_bytes": 1419048
},
{
"videoId": "iqN6MVzpJTk",
"oembed": {
"status": "PASS",
"title": "Hermes Agent by Nous Research: The Open-Source Agent Model Everyone Is Switching To",
"author": "Praveen Govindaraj"
},
"publishDate": "2026-02-26T21:56:54-08:00",
"uploadDate": "2026-02-26T21:56:54-08:00",
"lengthSeconds": 193,
"views": 3140,
"watchpage_bytes": 1305120
},
{
"videoId": "Ta2wg6xPaY4",
"oembed": {
"status": "PASS",
"title": "Learn 95% of Hermes Agent in 31 Minutes",
"author": "Sharbel A."
},
"publishDate": "2026-08-09T07:00:19-07:00",
"uploadDate": "2026-08-09T07:00:19-07:00",
"lengthSeconds": 1888,
"views": 129313,
"watchpage_bytes": 1468130
},
{
"videoId": "8tpuky8HpXw",
"oembed": {
"status": "PASS",
"title": "Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model",
"author": "Tonbi's AI Garage"
},
"publishDate": "2026-03-11T07:00:14-07:00",
"uploadDate": "2026-03-11T07:00:14-07:00",
"lengthSeconds": 907,
"views": 18167,
"watchpage_bytes": 1326925
},
{
"videoId": "tP6yf22OJdI",
"oembed": {
"status": "PASS",
"title": "Did Hermes Agent just kill OpenClaw? (full guide)",
"author": "Alex Finn"
},
"publishDate": "2026-03-31T06:15:10-07:00",
"uploadDate": "2026-03-31T06:15:10-07:00",
"lengthSeconds": 835,
"views": 132128,
"watchpage_bytes": 1390199
},
{
"videoId": "UWjh5Z4s8jY",
"oembed": {
"status": "PASS",
"title": "Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra",
"author": "Peter Yang"
},
"publishDate": "2026-08-02T06:00:12-07:00",
"uploadDate": "2026-08-02T06:00:12-07:00",
"lengthSeconds": 2804,
"views": 36613,
"watchpage_bytes": 1386664
},
{
"videoId": "5d02TYoOzfE",
"oembed": {
"status": "PASS",
"title": "Use This To Make The Hermes Agent Basically Free",
"author": "AI LABS"
},
"publishDate": "2026-07-01T07:00:26-07:00",
"uploadDate": "2026-07-01T07:00:26-07:00",
"lengthSeconds": 788,
"views": 54858,
"watchpage_bytes": 1412833
},
{
"videoId": "83nWNRKZTCE",
"oembed": {
"status": "PASS",
"title": "Meet the AI Agent That Grows With You Hermes Agent by Nous Research",
"author": "Eddy Says Hi #EddySaysHi"
},
"publishDate": "2026-03-21T13:00:09-07:00",
"uploadDate": "2026-03-21T13:00:09-07:00",
"lengthSeconds": 366,
"views": 465,
"watchpage_bytes": 1254068
},
{
"videoId": "6M2tItdARew",
"oembed": {
"status": "PASS",
"title": "The AI Agent That Never Forgets: Meet Hermes Agent by Nous Research",
"author": "Siggi"
},
"publishDate": "2026-03-11T13:53:17-07:00",
"uploadDate": "2026-03-11T13:53:17-07:00",
"lengthSeconds": 371,
"views": 1199,
"watchpage_bytes": 1267812
},
{
"videoId": "5_N84t1rUU0",
"oembed": {
"status": "PASS",
"title": "Hermes Agent Fundamentals In 29 Minutes",
"author": "Tina Huang"
},
"publishDate": "2026-07-20T09:37:02-07:00",
"uploadDate": "2026-07-20T09:37:02-07:00",
"lengthSeconds": 1780,
"views": 462764,
"watchpage_bytes": 1561142
},
{
"videoId": "UTZhvPXnmwA",
"oembed": {
"status": "PASS",
"title": "Hermes Agent: Agents that grow with you |Episode #357|",
"author": "Practical AI"
},
"publishDate": "2026-05-20T15:00:17-07:00",
"uploadDate": "2026-05-20T15:00:17-07:00",
"lengthSeconds": 2853,
"views": 1888,
"watchpage_bytes": 1329765
},
{
"videoId": "1UgXUjT-QtI",
"oembed": {
"status": "PASS",
"title": "Hermes Agent: Why Everyone's Ditching OpenClaw in 2026",
"author": "Luke Alexander AI"
},
"publishDate": "2026-03-26T07:10:29-07:00",
"uploadDate": "2026-03-26T07:10:29-07:00",
"lengthSeconds": 1083,
"views": 13625,
"watchpage_bytes": 1321848
},
{
"videoId": "6GtF_uHbGhw",
"oembed": {
"status": "PASS",
"title": "Every Level of Hermes Agent Explained",
"author": "Jack Roberts"
},
"publishDate": "2026-06-17T12:27:42-07:00",
"uploadDate": "2026-06-17T12:27:42-07:00",
"lengthSeconds": 1535,
"views": 162977,
"watchpage_bytes": 1577121
},
{
"videoId": "CwPUOVUdApE",
"oembed": {
"status": "PASS",
"title": "Hermes Agent: The Ultimate Beginner\u2019s Guide",
"author": "Metics Media"
},
"publishDate": "2026-04-24T06:59:04-07:00",
"uploadDate": "2026-04-24T06:59:04-07:00",
"lengthSeconds": 2228,
"views": 118612,
"watchpage_bytes": 1637578
},
{
"videoId": "sCa3BtpkziQ",
"oembed": {
"status": "PASS",
"title": "100 Days With Hermes Agent in 21 Minutes",
"author": "Sharbel A."
},
"publishDate": "2026-06-17T07:47:26-07:00",
"uploadDate": "2026-06-17T07:47:26-07:00",
"lengthSeconds": 1279,
"views": 57581,
"watchpage_bytes": 1454230
},
{
"videoId": "jmtpYUOr7_U",
"oembed": {
"status": "PASS",
"title": "Hermes Agent Just Killed OpenClaw (Full Tutorial)",
"author": "Leon van Zyl"
},
"publishDate": "2026-04-28T04:19:38-07:00",
"uploadDate": "2026-04-28T04:19:38-07:00",
"lengthSeconds": 1199,
"views": 15988,
"watchpage_bytes": 1510488
},
{
"videoId": "cu2fgknmemA",
"oembed": {
"status": "PASS",
"title": "Hermes Agent The 24/7 Self-Evolving AI Agent!",
"author": "WorldofAI"
},
"publishDate": "2026-04-07T00:01:34-07:00",
"uploadDate": "2026-04-07T00:01:34-07:00",
"lengthSeconds": 555,
"views": 46889,
"watchpage_bytes": 1545788
},
{
"videoId": "4sAmpcSOVEw",
"oembed": {
"status": "PASS",
"title": "Hermes Agent - Crash Course for Beginners (AI Agent)",
"author": "Adrian Twarog"
},
"publishDate": "2026-07-21T01:20:41-07:00",
"uploadDate": "2026-07-21T01:20:41-07:00",
"lengthSeconds": 1338,
"views": 46484,
"watchpage_bytes": 1585548
},
{
"videoId": "P2LIFtrRr2U",
"oembed": {
"status": "PASS",
"title": "Hermes Agent: New FREE OpenClaw Alternative!",
"author": "Julian Goldie SEO"
},
"publishDate": "2026-03-09T14:00:32-07:00",
"uploadDate": "2026-03-09T14:00:32-07:00",
"lengthSeconds": 735,
"views": 9975,
"watchpage_bytes": 1353286
},
{
"videoId": "8GjyOQy19so",
"oembed": {
"status": "PASS",
"title": "Hermes Agent Full Tutorial INSTALLATION + USECASES",
"author": "CodeHead"
},
"publishDate": "2026-05-14T08:00:23-07:00",
"lengthSeconds": 467,
"views": 64285
},
{
"videoId": "zwqhemjHq3E",
"oembed": {
"status": "PASS",
"title": "Hermes Agent vs OpenClaw",
"author": "Sharbel A."
},
"publishDate": "2026-04-20T07:14:00-07:00",
"lengthSeconds": 928,
"views": 35360
}
]
@@ -0,0 +1,10 @@
import re, sys
html = open(sys.argv[1], encoding='utf-8', errors='ignore').read()
ids = re.findall(r'"videoRenderer":\{"videoId":"([\w-]{11})"', html)
print("videoRenderer hits:", len(set(ids)))
for vid in dict.fromkeys(ids):
m = re.search(r'"videoId":"%s".{0,3000}?"title":\{"runs":\[\{"text":"(.*?)"\}' % vid, html, re.S)
ch = re.search(r'"videoId":"%s".{0,6000}?"ownerText":\{"runs":\[\{"text":"(.*?)"' % vid, html, re.S)
dur = re.search(r'"videoId":"%s".{0,4000}?"lengthText":\{"accessibility".{0,400}?"simpleText":"(.*?)"' % vid, html, re.S)
print((vid, m.group(1) if m else "?", ch.group(1) if ch else "?", dur.group(1) if dur else "?"))
@@ -0,0 +1,42 @@
# 06 — Sources
Access date for ALL entries: **2026-09-11** (via citation ledger `sources.py`; doc URLs additionally confirmed HTTP 200 by curl -L).
## Official docs (hermes-agent.nousresearch.com)
| # | URL | Supported |
|---|-----|-----------|
| 1 | https://hermes-agent.nousresearch.com/docs | Docs index; overall feature map |
| 3 | https://hermes-agent.nousresearch.com/docs/user-guide/configuration | Config sections, SOUL.md, checkpoints |
| 4 | https://hermes-agent.nousresearch.com/docs/reference/slash-commands | Slash command registry (03) |
| 5 | https://hermes-agent.nousresearch.com/docs/reference/tools-reference | Toolset inventory (02 §6) |
| 6 | https://hermes-agent.nousresearch.com/docs/user-guide/features/cron | Cron surface (02 §9) |
| 7 | https://hermes-agent.nousresearch.com/docs/user-guide/features/kanban | Kanban surface (02 §8) |
| 8 | https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp | MCP surface + `hermes mcp serve` (02 §5, 04 P3) |
| 9 | https://hermes-agent.nousresearch.com/docs/user-guide/features/memory | Memory surface (02 §2) |
| 10 | https://hermes-agent.nousresearch.com/docs/user-guide/profiles | Profiles (02 §13) |
| 11 | https://hermes-agent.nousresearch.com/docs/integrations/providers | Model/provider routing (02 §12) |
| 12 | https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ | Gateway platforms (02 §15) |
| 13 | https://hermes-agent.nousresearch.com/docs/user-guide/features/curator | Skill maintenance (02 §3) |
| 14 | https://hermes-agent.nousresearch.com/docs/reference/cli-commands | CLI command index cross-check (03) |
| 15 | https://hermes-agent.nousresearch.com/docs/reference/skills-catalog | Skills catalog (02 §3) |
## GitHub
| # | URL | Supported |
|---|-----|-----------|
| 2 | https://github.com/nousresearch/hermes-agent | Repo identity, learning-loop description (01, 02) |
## Local primary sources (not web URLs; verified on this host)
- Live CLI help output, Hermes Agent v0.21.1 (2026.9.7), reference install (Syslog kagentz): `hermes --help` + `hermes {chat,model,config,cron,kanban,skills,sessions,mcp,profile,memory,tools,project,gateway,computer-use,doctor,status,import-agent} --help` → raw dump `cli-help-dump.txt` (577 lines). Basis for all VERIFIED-LIVE tags in 01/02/03/04.
- `/home/hermes/.hermes/skills/autonomous-ai-agents/delegate-coding-agent/SKILL.md` (v1.0.0) + `references/claude-code.md` (v2.2.1) → 04 Patterns 1-2, 01 Claude Code column.
- `/home/hermes/.hermes/skills/autonomous-ai-agents/merge-reconciler/SKILL.md` → 04 Pattern 2.
- `/home/hermes/.hermes/skills/autonomous-ai-agents/computer-use/SKILL.md` (v2.0.0) → 04 Pattern 5, 02 §11.
## YouTube verification
- YouTube search results pages (scraped 2026-09-11): `https://www.youtube.com/results?search_query=hermes+agent+nous+research` and `...nous+research+hermes+agent+official`
- oEmbed endpoint per video: `https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=<ID>&format=json` — all 21 checked IDs returned PASS (HTTP 200, title/author match). Raw evidence incl. publishDate/lengthSeconds/viewCount per watch page: `video-verification.json`.
- 17 listed in 05-videos.md + 4 spares; 21/21 pass, 0 fail.
## Explicit gaps (could not close)
1. No official Nous Research-produced tutorial video was found — video list is third-party ecosystem content (disclosed in 05).
2. Slash commands were verified DOC-ONLY (https://hermes-agent.nousresearch.com/docs/reference/slash-commands); they require an interactive session to exercise, which this headless run does not have. CLI equivalents were verified live.
3. `hermes-agent.nousresearch.com/docs/developer-guide/` returned 404 — developer docs live in-repo (`AGENTS.md` in the GitHub repo), not as a docs site section.
+2 -67
View File
@@ -43,7 +43,7 @@ Docker hosts get special attention:
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
## Threat Levels (GUEST filesystems)
## Threat Levels
| Level | Threshold | Response | Escalation |
|-------|-----------|----------|------------|
@@ -53,39 +53,6 @@ Docker hosts get special attention:
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
## Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)
Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are **report-only** — no automatic deletion of media or datastore content ever.
| Level | Threshold | Response | Escalation |
|-------|-----------|----------|------------|
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
| **HOST-RED** | 95% | Name the volume + % + absolute free space; **media volumes: "capacity decision — owner to decide"; pbs-datastore/host-root: "immediate owner attention"** | Zulip DM + channel alert (state-change only) |
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
**State lives in a small JSON state file:** the scanner resolves it to an absolute path from the script's location — `$(dirname "$0")/state/host-disk-bands.json` (i.e., the `state/` directory next to the `scripts/` directory in the prose-contracts repo). The state file is gitignored runtime state — the scanner creates it on first run. Keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
```
storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN
```
**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.
**Action classes by volume type:**
- **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
- **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
- **pbs-datastore** (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.
**Justification (measured 2026-09-15, firstmate):** storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.
## Requires
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
@@ -160,17 +127,6 @@ from the `report_only_guests` YAML block above.
- `retry`: 2 attempts for SSH failures before marking a CT unreachable
## Execution
**Execution model**: This contract is executed by a host-scheduled cron job (see
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
### Host filesystems: report-only, NEVER auto-delete
Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
@@ -235,21 +191,6 @@ call summary-reporter
plan: plan
```
## GC SCHEDULE (PBS datastore only)
The PBS GC schedule is defined in ONE authoritative place: `/etc/cron.d/pbs-gc` on storepve.
The schedule is `0 20 * * *` (20:00 LOCAL = 00:00 UTC, since host timezone is America/New_York).
This applies **only** to the PBS datastore (`/tank/pbs-backup`), NOT to media volumes.
Media volumes (/media/*) are report-only at all threat levels.
The cron runs `/usr/local/bin/pbs-gc.sh` which executes:
```bash
proxmox-backup-manager garbage-collection start storepve-datastore
```
This is NOT a `prune` operation; it is a GC pass that reclaims unreferenced chunks.
There is no `--keep-daily` flag; retention is governed by jobs.cfg (keep-daily=35).
The GC does not touch media volumes or any other filesystem.
## GC Strategies by Host Type
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
@@ -302,12 +243,6 @@ done
## Alert Templates
### HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)
```
⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
Action: {volume_type-specific action}
```
### AMBER (75-84%)
```
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
@@ -414,7 +349,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| 108 | media | storepve | lxc | ✅ reachable |
| 110 | gitea | minipve | lxc | ✅ reachable |
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
| 112 | tanko | minipve | lxc | ✅ reachable |
| 112 | tanko | amdpve | lxc | ✅ reachable |
| 113 | baggy | amdpve | lxc | ✅ reachable |
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
| 116 | syslog-api | minipve | lxc | ✅ reachable |
-169
View File
@@ -1,169 +0,0 @@
# Contract execution pinning
Which copy of a contract script actually ran, and how that is proven.
## Why this exists
Three times on 2026-09-25 a contract reported a verdict from a copy that was
not the merged one:
1. The ops lane's own clone sat on the merged feature branch
`fix/search-stack-multi-engine-20260925` at `8b2eba4` with no `pve_auth`
fix, while it executed the daily digest from a different clone. Nothing in
the workflow noticed.
2. `scripts/search-stack-check.py` was deployed into the pinned runner clone
by hand rather than through git.
3. A stale local `origin/master` ref made an ancestry check report
"unlanded work" for a branch that had in fact merged — the same staleness
would have passed a stale script as current.
A contract verdict is only meaningful if it came from the merged copy. The
control is `scripts/revision-preflight.sh`.
## The rule
**Every contract pins exactly one clone for execution: the clone that
`scripts/contract-run.sh` itself lives in.**
`contract-run.sh` derives that from its own location (`SCRIPTS_DIR`) and checks
the script it is about to run against `origin/master` in the same clone. There
is no second path to configure, and no contract may be executed from a
hand-copied location.
| Contract | Script | Pinned clone |
| --- | --- | --- |
| `infrastructure-monitoring` | `scripts/infra-monitoring.sh` | the clone containing `contract-run.sh` |
| `proxmox-monitor` | `scripts/proxmox-monitor.sh` | same |
| `zulip-health` | `scripts/zulip-monitor.sh` | same |
| `agent-health-check` | `scripts/agent-health-check.py` | same |
| `litellm-health` | `scripts/litellm-health-check.py` | same |
| `disk-gc-threat-response` | `scripts/disk-gc-scan.py` | same |
| `pm2-self-heal` | `scripts/pm2-self-heal.sh` | same |
| `search-stack-visibility` | `scripts/search-stack-check.py` | same |
### The deployed runner
The scheduler on **CT 100 (abiba)** runs contracts from
**`/opt/contract-runner`** via `/etc/cron.d/contract-runner`. That clone is the
pinned execution copy for every scheduled contract, and it must be kept current
with `master` by fast-forward. Its `origin` is a local path to the upstream
working copy, not a network remote.
`daily-health-digest` is **not** in the table above because it has no contract
file and no mapping — it is dispatched by cron as
`fm-send.sh ops "run contract: daily-health-digest"` and was, until
2026-09-25, executed by hand from whichever clone the operator happened to be
in. Creating its contract file and pinning it to a clone is an open follow-up.
## How the check works
`scripts/revision-preflight.sh <script-path> <clone-path>`:
* resolves the **repo-relative** path of the executing script inside the clone;
* **fetches** the remote first, so a stale local ref cannot make a stale script
look current — bounded by `--fetch-timeout` (default 20s) so a hung remote
cannot block a scheduled contract;
* compares the script's sha256 against `<ref>:<repo-relative-path>`;
* **fails closed** — a path absent from the ref, an unresolvable ref, or a
failed fetch is a failure, never a warning.
### Exit codes and reason classes
The guard distinguishes **"I could not check"** from **"this copy is wrong"**,
and every non-zero exit prints a machine-readable `REASON=<class>` line before
the human text, because a warning nobody can classify is not actionable — and
the flip to `enforce` (below) depends on being able to read these apart.
| exit | `REASON=` | meaning |
| --- | --- | --- |
| 0 | — | verified match |
| 2 | `cannot-verify:fetch-failed` | remote unreachable, failed, or timed out |
| 2 | `cannot-verify:ref-unresolvable` | `<ref>` does not exist in the clone |
| 1 | `mismatch:path-absent` | the script does not exist in `<ref>` |
| 1 | `mismatch:content` | the script differs from `<ref>` |
| 1 | `mismatch:detached-head` | the clone is on a detached HEAD |
| 1 | `mismatch:clone-ahead` | local HEAD is strictly ahead of `<ref>` (mid-review) |
`detached-head` and `clone-ahead` are named separately on purpose: they are
*legitimate* states that merely fail to be "the merged copy", and they are far
less alarming than a hand-edited file. `clone-ahead` requires HEAD to be
**strictly** ahead — an uncommitted edit on a commit that *is* the ref is a
plain `content` mismatch.
## Modes in `contract-run.sh`
| `CONTRACT_REVISION_PREFLIGHT` | Behaviour |
| --- | --- |
| unset / **`warn` (default)** | log the refusal and its class, then still report |
| `enforce` | withhold the verdict, alert, exit `2` |
| `off` | skip the check entirely |
**The default is `warn`, deliberately.** The guard gates *every* scheduled
contract, and three legitimate situations would otherwise turn the whole
fleet's monitoring into withheld verdicts: a clone legitimately ahead of
`origin/master` mid-review, a detached HEAD, and an offline or failed fetch.
That is a bigger risk than the staleness the guard exists to catch. `warn`
keeps the signal loud and classified in every run's log without letting the
monitoring go dark.
### Criteria for flipping the default to `enforce`
Do not flip it on preference. Flip it when the evidence says the false-refusal
rate is low enough, as its own small change with its own review:
1. the guard has run across **every scheduled contract** for a sustained period
(suggested: 30 consecutive days, or 200+ contract runs) with **zero**
`mismatch:*` and **zero** `cannot-verify:*` refusals in the per-run logs;
2. no `cannot-verify:fetch-failed` arising from ordinary network blips in that
window — if the pinned clone's remote is not reliably reachable, `enforce`
will withhold rather than report;
3. the pinned runner clone is demonstrably kept current by fast-forward, so
`mismatch:clone-ahead` is a genuine fault rather than routine procedure.
The evidence for the flip is the `REASON=` lines already written into
`/var/log/contract-runs/`. Until then the default stays `warn`.
## Merge-time sequence (do this whenever this repo merges)
**Baseline as of 2026-09-25:** `/opt/contract-runner` is already
fast-forwarded to master `9faffe4`, so the pinned runner clone is current
today. This sequence exists to keep it that way.
After any merge to `master`:
```bash
# 1. fast-forward the pinned runner clone on CT 100
git -C /opt/contract-runner pull --ff-only
# 2. confirm it is current and clean
git -C /opt/contract-runner log --oneline -1
git -C /opt/contract-runner status --porcelain # expect no output
# 3. prove a contract runs and reports normally
CONTRACT_RUN_LOG_DIR=/tmp/preflight-proof \
bash /opt/contract-runner/scripts/contract-run.sh search-stack-visibility
echo "EXIT=$?" # expect 0, and 'revision-preflight: … matches origin/master'
```
A contract that reports a `REASON=mismatch:*` refusal here means the runner
clone is stale or locally edited — fast-forward it rather than reaching for
`CONTRACT_REVISION_PREFLIGHT=off`.
**Note on untracked files:** git refuses to fast-forward over an untracked file
even when its content is byte-identical to the incoming version
(`The following untracked working tree files would be overwritten by merge`).
A dirty clone will therefore block step 1. Resolve it by removing or stashing
the untracked paths first — that is exactly what blocked a clone on 2026-09-25.
## Operating notes
* Under the default `warn`, a stale pinned clone still produces verdicts but
every run logs the refusal and its class. Read those lines; do not ignore
them.
* Under `enforce`, a stale pinned clone **withholds**. That is the intended
failure. Recover by fast-forwarding:
`git -C /opt/contract-runner pull --ff-only`.
* When a contract legitimately changes, land it through the normal branch + PR
path and fast-forward the pinned clone. Do not copy files into it by hand.
* `--no-fetch` exists for offline inspection; it prints that freshness is
assumed rather than verified, and it is not used by `contract-run.sh`.
-9
View File
@@ -1,14 +1,5 @@
# Probe-drift round 2 — per-leg before/after evidence
> **Historical record** — 2026-09-28: The lines below that describe tanko as
> "DSH (DeepSeek Harness)" only reflect what the check reported when it was
> running. Tanko's runtime was later found to be **hybrid (DSH + Hermes)** —
> the check had a `/root/` hardcoding bug that made it probe the wrong home
> directory and report `wrapper-missing:tanko` for an agent with a working
> wrapper. This document records the observed output, not the underlying
> truth; see `fix/agent-health-root-hardcoding-20260928` for the correction.
**Date:** 2026-09-10
**Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`
**Branch:** `fm/probe-drift-round2-20260909`
-23
View File
@@ -174,29 +174,6 @@ Key notes:
## Execution
### Executor, schedule and dead-man's-switch
This contract is implemented by a real script, scheduled on the inference host:
| | |
| --- | --- |
| **Executor** | `/opt/inference-harness/scripts/gpu-self-heal.py` on **CT 116** |
| **Schedule** | `/etc/cron.d/gpu-self-heal` on CT 116 — `2 */6 * * *` |
| **Log** | `/var/log/litellm/gpu-self-heal.log` |
| **Posting** | the script calls `gitea-logger.sh gpu {RUN_ID}.json <report>` → `SyslogSolution/health-logs/gpu/{RUN_ID}.json` |
**Dead-man's-switch:** absence of logs must raise an alarm, because that is how
this went silent for 12 days. That alarm cannot live on the producer — a stopped
job cannot report that it stopped — so it lives off-host as the
**`health-log-freshness` contract on CT 100**, which fails when
`health-logs/gpu/` is older than 12 h. A failed run is also visible in the log
above, but only the off-host check catches a *missing* run.
**History (2026-09-26, relay-785):** the executor was never lost — only its
schedule was, dropped during a CT 116 `/etc/cron.d` rework on 2026-09-21. The
`gpu/` log was silent from `2026-09-14T18:02:03Z`. The schedule was restored and
the off-host freshness check added; do not treat either as optional.
```prose
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
-80
View File
@@ -1,80 +0,0 @@
---
kind: function
name: health-log-freshness
description: >
Dead-man's-switch for the health-logs posting jobs. Fails when the newest
commit in a watched SyslogSolution/health-logs directory is older than that
directory's threshold.
Exists because absence of logs raises no alarm: the gpu/ directory went silent
for 12 days (2026-09-14T18:02:03Z to 2026-09-26) and nothing noticed, because
the only thing that would have noticed was the job that had stopped. The check
therefore runs on CT 100, a DIFFERENT host from the producers on CT 116, so a
dead producer host or a deleted schedule still raises an alarm.
Watched: gpu/ (12h, producer 2 */6 * * *), litellm/ (18h, producer 0 */6 * * *).
Not watched: pm2/ — that contract's health-logs claim was retired 2026-09-26;
see pm2-self-heal.prose.md step 6.
Verified 2026-09-26: would have caught the real gap (at 2026-09-20 the newest
gpu/ entry was 126h old against a 12h limit).
version: 1.0.0
---
## Execution
Host-scheduled on CT 100 via `/etc/cron.d/contract-runner`:
```
20 */4 * * * root CONTRACT_RUN_LOG_DIR=/var/log/contract-runs /bin/bash /opt/contract-runner/scripts/contract-run.sh health-log-freshness >/dev/null 2>&1 || /root/abiba-workspace/bin/fm-inbox.sh note "contract-runner: health-log-freshness FAILED - see /var/log/contract-runs/" >/dev/null 2>&1
```
Every 4 hours, offset to `:20` to avoid the existing `:05`/`:15`/`:35` slots.
## Output shape
```
Health-log freshness — dead-man's-switch
==================================================================
✅ health-logs/gpu/ newest 2026-09-26T14:48:03Z (0.02h old, limit 12.0h)
last commit: gpu: gpu-self-heal-20260926-144802.json
producer: gpu-self-heal.py, CT116 cron 2 */6 * * *
==================================================================
VERDICT: PASS — every watched health-log directory is advancing
```
`--json` emits `{checked: {...}, failures: [...]}`.
## Exit codes
| exit | meaning |
| --- | --- |
| 0 | every watched directory is advancing |
| 1 | at least one is stale, or could not be read |
| 2 | the check could not run (no Gitea credential) |
A directory that **cannot be read** is a failure, not a skip: unreadable and
stopped are indistinguishable from the outside.
## Configuration
| variable | default | meaning |
| --- | --- | --- |
| `GITEA_URL` | `https://git.sysloggh.net` | Gitea base URL |
| `GITEA_TOKEN` / `GITEA_PAT` | — | API token; falls back to basic auth from `~/.git-credentials` |
| `HEALTH_LOG_MAX_AGE_GPU` | `12` | hours |
| `HEALTH_LOG_MAX_AGE_LITELLM` | `18` | hours |
## Thresholds
Sized for the producer cadence plus one missed run, so a single blip does not
page but a genuine stop does:
* `gpu/` — 6 h cadence, 12 h limit;
* `litellm/` — 6 h cadence, 18 h limit (proven healthy; a looser bound avoids noise).
## Maintains
- health-logs-gpu-freshness: { status: "ok|stale", last_check: timestamp }
- health-logs-litellm-freshness: { status: "ok|stale", last_check: timestamp }
+3 -3
View File
@@ -70,10 +70,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
## Config Pattern — Mandatory Fields
### For Hermes Agents (Mumuni, Koonimo, Tanko-hybrid)
### For Hermes Agents (Mumuni, Koonimo)
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
(Tanko is hybrid — runs both DSH and Hermes since 2026-08-27, so its Hermes config is also checked.)
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
### 1. Main Model
```yaml
@@ -297,7 +297,7 @@ Run the consolidated health check:
```bash
python3 /root/scripts/agent-health-check.py
```
This validates each agent's live LiteLLM key against the gateway, including tanko, which runs HYBRID (DSH + Hermes) since 2026-08-27; detects GPU port conflicts (ghost processes),
This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes),
verifies gateway liveness, confirms Zulip streaming (`edit_message` present),
and counts recent errors. Non-disruptive — never restarts anything.
+14 -77
View File
@@ -5,13 +5,7 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-09-27: Clarified the Auxiliary Tasks policy — light aux (vision,
web_extract/browsing) -> gpu-vision (RTX 5070); context-heavy aux (compression) ->
syslog-auto (2026-07-23 decision, Rule 7). Removed the false "one model for all
auxiliary" / "never syslog-auto" claim; stated gpu-dense + strix-moe are the reasoning
hosts and aux should not be pinned to them. Now matches audit-hermes-config.py line-for-line.
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
@@ -151,12 +145,6 @@ mcp_servers:
url: http://192.168.68.65:3100/mcp
timeout: 120
connect_timeout: 60
litellm:
url: https://litellm.sysloggh.net/mcp
headers:
x-litellm-api-key: "Bearer <AGENT_KEY>" # Rule 15: must be a REAL key (sk-...), not an env-var name
# Note: MCP endpoint requires Accept: application/json, text/event-stream header
# This is handled by the MCP client library; don't add to config
# ─── Compression ───
compression:
@@ -172,16 +160,13 @@ compression:
abort_on_summary_failure: false
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# Auxiliary tasks split into TWO model classes — do NOT assume one model for all:
# Light auxiliary (vision, web_extract/browsing) -> model: gpu-vision # RTX 5070
# Keeps the reasoning hosts (gpu-dense / strix-moe) free for agent prompts.
# Context-heavy auxiliary (compression) -> model: syslog-auto # weighted pool
# Deliberate per the 2026-07-23 OPERATIONAL DECISION in Rule 7: summarization
# runs against long histories and must be able to use the pool.
# Do NOT pin auxiliary work to the reasoning hosts (gpu-dense / strix-moe).
# All auxiliary services share identical ROUTING (base_url + api_key_env), not model:
# All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gpu-vision # stable alias (NOT a raw model name)
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
# api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
@@ -232,40 +217,6 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
3. **After update**: Restart Hermes on the agent host
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
## MCP Server Configuration
MCP server entries in `mcp_servers:` must follow the format shown in the Template section.
**Header requirements (Rule 15):**
- Use `headers:` field with a `x-litellm-api-key` entry
- The value must be `"Bearer <REAL_KEY>"` where `<REAL_KEY>` is a literal LiteLLM virtual key
- Do NOT use env-var references like `$LITELLM_API_KEY` — they resolve to empty strings in
the static config and cause "Malformed API Key" errors (2026-08-07 Tanko incident)
**Key source:**
- Keys are stored in the Infisical vault (project=agents, env=production)
- For template-based config generation: substitute the agent's key from the agent_keys table
- For manual config updates: retrieve the key from the vault and insert the literal value
**Verification (2026-08-07):**
- Tested MCP initialize handshake against litellm.sysloggh.net/mcp with agent virtual key
- Confirmed: 200 response with `serverInfo.name: "litellm-mcp-server"`
- Confirmed: tools/list returns 200 (MCP endpoint accessible with virtual keys)
- Note: Per-key MCP grants are now supported (verified 2026-09-18), resolving the earlier
contradiction with infrastructure-update.prose.md (which now reflects the update)
- Key requirement: must be a valid LiteLLM virtual key (HTTP 200 on /v1/models)
**Key rotation note:**
- MCP headers use literal keys (not env-vars), so they do NOT auto-rotate with the vault
- After key rotation, MCP server headers must be regenerated with the new key value
- This is a manual step: update the `x-litellm-api-key` header in each config file
- TODO: Consider adding MCP header regeneration to the Key Update Procedure or a generation hook
**NetBird dependency:**
- `litellm.sysloggh.net` is a NetBird endpoint (see Rule 5 and Infrastructure Stack table)
- NetBird outages cause 502 errors on MCP requests, not auth failures
- Diagnose: if MCP requests fail with 502, check NetBird status before investigating keys
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
@@ -504,28 +455,14 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
before and after any config change to catch this and all other rule violations.
### Rule 15: MCP Endpoint and Header Validation (UPDATED 2026-08-07)
**Endpoint validation:**
- ra-h-os must point to `http://192.168.68.65:3100/mcp`
- litellm must point to `https://litellm.sysloggh.net/mcp`
- Mismatched endpoints cause silent failures (e.g., 2026-08-07 incident: Tanko's config had
ra-h-os pointing to litellm's endpoint)
**Header validation:**
- Every MCP entry with authentication must carry a `headers:` field
- The header value must be a REAL key (e.g., `Bearer sk-abc123...`), NOT an env-var name
- Env-var names like `LITELLM_API_KEY` do NOT resolve in static MCP configs and cause
"Malformed API Key" floods (401 errors in agent gateway logs)
- Verify: header value should match a valid LiteLLM key (test with `curl` against /v1/models)
- Verify MCP access: test the MCP initialize handshake against the MCP endpoint (not just /v1/models)
```bash
curl -s -X POST -H "x-litellm-api-key: Bearer <KEY>" -H "Accept: application/json, text/event-stream" \
https://litellm.sysloggh.net/mcp -d '{"jsonrpc":"2.0","id":1,"method":"initialize",...}' \
| jq '.data.result.serverInfo' # should show serverInfo.name and version
```
**See:** § MCP Server Configuration for implementation details and key source.
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
- Every MCP server entry must point at the correct endpoint:
- ra-h-os = http://192.168.68.65:3100/mcp
- litellm = https://litellm.sysloggh.net/mcp
- MCP entries must carry a REAL key value in the header.
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
endpoints and result in "Malformed API Key" floods.
- Ensure the header value is the actual key (e.g., `sk-...`).
## Execution
+7 -36
View File
@@ -16,26 +16,7 @@ author: Abiba (pi agent)
## Rule (One Sentence)
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
## Model Access Tiers (2026-09-20)
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
| Tier | Models | Who gets it |
|------|--------|-------------|
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
**Rules:**
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
2. Agent keys MUST carry an **explicit local-only** `models` list.
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
4. The master key always bypasses scoping — it is admin-only, never for inference.
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with authenticated path `http://192.168.68.116/litellm/v1/responses` — hardcoded keys AND unauthenticated `/v1` direct access are both forbidden.**
## Scope
@@ -57,8 +38,8 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` non-canonical but working).
The public host `https://litellm.sysloggh.net` serves `/v1` ONLY (404 on `/litellm/v1`).
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid).
Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**.
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
@@ -120,7 +101,7 @@ auxiliary:
fallback_providers:
- provider: deepseek
base_url: https://api.deepseek.com
api_key: sk-synthetic-external-example # ← hardcoded OK (external, synthetic example)
api_key: sk-b7d9... # ← hardcoded OK (external)
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
```
@@ -128,11 +109,11 @@ fallback_providers:
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
model:
provider: harness
api_key: sk-synthetic-example-12345 # ← RULE VIOLATION: hardcoded key (synthetic example)
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION: hardcoded key
model:
provider: harness
base_url: http://192.168.68.116/v1 # ← NON-CANONICAL but WORKING (authenticated via nginx, WARN not FAIL)
base_url: http://192.168.68.116/v1 # ← RULE VIOLATION: unauthenticated path
api_key_env: LITELLM_API_KEY
```
@@ -158,16 +139,6 @@ scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
When reporting findings, separate POLICY observations from FAULT findings:
### ACCEPTABLE PATTERN
Agent keys live in `.env` or `.env.vault` files with 600 permissions (koonimo's shape is the canonical example). A plaintext key inside a `config.yaml` or any `config.yaml.bak-*` file is a violation — the backup files are not part of the runtime credential path and are not watched by the scanner, so a key in them is stale clutter that a future reader can mistake for a working key.
**Fix procedure** (when a backup file is found with a plaintext key):
1. Move the file out of the scanned tree (e.g., `mv /root/.hermes/config.yaml.bak-* /root/hermes-config-backups/`) — do NOT delete the file, just move it so the scanner pattern no longer matches.
2. Re-run the reachability check to confirm COMPLIANT.
3. Report the before/after check output and the commands you ran.
**Rationale**: Moving the file preserves history without leaving a credential where a scanner trips over it. Deleting the file loses the historical context. Keeping it in place means the next scan will report it as a finding and waste time re-deciding.
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
@@ -201,7 +172,7 @@ grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-…' /root/.config/systemd/ 2>/dev/null
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /root/.config/systemd/ 2>/dev/null
# 3. Verify running process env matches dedicated key
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
+4 -4
View File
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
| Param | Type | Required | Default | Description |
|-------|------|----------|---------|-------------|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — hybrid (DSH + Hermes) since 2026-08-27, no Hermes plugin) |
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) |
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
## Maintains
@@ -55,7 +55,7 @@ connectivity recovery including end-to-end DM validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------|
| Tanko | CT112 | minipve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(hybrid (DSH + Hermes) since 2026-08-27 — historical, plugin retired on this host)* |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
@@ -72,7 +72,7 @@ connectivity recovery including end-to-end DM validation.
### Step 1: Resolve Target
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
For CT112 route through `ssh root@minipve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
### Step 2: Pull Latest Plugin Source
@@ -121,7 +121,7 @@ cp plugins/platforms/zulip/adapter.py \
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
# Fix ownership (was Tanko-only, runs as jerome user)
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (hybrid: DSH + Hermes).
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
+3 -3
View File
@@ -24,7 +24,7 @@ gateway restart, and connection validation.
| Param | Type | Required | Default | Description |
|-------|------|----------|---------|-------------|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — hybrid (DSH + Hermes) since 2026-08-27) |
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) |
## Maintains
@@ -67,7 +67,7 @@ gateway restart, and connection validation.
### Step 1: Locate Target
Map `target` to connectivity parameters from the live-state table above.
For CT112 route through `ssh root@minipve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
### Step 2: Deploy Zulip Adapter
@@ -93,7 +93,7 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on hybrid (DSH + Hermes), no Hermes plugin)
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin)
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
# Clean up
+8 -81
View File
@@ -105,8 +105,8 @@ description: >
| Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | abiba, tanko, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | kagentz, baggy, scottdenya, adguard2 | Agents, compute |
| minipve | .12 | 16C | 30GB | abiba, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | kagentz, tanko, baggy, scottdenya, adguard2 | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
@@ -181,8 +181,8 @@ description: >
**Stirling-PDF** (deployed 2026-07-03, Authentik SSO 2026-07-03):
- URL: `https://pdf.sysloggh.net` (public) / `http://192.168.68.7:8989` (direct)
- Swagger: `http://192.168.68.7:8989/swagger-ui.html`
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_USER»` / `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
- API key: `«vault: infrastructure/production STIRLING_API_KEY»`
- Admin credentials: `admin` / `kakashi20stirling`
- API key: `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88`
- Authentik OAuth2: configured but disabled (requires paid Server license). Ready to enable: set `SECURITY_OAUTH2_ENABLED=true` + `SECURITY_LOGINMETHOD=all`
- Compose: `/opt/home_stack/docker-compose.yml`
- Control script: `/opt/home_stack/infra-control.sh`
@@ -364,79 +364,6 @@ For docker-vm specifically:
- No PBS backup in 48h → fail
```
### Backup Safety Preconditions (2026-09-15)
#### Background & Rationale
Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first:
1. **acerpve thin-pool VM 101** (acerpve, 192.168.68.9, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root.
2. **amdpve 0700 tmpdir** (amdpve, 192.168.68.15, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup.
3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood.
> ⚠️ **Hostname Resolution Warning (2026-09-15)**: The PVE node hostnames (acerpve, amdpve, minipve, storepve, ocupve) all resolve to the VPS (72.61.0.17, the Netbird VPS at srv1079750.hstgr.cloud) via the wildcard `*.dns.sysloggh.net` record, NOT to the actual nodes. So `ssh acerpve` lands on the VPS. **Nodes must be addressed by IP**: acerpve 192.168.68.9, amdpve 192.168.68.15, storepve 192.168.68.6, minipve 192.168.68.12, ocupve 192.168.68.5. Guest CTs are reached through their node (`pct exec`). Guest hostnames that resolve on the LAN (e.g. kagentz = 192.168.68.14) are fine. (The DNS address records are a separate decision — row: dag-daemon-node-hostnames-resolve-to-the-vps-20260915.)
#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts)
Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass:
```bash
# Check 1: Pool headroom (PRIMARY - yields percentages directly)
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname, see warning above
lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data
# Example output (acerpve, 192.168.68.9):
# LV Data% Meta% LSize
# data 29.95 1.22 <816.21g
# Required thresholds (documented minimum):
# data_percent < 90% (80% recommended for safety margin)
# metadata_percent < 70% (metadata fills faster than data)
# Check 2: Verify pool is not in error state (dmsetup shows the raw DM device)
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname
dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1
# dmsetup status pve-data-tpool field order (verified on 192.168.68.9):
# $1=start $2=length $3="thin-pool" $4=transaction-id
# $5=metadata_used/metadata_total (blocks) $6=data_used/data_total (sectors)
# remaining fields are flags ("-", "rw", "discard_passdown", "queue_if_no_space", ...)
# This is only used for the ERROR-STATE check; use the lvs command above for percentages.
# metadata_percent = $5 / ($5 split by /) [second number in pair]
```
**Minimum thresholds**: If either `data_percent >= 90%` or `metadata_percent >= 70%`, the backup MUST NOT start. State explicitly that these are hard stops, not warnings.
**Why this is a precondition**: A full or errored thin pool fails EVERY volume on the VG at once, including the host root. This is not a soft failure - it takes down the entire Proxmox host.
#### Staging Directory Requirement (2026-09-14 incident)
Any custom vzdump `tmpdir` MUST be world-traversable and writable exactly like `/var/tmp` (mode 1777). The archive step of vzdump runs in an unprivileged user namespace and cannot traverse a root-owned 0700 directory.
**Symptom to recognize**: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES` on every container in the backup run.
**Fix**: `chmod 1777 <custom-tmpdir>` before starting vzdump.
#### Task Start Rule for Truncating Shells
When starting a backup task from a shell that may truncate output (e.g., pipes, `head`), always use:
```bash
pvesh create /storage/backup --output-format json -- ... | head -2
# ❌ Can kill the backup task ("broken pipe" status)
```
Instead, capture JSON output without piping to truncating commands:
```bash
# Use --output-format json and capture to variable
result=$(pvesh create /storage/backup --output-format json -- ...)
# Then parse result if needed
```
#### GPU Host Backup Status (acerpve VM 101)
VM 101 (llm-gpu) and VM 103 (ocu-llm) have NO scheduled backup job. Coverage is manual one-off runs only. This is intentional for VM 101 until the thin-pool failure mechanism is understood and documented.
## Section 5: Network Services — Monitoring
### 5.1 Service Inventory
@@ -636,7 +563,7 @@ monitor, or integration breaks.
```bash
# Full cluster status
PVE="https://minipve.sysloggh.net"
AUTH="Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
# Docker health from Abiba
@@ -682,7 +609,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | **.17** | Git | ❌ |
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
| 112 | tanko | minipve | .122 | hybrid (DSH + Hermes) agent | ✅ |
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
@@ -712,7 +639,7 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| 100 | abiba | minipve | `pct-run 100` |
| 105 | kagentz | amdpve | `pct-run 105` |
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
| 112 | tanko | minipve | `pct-run 112` |
| 112 | tanko | amdpve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` |
| 104 | authentik | minipve | `pct-run 104` |
@@ -759,4 +686,4 @@ which was kill+nohup outside systemd) are banned by policy.
| Script | Why Disabled |
|--------|-------------|
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
| `zulip-monitor.sh` (Abiba) | Standalone cron replaced by agent-health-check.py + PM2 auto-restart; the script itself remains active as the execution step of the `zulip-health` responsibility contract (see `zulip-health.prose.md`) |
| `zulip-monitor.sh` (Abiba) | Replaced by agent-health-check.py + PM2 auto-restart |
+3 -39
View File
@@ -102,13 +102,6 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
- Stack persists across reboots (systemd for exporters, Docker restart policy)
## Execution
**Execution model**: This contract is executed by a host-scheduled cron job (see
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
### Liveness rule (scoped)
@@ -145,19 +138,11 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
tool calls; never repeat a prior report unless a live probe fails.**
**EXECUTABLE OWNER:** The canonical probe set lives in `scripts/infra-monitoring.sh`.
A check run is a single command: `bash scripts/infra-monitoring.sh` (from the
repository root). Paste its raw output verbatim into the report. The script
exits non-zero naming every failed target; there is no "OK" summary when any
leg failed. Port drift is caught by `scripts/test_infra_monitoring.sh` which
asserts every probed port matches the documented value.
**PROBE SHAPE (per standing rules above):**
- Every probe prints: `✅ <name>: alive` on success, or `🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)` on failure
- PVE API failures include `(any-HTTP liveness, -k for self-signed)` to distinguish TLS vs connection
- Retry once on connection failure at longer timeout (25s connect, 30s max)
- Every probe prints the target name + URL + HTTP code (or failure kind)
- Retry once on connection failure at longer timeout
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
- Report the actual probe output, not a summary verdict
- Report the actual probe command and its result, not a summary verdict
```bash
# Provenance — run first; paste the absolute path into the report
@@ -284,27 +269,6 @@ code (or failure kind with retry details). Apply the standing probe rules: any
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
a failure.
### Docker Stats and PVE Exporter Ports
These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:
| Exporter | Port | Container | Metrics |
|----------|------|-----------|---------|
| **Docker Stats** | **9324** | harness-docker-stats | `docker_container_*` (per-container CPU/mem/network) |
| **PVE Exporter** | **9221** | harness-pve-exporter | `pve_*` (5 cluster-level metrics) |
**IMPORTANT**: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (`builder_builds_*`, `containerd_build_info_*`).
```bash
# Docker Stats (harness-docker-stats)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
# Expected: 200 or 404 (any HTTP status = ALIVE)
# PVE Exporter (harness-pve-exporter)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
# Expected: 200 or 404 (any HTTP status = ALIVE)
```
### Phase 1: GPU Exporters
**NVIDIA (.8 and .110)**:
+8 -8
View File
@@ -59,7 +59,7 @@ Before ANY update wave:
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, minipve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
@@ -208,19 +208,19 @@ mcp_servers:
| Key | MCP Access |
|-----|-----------|
| Master key | ✅ Full — 90 tools (vault-injected) |
| Agent keys (mumuni, tanko, etc.) | ✅ Per-key grants supported (as of 2026-09-18 verification) |
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.99.1 |
### Known Limitations
- ~~Per-key MCP server grants not functional — only master key has access~~ (resolved 2026-09-18: per-key grants now work)
- Per-key MCP server grants not functional — only master key has access
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
### Migration Path (COMPLETED 2026-09-18)
Per-key MCP grants are now supported:
1. ✅ Agent keys granted MCP access via `allowed_mcp_servers` field
2. ✅ Hermes `mcp_servers.litellm.url` set to `https://litellm.sysloggh.net/mcp`
3. ✅ `headers: {x-litellm-api-key: "Bearer <literal_key>"}` added to MCP config
### Migration Path
When LiteLLM is upgraded to a version supporting per-key MCP grants:
1. Grant agent keys `mcp_servers: ["ra_h_os"]`
2. Update Hermes `mcp_servers.ra-h-os.url` from `http://192.168.68.65:3100/mcp` → `http://192.168.68.116:4000/mcp/`
3. Add `headers: {x-litellm-api-key: "Bearer $LITELLM_API_KEY"}` to MCP config
## Security-Specific Updates
+8 -72
View File
@@ -71,10 +71,6 @@ description: >
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
In LiteLLM Community both silently grant access to EVERY model, including cloud.
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
- Return the new key
5. **If action == "rotate"**:
@@ -91,66 +87,6 @@ description: >
- Confirm key alias matches agent_name in LiteLLM key list
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
## Cloud Provider Consolidation (2026-09-20)
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
### Provider map (per-account namespacing)
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
stay separate:
| Prefix | Upstream | Auth | Vault secret |
|--------|----------|------|--------------|
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
### Access tiers (MUST be enforced per key)
| Tier | Model names | Granted to |
|------|-------------|------------|
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
### Creating the cloud-enabled key
```bash
# ALWAYS read the live roster first (key-scoped):
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
| jq -r '.data[].id'
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
```
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
any `<prefix>/` cloud model.
### Adding a new cloud provider
1. Add the upstream key to Infisical `infrastructure/production/root`.
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
3. Restart the `harness-litellm` container.
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
captain approval.
5. Update this table and the access-tier section.
## Production Vault Access Process (canonical, 2026-07-17)
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
@@ -208,8 +144,8 @@ through its agent wrapper.
safety net for vault outage or token revocation. Must be kept in sync on rotation.
Example:
```bash
MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY»
MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY»
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
```
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
```ini
@@ -255,7 +191,7 @@ through its agent wrapper.
### Tanko migration (COMPLETED 2026-07-17)
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-synthetic-tanko-example…`)
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
@@ -335,13 +271,13 @@ not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update ma
UI bootstrap, model selection) is built around OpenRouter's native authentication.
**Key Storage:**
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»…`)
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
- **Fallback**: The container's .env is the primary source; vault sync is optional
(unlike fleet agents which require vault injection)
**Current Key (2026-09-01):**
- **Prefix**: `«vault: agents/production OPENROUTER_API_KEY»`
- **Prefix**: `sk-or-v1-0af3f3…`
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
- **Plan**: Paid (not free tier)
- **Usage**: 0 (as of 2026-09-01)
@@ -386,10 +322,10 @@ directly call OpenRouter via Python's requests library. Converting would require
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
```bash
# PRIMARY (proven, runs on CT 116 with no extra tooling):
# Read at runtime from the container's environment:
docker exec harness-litellm printenv LITELLM_MASTER_KEY
# Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching)
# The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it)
# Or from Infisical vault (project=infrastructure env=prod) - NOTE: --plain is broken on CLI 0.43.110 (prints nothing):
infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production | awk '$1=="LITELLM_MASTER_KEY"{print $NF}'
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
```
+1 -8
View File
@@ -19,7 +19,7 @@ description: >
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
- Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
---
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
@@ -111,13 +111,6 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
## Execution
**Execution model**: This contract is executed by a host-scheduled cron job (see
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
1. **Read parameters** — Use provided values or defaults
+1 -1
View File
@@ -116,7 +116,7 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (hardcoded key → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM (deprecated key, no live usage).
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
## Maintains
+6 -44
View File
@@ -6,19 +6,13 @@ name: memory-fixer
description: >
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
version: 2.2.0
version: 2.1.0
---
---
# Memory Fixer
> **Executable copy:** the `okyeame-memory-fixer` cron job on kagentz (`hermes cron list`) holds its instruction
> set **inline in `~/.hermes/cron/jobs.json`** (`hermes cron edit <id> --prompt …`; there is no `--prompt-file`, and
> `~/.hermes/cron/memory-fixer-prompt.md` is a synced draft, not the live instruction). This file is the institutional
> record of the same contract; when the two diverge, the job prompt is what actually runs — diff it against this file
> before claiming a prompt change landed.
> ⚠️ Corrected 2026-09-26: the previous pointer (`/root/.hermes/contracts/memory-fixer-v3.md`) does not exist on
> kagentz — no `/root` access from this container — and was verified unreachable, not merely stale.
> **Canonical copy:** `/root/.hermes/contracts/memory-fixer-v3.md` (used by the `memory-fixer-daily` cron job). This file is the institutional record of the same contract. When the two diverge, treat the v3 source in `/root/.hermes/contracts/` as executable truth.
## Purpose
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever.
@@ -131,44 +125,15 @@ updateNode(id, {
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
### 5. Duplicate-Node Detection (Level 1 — read-only, every run)
The graph's duplicate problem is rarely an agent mistyping a title: it is **recurring writers creating a new
node per run instead of updating one**. This phase detects that class and reports it. It is read-only and
**never merges**.
```bash
python3 /home/hermes/.hermes/scripts/memory_dup_detect.py --json
```
Read-only, ~15s over the whole graph, exit 0. That script is the source of truth for the clustering logic —
do not re-implement it in the prompt or hand-count "duplicates" from titles.
Consume each `items[]` entry's `verdict` field; do not invent your own:
| `verdict` | Meaning | Required action |
|---|---|---|
| `WRITER-DEFECT` (`run_family: true`) | ONE scheduled task writes a new node per run | Report the ids, the `agents` (the writer) and `span_days`. **Never merge** — each node is that run's audit record. If the family grew since the last report, say `UNFIXED` and name the writer. |
| `SAFE-MERGE` | Bodies identical | Still requires an explicit `merge #A into #B` decision from Kwame. |
| `HUMAN-DECISION` | Same subject, bodies differ | Propose **connect (an edge)**, never merge. |
- **Title overlap alone is not duplication.** Four distinct client workflows of one family (#357-#361) and two
different machines' migrations (#1792/#1793) both score high on title tokens while their bodies sit 0.1-0.3
apart. Confirm against body similarity before calling anything a duplicate.
- Report clusters as **candidates for Kwame's decision**, never as established duplicates — a wrong auto-merge
destroys distinct content irrecoverably.
- Per-run history nodes are kept deliberately. Bulk-merging a run family destroys the audit trail the family exists for.
## Level 2 Escalations (Kwame Decision Required)
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
2. **Duplicate Nodes** — as detected by fix 5, by `verdict`, never by raw title overlap. `WRITER-DEFECT` is a writer fix (update one canonical node), not a merge decision; `SAFE-MERGE` and `HUMAN-DECISION` clusters are escalated for merge-or-connect.
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
3. **Orphan Nodes >90 days old** — Archive or connect?
## Reporting Format
The fixer does **not** send anything. Under the single-egress model (2026-09-21) every report leaves the node
through Mumuni's gate (`comms_drop.py` for the queue, `comms_gate.py` to release and read-back verify), so
exit 0 means QUEUED, never delivered. A report body is written to a file and handed to the outbox helper:
The fixer reports to Kwame via this Zulip DM:
```
🦅 Memory Fixer — [HH:MM UTC]
@@ -182,10 +147,8 @@ Stale nodes needing review (max 10):
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
...
Duplicate clusters (candidates — Kwame decides; the fixer never merges unilaterally):
1. [WRITER-DEFECT] #AAA/#BBB/#CCC — writer <agent>, N nodes, span Nd (UNFIXED if it grew since the last report)
2. [HUMAN-DECISION] #DDD/#EEE — same subject, bodies differ, SUGGEST: connect
3. "none" when the scan returned no clusters
Duplicates needing decision:
1. [Node #AAA] vs [Node #BBB] — Same title
Orphans >90 days:
1. [Node #EEE] Title — X days stale, orphaned
@@ -232,7 +195,6 @@ The result must be 0 rows when all decisions are executed. Report what was done.
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
- **Duplicate scan ran:** every report carries the fix 5 block (`none` when there were no clusters). A report with no duplicate section means phase 5 was skipped — a silently skipped detection phase is the failure this phase exists to prevent.
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
## Logging
+15 -39
View File
@@ -2,10 +2,16 @@
kind: responsibility
name: pm2-self-heal
description: >
Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
and auto-restarts any that are stopped or errored. Logs every action to
Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary).
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
errored. Logs every action to the knowledge graph and alerts the owner via
Zulip DM on failures.
CRITICAL: Never restart abiba-zulip — it runs this contract.
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
2026-07-04 'removed/decommissioned' note was stale and is removed).
---
## Maintains
@@ -16,8 +22,6 @@ description: >
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed.
## Continuity
@@ -41,9 +45,6 @@ description: >
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
spoton incident). Alerts include the restart count.
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton
reference above is historical context for the crash-loop guard, not a live process.
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
@@ -51,44 +52,19 @@ description: >
and PM2 counter reset on 2026-06-28.
## Execution
**Execution model**: This contract is executed by a host-scheduled cron job (see
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
2. **Check abiba-telegram** (safe to auto-restart):
2. **Check abiba-telegram**:
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
- If restarts > 5 → alert owner
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
- If status is "online" → pass, log restarts count
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
- If restarts > 5 in last hour → alert owner with full diagnostics
4. **Check gitea-runner**:
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 5 → alert owner
5. **Check zulip-watchdog**:
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 5 → alert owner
6. **Log results** — the durable per-run record is `/var/log/contract-runs/pm2-self-heal-<UTCstamp>.log` on CT 100, written by `scripts/contract-run.sh` from `/etc/cron.d/contract-runner` every 4 hours, with a firstmate inbox note raised on any non-zero exit.
**CORRECTED 2026-09-26 (relay-785):** this step previously required appending to
`SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea and called it a "hard rule".
That posting was **never implemented** — `scripts/pm2-self-heal.sh` contains no
Gitea or git-push code — so the requirement was a coverage claim the executor did
not honour, and `health-logs/pm2/` has held only its init commit since 2026-07-28.
The claim is retired rather than implemented: the contract-runner's per-run logs
plus its failure note already give a durable record and a working alarm, and a
second posting path would add work without adding a signal. `health-logs/pm2/`
is left as historical evidence, not as a live obligation.
7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
8. **Wait 5 min** → repeat from step 1
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
6. **Wait 5 min** → repeat from step 1
## Example Output (when healthy)
-42
View File
@@ -89,48 +89,6 @@ agent: abiba
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
## PBS GC (Proxmox Backup Server)
### Schedule
Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**.
**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.
### What Actually Runs
Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`.
**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC.
### Datastore Location
- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free)
- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume)
### Liveness Check
The new `proxmox-monitor.sh` leg checks storepve-datastore GC health:
- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json`
- **FAILS** if `last-run-endtime` is older than 48 hours
- Reports age in hours and pending-bytes
**All six verdict shapes** (exactly as emitted by the script):
1. **Healthy** (fresh GC, 0 B pending):
`✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)`
2. **Stale** (GC ran >48h ago):
`🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)`
3. **Probe-failed: empty read** (000/timeout/unreadable):
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)`
4. **Probe-failed: unparseable** (non-empty but invalid JSON — the "command not found" case):
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)`
5. **Never-run: datastore absent** (valid JSON but storepve-datastore not in list):
`🔴 PBS GC: never-run (storepve-datastore not found in GC list)`
6. **Never-run: no endtime** (valid JSON with datastore present but last-run-endtime is null/0):
`🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)`
## Operations
### view-dashboards
+30 -59
View File
@@ -50,11 +50,6 @@ Changelog:
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
from her side. This script must not probe mumuni or .24 — the v2 changelog
roster line was the last reference still placing her at .24 / CT100.
v6 (2026-09-28): .8 GPU health probe now runs as `llmuser` instead of `root`.
Root SSH to .8 was lost when the guest was rebuilt, so every .8 leg read as
UNREACHABLE for a healthy host. llmuser owns llama-server and can read
`systemctl is-active`, `systemctl show -p MainPID`, and the :8080 pid.
.110 and .15 keep the default `root` user.
"""
import subprocess, json, sys, os, time, re, io, contextlib
@@ -75,7 +70,7 @@ PVE_NODES = {
# Agent definitions: ct, host, user, pve_node, vault_key_name
AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "minipve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "hybrid"},
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
# abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its
# local env file (key_env below), not from the shared vault or .bashrc.
# runtime=pi: abiba has run pi-only since the harness purge. There is no
@@ -100,7 +95,7 @@ AGENTS = {
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
# .15 strixhalo (amdpve) -> strix-server.service (active)
GPU_HOSTS = {
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service", "user": "llmuser"},
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"},
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
}
@@ -127,8 +122,10 @@ def _fail(key, agent_name=None):
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
# Fallback: if no env token, read the shared vault token file
if not INFISICAL_TOKEN:
# Fallback: read the shared vault token file
_token_path = os.path.expanduser("~/.infisical-token")
if os.path.isfile(_token_path):
try:
@@ -136,7 +133,6 @@ if not INFISICAL_TOKEN:
INFISICAL_TOKEN = _f.read().strip()
except (OSError, UnicodeDecodeError):
pass
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
# ── Helpers ──────────────────────────────────────────────────────────
@@ -152,18 +148,6 @@ def ssh(host, cmd, user="root"):
except:
return None
def get_user_home(user):
"""Resolve the home directory for a user.
For 'root', returns '/root'. For any other user, returns '/home/<user>'.
This is used to construct paths that reference a user's home directory
(e.g., ~/.local/bin/hermes, ~/.hermes/config.yaml) instead of hardcoding /root/.
"""
if user == "root":
return "/root"
else:
return f"/home/{user}"
def http_get(url, headers=None, timeout=5):
"""Return HTTP status code as string."""
try:
@@ -317,14 +301,13 @@ def check_gpu_ports():
host = gpu["host"]
port = gpu["port"]
svc = gpu["service"]
user = gpu.get("user", "root") # default root, overridden per-host where needed
# `systemctl is-active` exits non-zero when the unit is inactive or
# missing, which the ssh() helper would swallow as an SSH failure and
# report as UNREACHABLE. `|| true` keeps the real state word so we can
# tell "unit inactive" from "host unreachable".
svc_status = ssh(host, f"systemctl is-active {svc} || true", user=user)
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1", user=user)
svc_status = ssh(host, f"systemctl is-active {svc} || true")
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
if not svc_status:
print(f" ❌ {label}: UNREACHABLE")
@@ -335,14 +318,14 @@ def check_gpu_ports():
print(f" ❌ {label}: PORT {port} NOT LISTENING (svc={svc_status})")
FAIL.append(f"gpu-no-port:{label}")
elif svc_status != "active":
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2", user=user)
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2")
if svc_pid and port_owner != svc_pid:
print(f" ❌ {label}: GHOST PROCESS — port owned by pid {port_owner}, svc pid {svc_pid} (svc={svc_status})")
FAIL.append(f"gpu-ghost:{label}:{port_owner}")
else:
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
else:
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health", user=user)
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health")
if health and '"status":"ok"' in health:
print(f" ✅ {label}: healthy (pid={port_owner})")
elif health and '"status":"no slot available"' in health:
@@ -394,10 +377,10 @@ def check_agents():
ct = agent["ct"]
report_only = agent.get("report_only", False)
# Tanko runs hybrid (DSH + Hermes) since 2026-08-27 — it runs both DSH and Hermes gateway.
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
# Non-Hermes runtimes have no gateway to probe. dsh/pi-only skip the check;
# hybrid runs both DSH and Hermes and is checked normally.
# Non-Hermes runtimes have no gateway to probe. dsh = Tanko since
# 2026-08-27; pi = abiba since the harness purge (.24 is pi-only).
if agent.get("runtime") in ("dsh", "pi"):
is_dsh = agent.get("runtime") == "dsh"
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
@@ -518,7 +501,7 @@ def check_ct_liveness():
def check_config_integrity():
"""Verify agent config.yaml parses as valid YAML."""
for name, agent in AGENTS.items():
# DSH/pi-only runtimes have no Hermes config.yaml; hybrid has both.
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
if agent.get("runtime") == "dsh":
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
continue
@@ -531,11 +514,11 @@ def check_config_integrity():
print(f" ⬜ {name}: cannot SSH — skip config check")
continue
home = get_user_home(user)
# Check YAML parses
yaml_ok = ssh(host,
f"python3 -c \"import yaml; yaml.safe_load(open('{home}/.hermes/config.yaml')); print('OK')\" 2>&1 || echo 'FAIL'",
"python3 -c "
'"import yaml; yaml.safe_load(open(\'/root/.hermes/config.yaml\')); print(\'OK\')" '
"2>&1 || echo 'FAIL'",
user=user)
if not yaml_ok:
print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)")
@@ -574,7 +557,7 @@ def _infisical_invocation_paths(wrapper_body):
def check_wrapper_integrity():
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
for name, agent in AGENTS.items():
# DSH/pi-only runtimes have no hermes CLI wrapper; hybrid has both.
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
if agent.get("runtime") == "dsh":
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
continue
@@ -588,8 +571,7 @@ def check_wrapper_integrity():
continue
# Check wrapper exists
home = get_user_home(user)
wrapper = ssh(host, f"ls -la {home}/.local/bin/hermes 2>/dev/null", user=user)
wrapper = ssh(host, "ls -la /root/.local/bin/hermes 2>/dev/null", user=user)
if not wrapper:
# Check alternate wrapper locations
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
@@ -612,7 +594,7 @@ def check_wrapper_integrity():
# a removed path (litellm-api-keys.prose.md documents
# `rm -f /usr/local/bin/infisical`) must neither produce a dangling path
# nor trigger the PATH check — it is not an invocation.
wrapper_body = ssh(host, f"cat {home}/.local/bin/hermes 2>/dev/null", user=user) or ""
wrapper_body = ssh(host, "cat /root/.local/bin/hermes 2>/dev/null", user=user) or ""
wrapper_code = "\n".join(line.split("#", 1)[0] for line in wrapper_body.splitlines())
invoked_paths = _infisical_invocation_paths(wrapper_body)
if "infisical" in wrapper_code:
@@ -646,35 +628,24 @@ def check_wrapper_integrity():
else:
print(f" ℹ️ {name}: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK")
# Check that the wrapper's target resolves. The fleet's wrappers do NOT
# all use a hermes-real indirection — some exec the venv module directly.
# Verify the wrapper actually points to something runnable.
# Check hermes-real exists
hermes_real = ssh(host,
f"ls -la {home}/.local/bin/hermes-real 2>/dev/null || echo MISS",
"ls -la /root/.local/bin/hermes-real 2>/dev/null || echo MISS",
user=user)
if hermes_real and hermes_real.strip() != "MISS":
print(f" ✅ {name}: wrapper shape: hermes-real at {home}/.local/bin/hermes-real")
else:
# Try the venv under home
venv_home = ssh(host,
f"test -x {home}/.hermes/hermes-agent/venv/bin/python && echo OK || echo MISS",
if not hermes_real or hermes_real.strip() == "MISS":
# Check venv path
hermes_real = ssh(host,
"ls -la /usr/local/lib/hermes-agent/venv/bin/hermes 2>/dev/null || echo MISS",
user=user)
if venv_home and venv_home.strip().splitlines()[-1] == "OK":
print(f" ✅ {name}: wrapper shape: direct venv exec ({home}/.hermes/hermes-agent/venv/bin/python)")
if not hermes_real or hermes_real.strip() == "MISS":
print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)")
_fail(f"wrapper-no-hermes-real:{name}", name)
else:
# Try the system-wide venv
venv_sys = ssh(host,
"test -x /usr/local/lib/hermes-agent/venv/bin/python && echo OK || echo MISS",
user=user)
if venv_sys and venv_sys.strip().splitlines()[-1] == "OK":
print(f" ✅ {name}: wrapper shape: system venv (/usr/local/lib/hermes-agent/venv/bin/python)")
else:
print(f" ❌ {name}: wrapper target NOT RESOLVABLE (no hermes-real, no venv)")
_fail(f"wrapper-no-hermes-real:{name}", name)
print(f" ✅ {name}: hermes-real at alt path")
# Check the .env file has the key
env_has_key = ssh(host,
f"grep -c 'LITELLM_API_KEY' {home}/.hermes/.env 2>/dev/null || echo 0",
"grep -c 'LITELLM_API_KEY' /root/.hermes/.env 2>/dev/null || echo 0",
user=user)
if env_has_key and env_has_key.strip() not in ("", "0"):
print(f" ✅ {name}: wrapper + .env key present")
-225
View File
@@ -1,225 +0,0 @@
#!/bin/bash
# contract-run.sh — Deterministic contract execution from machine scheduler
#
# Takes a contract name, resolves its script, runs it with timeout,
# logs output to $CONTRACT_RUN_LOG_DIR (default: /var/log/contract-runs/),
# and alerts on failure.
#
# Environment:
# CONTRACT_RUN_LOG_DIR Override the log directory (default: /var/log/contract-runs)
#
# Usage: bash scripts/contract-run.sh <contract-name>
#
# Contract names map to scripts as follows:
# infrastructure-monitoring -> scripts/infra-monitoring.sh
# proxmox-monitor -> scripts/proxmox-monitor.sh
# zulip-health -> scripts/zulip-monitor.sh
# agent-health-check -> scripts/agent-health-check.py
# litellm-health -> scripts/litellm-health-check.py
# disk-gc-threat-response -> scripts/disk-gc-scan.py
# pm2-self-heal -> scripts/pm2-self-heal.sh
# search-stack-visibility -> scripts/search-stack-check.py
# health-log-freshness -> scripts/health-log-freshness.py
#
# Execution copy: every contract pins the clone this script lives in (see
# docs/contract-execution-pinning.md). Before a contract runs, this wrapper
# proves the script it is about to execute byte-matches origin/master:
# CONTRACT_REVISION_PREFLIGHT=warn (default) log a refusal, still report
# CONTRACT_REVISION_PREFLIGHT=enforce refuse to report on a mismatch
# CONTRACT_REVISION_PREFLIGHT=off skip the check entirely
# A refusal names its class: cannot-verify:fetch-failed|ref-unresolvable,
# or mismatch:content|path-absent|detached-head|clone-ahead.
#
# Exit codes:
# 0 = contract passed
# 1 = contract failed (alert sent)
# 2 = probe failed (script missing, timeout, etc.)
set -uo pipefail
CONTRACT_NAME="$1"
SCRIPTS_DIR="$(cd "$(dirname "$0")" && pwd)"
LOG_DIR="${CONTRACT_RUN_LOG_DIR:-/var/log/contract-runs}"
TIMESTAMP=$(date -u '+%Y%m%d-%H%M%S')
LOG_FILE="${LOG_DIR}/${CONTRACT_NAME}-${TIMESTAMP}.log"
# Ensure log directory exists
mkdir -p "$LOG_DIR"
# Map contract name to script path
case "$CONTRACT_NAME" in
infrastructure-monitoring)
SCRIPT_PATH="${SCRIPTS_DIR}/infra-monitoring.sh"
INTERPRETER="bash"
;;
proxmox-monitor)
SCRIPT_PATH="${SCRIPTS_DIR}/proxmox-monitor.sh"
INTERPRETER="bash"
;;
zulip-health)
SCRIPT_PATH="${SCRIPTS_DIR}/zulip-monitor.sh"
INTERPRETER="bash"
;;
agent-health-check)
SCRIPT_PATH="${SCRIPTS_DIR}/agent-health-check.py"
INTERPRETER="python3"
;;
litellm-health)
SCRIPT_PATH="${SCRIPTS_DIR}/litellm-health-check.py"
INTERPRETER="python3"
;;
pm2-self-heal)
SCRIPT_PATH="${SCRIPTS_DIR}/pm2-self-heal.sh"
INTERPRETER="bash"
;;
disk-gc-threat-response)
SCRIPT_PATH="${SCRIPTS_DIR}/disk-gc-scan.py"
INTERPRETER="python3"
;;
search-stack-visibility)
SCRIPT_PATH="${SCRIPTS_DIR}/search-stack-check.py"
INTERPRETER="python3"
;;
health-log-freshness)
SCRIPT_PATH="${SCRIPTS_DIR}/health-log-freshness.py"
INTERPRETER="python3"
;;
*)
echo "Unknown contract: $CONTRACT_NAME" | tee -a "$LOG_FILE"
# Send alert for unknown contract
ALERT_MSG="🔴 Contract $CONTRACT_NAME: unknown contract name. Log: $LOG_FILE"
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=private" \
-d "to=9" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
fi
exit 2
;;
esac
# Check if script exists
if [ ! -f "$SCRIPT_PATH" ]; then
echo "Script not found: $SCRIPT_PATH" | tee -a "$LOG_FILE"
# Send alert for missing script
ALERT_MSG="🔴 Contract $CONTRACT_NAME: script not found at $SCRIPT_PATH. Log: $LOG_FILE"
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=private" \
-d "to=9" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
fi
exit 2
fi
# Run the script with timeout and capture output
echo "=== Contract: $CONTRACT_NAME ===" | tee "$LOG_FILE"
echo "Started: $(date -u '+%Y-%m-%d %H:%M:%S UTC')" | tee -a "$LOG_FILE"
echo "Script: $SCRIPT_PATH" | tee -a "$LOG_FILE"
echo "" | tee -a "$LOG_FILE"
# ── Revision preflight ───────────────────────────────────────────────────────
# A verdict is only meaningful if it came from the merged copy. This reports
# whether the executing copy matches, and distinguishes "could not check" from
# "this copy is wrong" so an operator can tell them apart.
# See docs/contract-execution-pinning.md.
#
# Default is WARN, not enforce: the guard gates every scheduled contract, and a
# legitimate state (branch mid-review, detached HEAD, briefly offline) would
# otherwise turn the whole fleet's monitoring into withheld verdicts. The
# criteria for flipping the default to enforce are written down in the doc.
REVISION_PREFLIGHT_MODE="${CONTRACT_REVISION_PREFLIGHT:-warn}"
REPO_ROOT="$(cd "${SCRIPTS_DIR}/.." && pwd)"
if [ "$REVISION_PREFLIGHT_MODE" != "off" ] && [ -x "${SCRIPTS_DIR}/revision-preflight.sh" ]; then
PREFLIGHT_OUT="$(mktemp)"
if "${SCRIPTS_DIR}/revision-preflight.sh" "$SCRIPT_PATH" "$REPO_ROOT" >"$PREFLIGHT_OUT" 2>&1; then
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
else
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
PREFLIGHT_REASON="$(grep -m1 '^REASON=' "$PREFLIGHT_OUT" | cut -d= -f2-)"
[ -n "$PREFLIGHT_REASON" ] || PREFLIGHT_REASON="unclassified"
if [ "$REVISION_PREFLIGHT_MODE" = "enforce" ]; then
echo "🚫 VERDICT WITHHELD: $PREFLIGHT_REASON" | tee -a "$LOG_FILE"
ALERT_MSG="🔴 Contract $CONTRACT_NAME: revision preflight REFUSED ($PREFLIGHT_REASON) — verdict withheld. Log: $LOG_FILE"
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=private" \
-d "to=9" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
fi
rm -f "$PREFLIGHT_OUT"
exit 2
fi
echo "⚠️ revision preflight: $PREFLIGHT_REASON — continuing because CONTRACT_REVISION_PREFLIGHT=$REVISION_PREFLIGHT_MODE" | tee -a "$LOG_FILE"
fi
rm -f "$PREFLIGHT_OUT"
fi
# Use timeout to prevent hangs (10 minutes default)
TIMEOUT=600
timeout "$TIMEOUT" $INTERPRETER "$SCRIPT_PATH" 2>&1 | tee -a "$LOG_FILE"
EXIT_CODE=${PIPESTATUS[0]}
# If timeout killed the process, EXIT_CODE will be 124
if [ $EXIT_CODE -eq 124 ]; then
echo "⏰ TIMEOUT: script exceeded ${TIMEOUT}s limit" | tee -a "$LOG_FILE"
fi
echo "" | tee -a "$LOG_FILE"
if [ $EXIT_CODE -eq 0 ]; then
echo "✅ VERDICT: PASS" | tee -a "$LOG_FILE"
exit 0
else
echo "🔴 VERDICT: FAIL (exit code $EXIT_CODE)" | tee -a "$LOG_FILE"
# Send alert (Zulip DM to user 9 + stream agent-hub topic alerts-infra)
# Using the same alert path as other monitors
ALERT_MSG="🔴 Contract $CONTRACT_NAME failed (exit $EXIT_CODE). Log: $LOG_FILE"
ALERT_SENT=false
# Take credentials from environment (ZULIP_API_KEY required)
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
# DM to user 9
DM_EXIT=0
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=private" \
-d "to=9" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || DM_EXIT=$?
# Stream agent-hub topic alerts-infra
STREAM_EXIT=0
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=stream" \
-d "to=agent-hub" \
-d "topic=alerts-infra" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || STREAM_EXIT=$?
if [ $DM_EXIT -eq 0 ] || [ $STREAM_EXIT -eq 0 ]; then
ALERT_SENT=true
else
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT FAILURE: DM exit=$DM_EXIT, stream exit=$STREAM_EXIT" >> "$LOG_FILE"
fi
else
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT SKIPPED: no ZULIP_API_KEY or curl" >> "$LOG_FILE"
fi
exit 1
fi
+40 -230
View File
@@ -16,42 +16,14 @@ from email.mime.text import MIMEText
from email.mime.multipart import MIMEMultipart
PVE = "https://192.168.68.12:8006"
# The HTTP header prefix is a protocol constant, not a credential. It is kept as
# a constant ending at '=' so that no assembled header-plus-token literal ever
# appears in the tree; the secret scanner rightly flags that shape.
PVE_AUTH_HEADER = "PVEAPIToken="
def pve_auth():
"""PVE API auth header, resolved at call time from the injected environment.
The token is injected by ``infisical run --env=prod`` as ``PVE_TOKEN``
(format ``user@realm!tokenid=secret``). It must never be hardcoded: a
placeholder literal authenticates as nobody, which is how this probe
reported zero nodes while still exiting 0. Raise loudly instead.
"""
token = os.environ.get("PVE_TOKEN")
if not token:
raise RuntimeError("PVE_TOKEN is not set (run under `infisical run --env=prod`)")
return f"Authorization: {PVE_AUTH_HEADER}{token}"
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
# ── Shared credentials —─
ZULIP_SITE = "https://chat.sysloggh.net"
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
# Note: /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential).
# No Zulip API key is required for this call. If a future leg genuinely needs abiba-bot's key,
# it must prove it with a 200 from /api/v1/users/me as abiba-bot and label itself degraded when it cannot.
# Never fall back to the vault's shared ZULIP_API_KEY.
ZULIP_AUTH = None
DEGRADED_LEGS = []
# Probe failures are different from degraded legs. A missing credential is an
# expected, survivable state (stays exit 0). A probe that cannot reach the API
# means the report has NO data for that section, which is a monitoring loss and
# must exit non-zero so it cannot pass unnoticed.
PROBE_FAILURES = []
ZULIP_KEY = "cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_KEY}"
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
LITELLM_BACKEND = "192.168.68.116"
@@ -72,13 +44,9 @@ TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
# ── Helpers ──
def pve_get(path):
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list).
A missing PVE_TOKEN is caught here and reported as ``None`` so the caller
records a probe failure; it must not escape as an unhandled exception.
"""
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list)."""
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
try:
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{pve_auth()}"'
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
if r.returncode != 0:
return None
@@ -146,7 +114,6 @@ def collect():
report["node_count"] = 0
report["nodes_online"] = 0
report["pve_probe_status"] = "unreachable"
PROBE_FAILURES.append("proxmox: node list unreachable (PVE_TOKEN missing or API down)")
else:
report["nodes"] = {n["node"]: {
"cpu_pct": round(n.get('cpu',0)*100, 1),
@@ -166,7 +133,6 @@ def collect():
if resources is None:
vms = []
report["resources_probe_status"] = "unreachable"
PROBE_FAILURES.append("proxmox: cluster resources unreachable")
else:
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
report["resources_probe_status"] = "ok"
@@ -190,7 +156,7 @@ def collect():
# ── Storage ──
storages = pve_get("/api2/json/nodes/storepve/storage")
report["storage"] = []
for s in (storages or []):
for s in storages:
total = s.get("total",0) or 1
used = s.get("used",0)
pct = used/total*100
@@ -243,7 +209,7 @@ def collect():
("Pulse", "https://pulse.sysloggh.net"),
("Proxmox", "https://192.168.68.12:8006"),
("SearXNG", "http://192.168.68.7:8888"),
("Firecrawl", "http://192.168.68.7:3002/"), # Firecrawl serves no /health - the root is its liveness endpoint
("Firecrawl", "http://192.168.68.7:3002/health"),
]
report["endpoints"] = []
for name, url in endpoints:
@@ -389,28 +355,6 @@ def collect():
# ── HTML Dashboard ──
def classify_endpoint(code):
"""Classify an endpoint probe per the fleet's probe policy.
Codified 2026-09-14 in the monitoring contracts: ANY HTTP status proves the
service answered, so the service is ALIVE - 200/301/302/401/403/404 alike.
Only a failed CONNECTION (000 / timeout / refused) is a failed probe. A 404
from a wrong path is not a service fault and must not render as one.
This replaces a string comparison that was wrong in both directions
(`ep["code"] >= "400"`): it rendered 301 as red, 404 as yellow, and a real
500 as yellow. 5xx is kept as its own "server error" signal rather than
being merged with 4xx.
"""
if not code or code == "000":
return "red", "no connection"
if code.startswith("5"):
return "yellow", "server error"
if code.startswith(("2", "3", "4")):
return "green", "alive"
return "yellow", f"unexpected {code}"
def build_html(r):
issues = []
@@ -646,7 +590,7 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']
# ── Network Endpoints ──
html += '<div class="card"><h2>🌐 Network Endpoints</h2><table><tr><th>Service</th><th>Status</th></tr>'
for ep in r["endpoints"]:
color = classify_endpoint(ep["code"])[0]
color = "green" if ep["code"] in ("200","302","401") else ("yellow" if ep["code"] >= "400" else "red")
html += f'<tr><td>{ep["name"]}</td><td class="{color}">HTTP {ep["code"]}</td></tr>'
html += '</table></div>'
@@ -710,188 +654,60 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']
return html
# ── Delivery: Zulip DM carrying the report as an HTML ATTACHMENT ──
#
# Captain's decision, clarified 2026-09-26: the report is sent as an HTML FILE,
# i.e. an attachment - NOT HTML rendered in the message body, and NOT a Markdown
# translation of it. So the styled dashboard is built exactly as before, uploaded
# through Zulip's file-upload API, and the message body stays short: subject,
# top-line status, and a pointer to the attachment.
#
# This removes the Google dependency entirely (no SMTP, no EMAIL_PASSWORD).
# The 10,000-character message cap does not apply: it bounds message TEXT only,
# and the report travels as a file.
# ── Send Email ──
ZULIP_SITE = "https://chat.sysloggh.net"
ZULIP_BOT_EMAIL = "abiba-bot@chat.sysloggh.net"
CAPTAIN_USER_ID = 9
ZULIP_KEY_FILE = "/root/.pi/agent/extensions/zulip/.env"
REPORT_ARTIFACT_DIR = "/var/log/daily-infra-report"
def zulip_key():
"""abiba-bot's Zulip key, from the env or the on-host 600 file."""
key = os.environ.get("ABIBA_ZULIP_API_KEY")
if key:
return key.strip()
def send_email(html_content, subject_prefix=""):
FROM = "abiba@sysloggh.com"
TO = "jerome@sysloggh.com"
SUBJECT = f"{subject_prefix}{'🏗️ Infrastructure Report — ' + DATE_STR}"
msg = MIMEMultipart("alternative")
msg["From"] = FROM
msg["To"] = TO
msg["Subject"] = SUBJECT
msg.attach(MIMEText("Infrastructure report in HTML format — enable images to view.", "plain"))
msg.attach(MIMEText(html_content, "html"))
try:
with open(ZULIP_KEY_FILE) as fh:
for line in fh:
if line.startswith("ABIBA_ZULIP_API_KEY="):
return line.split("=", 1)[1].strip()
except OSError:
return None
return None
def build_summary(r, filename, test=False):
"""Short Markdown body: subject, top-line status, pointer to the attachment.
Deliberately NOT a reproduction of the report - the attachment is the report.
"""
nodes = f"{r.get('nodes_online', 0)}/{r.get('node_count', 0)} nodes online"
guests = f"{r.get('running_vms', 0)}/{r.get('total_vms', 0)} guests running"
lines = [
("\U0001F9EA **TEST — **" if test else "") + "\U0001F3D7\uFE0F **Infrastructure Report — " + DATE_STR + "**",
f"**{nodes}** \u00b7 **{guests}** \u00b7 generated {TIME_STR}",
]
problems = []
if r.get("pve_probe_status") != "ok":
problems.append(f"\u274c Proxmox probe: {r.get('pve_probe_status')}")
if r.get("resources_probe_status") != "ok":
problems.append(f"\u274c Resources probe: {r.get('resources_probe_status')}")
lit = r.get("litellm", {}) or {}
checks = lit.get("checks", []) or []
if checks:
passed = sum(1 for c in checks if c.get("status") == "pass")
if passed != len(checks):
problems.append(f"\u274c LiteLLM: {passed}/{len(checks)} checks pass")
if not (r.get("zulip_ext", {}) or {}).get("connected"):
problems.append("\u274c Zulip extension: not connected")
for leg in DEGRADED_LEGS:
problems.append(f"\u26a0\uFE0F degraded: {leg}")
lines.append("\n".join(problems) if problems else "\u2705 All monitored services healthy")
lines.append(f"\U0001F4CE **Full report attached:** `{filename}`")
return "\n\n".join(lines)
def _curl(args, timeout=60):
r = subprocess.run(["curl", "-s", "-m", str(timeout)] + args,
capture_output=True, text=True)
try:
return json.loads(r.stdout or "{}"), r.stdout
except json.JSONDecodeError:
return {}, r.stdout
def _curl_json(args, timeout=90):
r = subprocess.run(["curl", "-s", "-m", str(timeout)] + args,
capture_output=True, text=True)
try:
return json.loads(r.stdout or "{}"), r.stdout
except json.JSONDecodeError:
return {}, r.stdout
def send_zulip(html_content, report, test=False):
"""Upload the styled HTML and post a short pointer to the captain's DM.
Returns (ok, message). On ANY failure the report body is also printed to
stdout and persisted to disk, so a delivery failure can never swallow the
content - the defect this folds in.
"""
os.makedirs(REPORT_ARTIFACT_DIR, exist_ok=True)
stamp = NOW.strftime("%Y%m%d-%H%M%S")
filename = f"infra-report-{stamp}.html"
html_path = os.path.join(REPORT_ARTIFACT_DIR, filename)
try:
with open(html_path, "w") as fh:
fh.write(html_content)
except OSError as e:
print(f" \u26a0\uFE0F could not persist report artifact: {e}", file=sys.stderr)
key = zulip_key()
if not key:
print(html_content) # never swallow the content
return False, ("\u274c Delivery FAILED: no Zulip credential "
"(ABIBA_ZULIP_API_KEY unset and "
f"{ZULIP_KEY_FILE} unreadable). Report persisted to {html_path}")
auth = ["-u", f"{ZULIP_BOT_EMAIL}:{key}"]
# 1. Upload the report as a file.
up, up_raw = _curl_json(auth + [
"-X", "POST", f"{ZULIP_SITE}/api/v1/user_uploads",
"-F", f"file=@{html_path};type=text/html",
])
if up.get("result") != "success" or not up.get("uri"):
print(html_content)
return False, (f"\u274c Delivery FAILED at upload: {up.get('msg') or up_raw[:160]} "
f"(report persisted to {html_path})")
uri = up["uri"]
size = os.path.getsize(html_path)
# 2. Post a short message pointing at it.
body = build_summary(report, filename, test=test)
link = f"[{filename}]({uri})"
body = body.replace(f"`{filename}`", link)
payload, raw = _curl_json(auth + [
"-X", "POST", f"{ZULIP_SITE}/api/v1/messages",
"-d", "type=private",
"-d", f"to=[{CAPTAIN_USER_ID}]",
"--data-urlencode", f"content={body}",
])
if payload.get("result") == "success":
return True, (f"\u2705 Delivered to Zulip DM (user {CAPTAIN_USER_ID}), "
f"message id {payload.get('id')}, attachment {size} bytes at {uri}")
print(html_content)
return False, (f"\u274c Delivery FAILED at message post: {payload.get('msg') or raw[:160]} "
f"(uploaded {uri}; report persisted to {html_path})")
EMAIL_PASSWORD = "rgbuomwcydxwbszd"
GMAIL_EMAIL = "jtabiri@gmail.com"
server = smtplib.SMTP("smtp.gmail.com", 587)
server.starttls()
server.login(GMAIL_EMAIL, EMAIL_PASSWORD)
server.sendmail(FROM, [TO], msg.as_string())
server.quit()
return True, "✅ Email sent to jerome@sysloggh.com"
except Exception as e:
return False, f"❌ Email failed: {e}"
# ── Main ──
if __name__ == "__main__":
is_test = ("--test-email" in sys.argv) or ("--test-zulip" in sys.argv)
is_test = "--test-email" in sys.argv
print(f"{'🧪 TEST MODE' if is_test else '📊'} Collecting infrastructure data...")
report = collect()
if "--json" in sys.argv:
print(json.dumps(report, indent=2, default=str))
if PROBE_FAILURES:
for leg in PROBE_FAILURES:
print(f"PROBE FAILURE: {leg}", file=sys.stderr)
sys.exit(1)
sys.exit(0)
print(" Building dashboard...")
html = build_html(report)
print(f" report ready: {len(html)} chars of HTML (delivered as a file attachment)")
if is_test:
print(" Sending TEST message to the captain's Zulip DM...")
prefix = "🧪 TEST — "
print(" Sending test email...")
else:
print(" Sending to the captain's Zulip DM...")
ok, msg = send_zulip(html, report, test=is_test)
prefix = ""
print(" Sending email...")
ok, msg = send_email(html, subject_prefix=prefix)
print(f" {msg}")
# Show summary
if DEGRADED_LEGS:
print(f"\n⚠️ Degraded legs ({len(DEGRADED_LEGS)}):")
for leg in DEGRADED_LEGS:
print(f" - {leg}")
else:
print("\n✅ All legs fully credentialed")
# A failed send must exit non-zero; a degraded leg (no credential) must stay exit 0
if not ok:
sys.exit(1)
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
print(f"\n📋 Summary:")
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
@@ -902,9 +718,3 @@ if __name__ == "__main__":
for k,v in report.get('agents',{}).items():
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
print(f" Agents: {', '.join(agent_parts)}")
if PROBE_FAILURES:
print(f"\n❌ Probe failures ({len(PROBE_FAILURES)}):")
for leg in PROBE_FAILURES:
print(f" - {leg}")
sys.exit(1)
+8 -238
View File
@@ -101,6 +101,8 @@ GUESTS: list[Guest] = [
# amdpve (192.168.68.15)
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
@@ -108,8 +110,6 @@ GUESTS: list[Guest] = [
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
# minipve (192.168.68.12)
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="minipve",
access_method="pct-run", probe_target="tanko (CT 112, minipve)"),
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
@@ -153,25 +153,6 @@ GPU_HOSTS = [
CONNECT_TIMEOUT = 5
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
# Host filesystem thresholds (from contract)
HOST_THRESHOLDS = {
"WARN": 85,
"AMBER": 90,
"RED": 95,
}
# State file path (absolute, so execution context doesn't matter)
STATE_FILE = pathlib.Path(__file__).resolve().parent.parent / "state" / "host-disk-bands.json"
# PVE nodes to probe for host filesystems
HOST_NODES = [
{"hostname": "acerpve", "ip": "192.168.68.9"},
{"hostname": "amdpve", "ip": "192.168.68.15"},
{"hostname": "storepve", "ip": "192.168.68.6"},
{"hostname": "minipve", "ip": "192.168.68.12"},
{"hostname": "ocupve", "ip": "192.168.68.5"},
]
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
"""Run a command and return (exit_code, stdout, stderr)."""
@@ -317,151 +298,6 @@ def scan_fleet() -> list[dict]:
return results
def classify_band(usage_pct: float) -> str:
"""Classify a percentage into a band."""
if usage_pct >= HOST_THRESHOLDS["RED"]:
return "HOST-RED"
elif usage_pct >= HOST_THRESHOLDS["AMBER"]:
return "HOST-AMBER"
elif usage_pct >= HOST_THRESHOLDS["WARN"]:
return "HOST-WARN"
else:
return "GREEN"
def probe_host_filesystems() -> tuple[list[dict], dict[str, str]]:
"""Probe host filesystems on all PVE nodes.
Returns:
- List of host filesystem results
- Dict of volume_key -> current_band (for state file)
"""
results = []
current_bands = {}
for node in HOST_NODES:
ip = node["ip"]
hostname = node["hostname"]
# Probe df for host filesystems
probe_cmd = f'ssh {SSH_OPTS} root@{ip} "df -hP / /media/* tank 2>/dev/null | tail -n +2"'
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
if exit_code != 0:
results.append({
"target": f"{hostname} ({ip})",
"hostname": hostname,
"ip": ip,
"reachable": False,
"volumes": [],
"probe_cmd": probe_cmd,
"failure_kind": "ssh-error",
})
continue
# Parse df output and classify each volume
volumes = []
for line in stdout.splitlines():
parts = line.split()
if len(parts) < 6:
continue
dev, size, used, avail, pct_str, mount = parts[:6]
pct = float(pct_str.rstrip("%"))
band = classify_band(pct)
# Volume type classification
if mount.startswith("/media/"):
vol_type = "media"
elif mount == "/" or "pve" in dev:
vol_type = "host-root"
elif mount == "tank" or "tank" in mount:
vol_type = "pbs-datastore"
else:
vol_type = "other"
# Volume key for state file (host/volume)
volume_key = f"{hostname}/{mount}"
current_bands[volume_key] = band
volumes.append({
"mount": mount,
"device": dev,
"size": size,
"used": used,
"avail": avail,
"pct": pct,
"band": band,
"type": vol_type,
})
results.append({
"target": f"{hostname} ({ip})",
"hostname": hostname,
"ip": ip,
"reachable": True,
"volumes": volumes,
"probe_cmd": probe_cmd,
"failure_kind": None,
})
return results, current_bands
def read_state_file() -> Optional[dict[str, str]]:
"""Read the state file if it exists."""
if not STATE_FILE.exists():
return None
try:
with open(STATE_FILE) as f:
return json.load(f)
except (json.JSONDecodeError, IOError) as e:
print(f"⚠️ State file exists but unreadable: {e}", file=sys.stderr)
return {}
def write_state_file(bands: dict[str, str]) -> None:
"""Write the state file."""
STATE_FILE.parent.mkdir(parents=True, exist_ok=True)
try:
with open(STATE_FILE, "w") as f:
json.dump(bands, f, indent=2)
except IOError as e:
print(f"⚠️ State file write failed: {e}", file=sys.stderr)
def detect_transitions(current_bands: dict[str, str], prior_bands: Optional[dict[str, str]]) -> list[dict]:
"""Detect band transitions (current vs. prior)."""
if prior_bands is None:
# First run — no transitions, just establish baseline
return []
transitions = []
# Check for volumes that moved to a higher band (escalation)
for volume, current_band in current_bands.items():
prior_band = prior_bands.get(volume, "GREEN")
# Band ordering: GREEN < HOST-WARN < HOST-AMBER < HOST-RED
band_order = {"GREEN": 0, "HOST-WARN": 1, "HOST-AMBER": 2, "HOST-RED": 3}
if band_order[current_band] > band_order[prior_band]:
transitions.append({
"type": "escalation",
"volume": volume,
"from": prior_band,
"to": current_band,
})
elif band_order[current_band] < band_order[prior_band]:
transitions.append({
"type": "recovery",
"volume": volume,
"from": prior_band,
"to": current_band,
})
return transitions
def render_results(results: list[dict]) -> str:
"""Render scan results in human-readable format."""
lines = []
@@ -480,88 +316,22 @@ def render_results(results: list[dict]) -> str:
return "\n".join(lines)
def render_host_results(results: list[dict], transitions: list[dict], prior_bands: Optional[dict[str, str]]) -> str:
"""Render host filesystem results in human-readable format."""
lines = []
lines.append("")
lines.append("=== Host Filesystem Bands ===")
lines.append("")
# Render transitions first (they're the actionable alerts)
if prior_bands is None:
lines.append(" (first run — recording baseline, no alerts)")
elif not transitions:
lines.append(" (no band changes since last scan)")
else:
for t in transitions:
volume, from_band, to_band = t["volume"], t["from"], t["to"]
if t["type"] == "escalation":
lines.append(f" ⚠️ {volume}: {from_band} → {to_band} (ESCALATION)")
else:
lines.append(f" ✅ {volume}: {from_band} → {to_band} (RECOVERY)")
# Render all volumes with their bands
lines.append("")
for node_result in results:
if not node_result["reachable"]:
lines.append(f" ❌ {node_result['target']}: UNREACHABLE ({node_result['failure_kind']})")
continue
lines.append(f" {node_result['target']}:")
for vol in node_result["volumes"]:
lines.append(f" {vol['mount']} ({vol['type']}): {vol['pct']}% ({vol['used']}/{vol['size']}, {vol['avail']} free) -> {vol['band']}")
return "\n".join(lines)
def main() -> int:
import argparse
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests and host filesystems.")
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
ap.add_argument("--json", action="store_true", help="machine-readable output")
ap.add_argument("--hosts-only", action="store_true", help="scan host filesystems only")
ap.add_argument("--guests-only", action="store_true", help="scan guests only (skip host filesystems)")
args = ap.parse_args()
# Scan guests (unless --hosts-only)
guest_results = []
if not args.hosts_only:
guest_results = scan_fleet()
# Scan host filesystems (unless --guests-only)
host_results = []
current_bands = {}
if not args.guests_only:
host_results, current_bands = probe_host_filesystems()
# Read prior state and detect transitions
prior_bands = read_state_file()
transitions = detect_transitions(current_bands, prior_bands)
# Write new state
write_state_file(current_bands)
else:
prior_bands = None
transitions = []
results = scan_fleet()
if args.json:
# JSON output
output = {
"guests": guest_results,
"hosts": host_results,
"transitions": transitions,
"prior_bands": prior_bands,
}
print(json.dumps(output, indent=2))
print(json.dumps(results, indent=2))
else:
# Human-readable output
if guest_results:
print(render_results(guest_results))
if host_results:
print(render_host_results(host_results, transitions, prior_bands))
print(render_results(results))
# Exit 0 if all probed (reachable or not), 1 if any probe error
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
# (a probe error means the probe itself failed, not just that the guest was unreachable)
return 0
-227
View File
@@ -1,227 +0,0 @@
#!/usr/bin/env python3
"""Health-log freshness watchdog (dead-man's-switch).
Absence of logs raises no alarm. On 2026-09-26 the `gpu/` directory of
SyslogSolution/health-logs went silent for 12 days unnoticed, because the only
thing that would have noticed was the job that had stopped. This check lives on
a DIFFERENT host (CT 100) from the producers, so a producer host that is dead,
or a schedule that was deleted, still raises an alarm.
It asks the Gitea API for the newest commit touching each watched directory and
fails when that commit is older than the directory's threshold.
Usage:
health-log-freshness.py # check every watched directory
health-log-freshness.py --json # machine-readable output
Exit: 0 = all fresh, 1 = at least one stale or unreachable, 2 = check could not run.
Environment:
GITEA_URL default https://git.sysloggh.net
GITEA_TOKEN API token (falls back to GITEA_PAT, then to the token
embedded in ~/.git-credentials for that host)
HEALTH_LOG_MAX_AGE_H override thresholds, e.g. HEALTH_LOG_MAX_AGE_GPU=12
"""
from __future__ import annotations
import base64
import json
import os
import re
import sys
import urllib.error
import urllib.request
from datetime import datetime, timezone
REPO = "SyslogSolution/health-logs"
GITEA_URL = os.environ.get("GITEA_URL", "https://git.sysloggh.net").rstrip("/")
# dir -> (threshold_hours, why)
# Thresholds are sized for the producer cadence plus one missed run:
# gpu/ runs every 6h -> 12h tolerates one miss, catches a second
# litellm/ runs every 6h -> 18h (it is proven healthy; a looser bound avoids
# noise while still catching a real stop)
WATCHED: dict[str, tuple[float, str]] = {
"gpu": (12.0, "gpu-self-heal.py, CT116 cron 2 */6 * * *"),
"litellm": (18.0, "litellm-health-check.sh, CT116 cron 0 */6 * * *"),
}
def _threshold(directory: str, default: float) -> float:
"""Allow HEALTH_LOG_MAX_AGE_<DIR> to override a threshold.
Documented override; used both operationally (tighten a bound while
investigating) and in tests (force a stale verdict deterministically).
"""
raw = os.environ.get(f"HEALTH_LOG_MAX_AGE_{directory.upper()}")
if raw is None:
return default
try:
return float(raw)
except ValueError:
print(f"WARN: ignoring non-numeric HEALTH_LOG_MAX_AGE_{directory.upper()}={raw!r}")
return default
# pm2/ is deliberately NOT watched: pm2-self-heal.prose.md claimed a health-logs
# posting that its executor never implemented, and the claim was retired on
# 2026-09-26 in favour of the contract-runner's per-run logs. The directory is
# left as historical evidence, not as a live obligation.
def _auth_candidates() -> list[str]:
"""Authorization header values to try, in order.
Gitea accepts either an API token (``token <tok>``) or HTTP basic auth, and
the value held under ``GITEA_TOKEN`` is not reliably an API token - on this
host it is the git account's *password*, which sent as a token returns HTTP
401. So return every candidate and let the caller use the first that works,
rather than guessing and failing.
"""
candidates: list[str] = []
for var in ("GITEA_TOKEN", "GITEA_PAT"):
if os.environ.get(var):
candidates.append(f"token {os.environ[var]}")
candidates.append(
"Basic "
+ base64.b64encode(
f"{_git_user()}:{os.environ[var]}".encode()
).decode()
)
# ~/.git-credentials may hold this Gitea under several hostnames (the public
# name and the internal IP both appear in this fleet), so accept any of them
# rather than filtering on the configured host - that filter produced zero
# candidates whenever GITEA_URL pointed at the internal address.
try:
with open(os.path.expanduser("~/.git-credentials")) as fh:
for line in fh:
m = re.match(r"https://([^:]+):([^@]+)@", line.strip())
if m:
raw = f"{m.group(1)}:{m.group(2)}".encode()
candidates.append("Basic " + base64.b64encode(raw).decode())
except OSError:
pass
seen, out = set(), []
for c in candidates:
if c not in seen:
seen.add(c)
out.append(c)
return out
def _git_user() -> str:
return os.environ.get("GITEA_USER", "abiba-bot")
def newest_commit_iso(directory: str, auth: list[str] | None) -> tuple[str | None, str]:
"""Return (iso_timestamp, detail) for the newest commit touching `directory`."""
url = (
f"{GITEA_URL}/api/v1/repos/{REPO}/commits"
f"?path={directory}&limit=1&stat=false"
)
req = urllib.request.Request(url, headers={"Accept": "application/json"})
headers = list(auth or [])
last = "no credential"
for i, hdr in enumerate(headers or [None]):
r = urllib.request.Request(url, headers={"Accept": "application/json"})
if hdr:
r.add_header("Authorization", hdr)
try:
with urllib.request.urlopen(r, timeout=20) as resp:
data = json.loads(resp.read().decode("utf-8", "replace"))
break
except urllib.error.HTTPError as exc:
last = f"HTTP {exc.code}"
if exc.code not in (401, 403):
return None, last
except Exception as exc: # noqa: BLE001
return None, repr(exc)
else:
return None, last
if not data:
return None, "no commits"
commit = data[0].get("commit", {})
when = (
(commit.get("committer") or {}).get("date")
or (commit.get("author") or {}).get("date")
)
msg = (commit.get("message") or "").splitlines()[0][:60]
return when, msg
def main() -> int:
as_json = "--json" in sys.argv
auth = _auth_candidates()
now = datetime.now(timezone.utc)
failures: list[str] = []
report: dict[str, dict] = {}
if not auth:
print("FAIL: no Gitea credential available (GITEA_TOKEN/GITEA_PAT/~/.git-credentials)")
return 2
for directory, (default_age_h, why) in WATCHED.items():
max_age_h = _threshold(directory, default_age_h)
when, detail = newest_commit_iso(directory, auth)
entry: dict = {"directory": directory, "producer": why, "max_age_h": max_age_h}
if when is None:
entry.update(ok=False, reason=f"could not read newest commit: {detail}")
failures.append(
f"health-logs/{directory}/ could not be read ({detail}) — "
f"a directory that cannot be read is indistinguishable from one that stopped"
)
else:
try:
ts = datetime.fromisoformat(when.replace("Z", "+00:00"))
except ValueError:
entry.update(ok=False, reason=f"unparseable timestamp {when!r}")
failures.append(f"health-logs/{directory}/ timestamp unparseable: {when!r}")
report[directory] = entry
continue
age_h = (now - ts).total_seconds() / 3600.0
stale = age_h > max_age_h
entry.update(
ok=not stale,
newest=when,
age_h=round(age_h, 2),
newest_commit=detail,
)
if stale:
entry["reason"] = f"stale: {age_h:.1f}h > {max_age_h}h"
failures.append(
f"health-logs/{directory}/ is STALE: newest entry {when} "
f"({age_h:.1f}h old, limit {max_age_h}h) from {why}"
)
report[directory] = entry
if as_json:
print(json.dumps({"checked": report, "failures": failures}, indent=2))
return 1 if failures else 0
print("Health-log freshness — dead-man's-switch")
print("=" * 66)
for directory, entry in report.items():
mark = "✅" if entry.get("ok") else "❌"
if entry.get("newest"):
print(
f"{mark} health-logs/{directory}/ newest {entry['newest']} "
f"({entry['age_h']}h old, limit {entry['max_age_h']}h)"
)
print(f" last commit: {entry.get('newest_commit')}")
else:
print(f"{mark} health-logs/{directory}/ {entry.get('reason')}")
print(f" producer: {entry['producer']}")
print("=" * 66)
if failures:
print("VERDICT: FAIL")
for f in failures:
print(f" - {f}")
return 1
print("VERDICT: PASS — every watched health-log directory is advancing")
return 0
if __name__ == "__main__":
sys.exit(main())
-234
View File
@@ -1,234 +0,0 @@
#!/bin/bash
# infrastructure-monitoring.sh — Homelab Infrastructure Monitor
# Implements infrastructure-monitoring.prose.md (check-health section)
#
# Legs: Grafana, Prometheus, LiteLLM, PVE API (5 nodes), GPU exporters,
# Docker Stats, PVE Exporter
#
# Design:
# - Every target, port, path, and expected status is defined in code
# - Liveness rule: any HTTP status = ALIVE for auth-gated/redirect endpoints;
# only connection failures (000/timeout) = probe-failed
# - Bare-200 rule: expected status must match exactly (200); anything else = alert
# - PVE API uses -k flag (self-signed certs), probes /api2/json/version
# - Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116, probed via SSH
# with one retry at longer timeout (25s connect, 30s max) to distinguish
# transient timeout from host-down
# - Non-zero exit naming every failed target; no "OK" summary when any leg failed
#
# Output shape per leg:
# ✅ <name>: alive
# 🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)
#
# Failure kinds: timeout | refused | tls | unexpected:<code> (printed in the failure line)
set -uo pipefail
# ── Configuration (documented in infrastructure-monitoring.prose.md) ────────
# Change these in ONE place; test_infra_monitoring.sh asserts against these.
GRAFANA_HOST="192.168.68.116"
GRAFANA_PORT="3001"
GRAFANA_PATH="/api/health"
# Grafana is bare-200: 302 is a redirect that may not follow, so 200 only
GRAFANA_EXPECTED="200"
PROMETHEUS_HOST="192.168.68.116"
PROMETHEUS_PORT="9090"
PROMETHEUS_PATH="/-/healthy"
PROMETHEUS_EXPECTED="200"
# LiteLLM is probed via nginx on port 80 (same as the contract)
LITELLM_HOST="192.168.68.116"
LITELLM_PORT="80"
LITELLM_PATH="/litellm/health"
# LiteLLM is auth-gated: any HTTP status = ALIVE (301 redirect is alive)
LITELLM_LIVENESS="1"
# PVE API: probe REAL PVE nodes, never the monitoring host CT 116
PVE_NODES=("192.168.68.9" "192.168.68.12" "192.168.68.6" "192.168.68.15" "192.168.68.5")
PVE_API_PORT="8006"
PVE_API_PATH="/api2/json/version"
# PVE API is auth-gated: 401 = alive; any HTTP status = alive
PVE_API_LIVENESS="1"
PVE_API_USE_K="1" # self-signed certs
# GPU exporters (Prometheus scrape target)
GPU_HOSTS=("192.168.68.8" "192.168.68.110" "192.168.68.15")
GPU_PORT="9400"
GPU_PATH="/metrics"
GPU_EXPECTED="200"
# Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116
DOCKER_STATS_PORT="9324" # harness-docker-stats (docker_container_* metrics)
PVE_EXPORTER_PORT="9221" # harness-pve-exporter (5 pve_* metrics)
CT116_SSH_HOST="192.168.68.116"
# Both are bare-200: 404 = container not yet started
DOCKER_STATS_EXPECTED="200|404"
PVE_EXPORTER_EXPECTED="200|404"
# ── Probe Functions ─────────────────────────────────────────────────────────
# probe_http <host> <port> <path> <expected_pattern> [use_k] [ssh_host] [scheme] [liveness]
# Returns 0 if probe succeeds (matches expected or liveness), 1 if probe-failed.
# Prints the result line.
#
# FIX C1: The kind value is computed and printed in the failure line.
# FIX C2: SSH retry logic is in the first attempt branch (not unreachable).
LAST_KIND=""
probe_http() {
local host="$1" port="$2" path="$3" expected="$4"
local use_k="${5:-}" ssh_host="${6:-}" scheme="${7:-http}" liveness="${8:-0}"
local url="${scheme}://${host}:${port}${path}"
local code="" kind=""
LAST_KIND=""
# Single invocation that captures both output and status
if [ -n "$ssh_host" ]; then
out=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} ${url}" 2>/dev/null)
rc=$?
else
out=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} "$url" 2>/dev/null)
rc=$?
fi
code=$(printf '%s' "$out" | tr -d '[:space:]')
# Classify failure kind and retry if needed
if [ -z "$code" ] || [ "$code" = "000" ]; then
# Distinguish timeout from TLS error from refused
case "$rc" in
35|51|58|59|60|77|83) kind="tls" ;;
*) kind="timeout" ;;
esac
# Retry once at longer timeout (25s connect, 30s max)
if [ -n "$ssh_host" ]; then
code=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} ${url}" 2>/dev/null)
rc=$?
code=$(printf '%s' "$code" | tr -d '[:space:]')
else
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} "$url" 2>/dev/null)
rc=$?
code=$(printf '%s' "$code" | tr -d '[:space:]')
# On retry, classify: still 000 = keep existing kind (or timeout if empty), unexpected status = refused
if [ -z "$code" ] || [ "$code" = "000" ]; then
[ -z "$kind" ] && kind="timeout"
elif ! echo "$code" | grep -qE "^(${expected})$"; then
kind="refused"
fi
fi
fi
# Check result
if [ -n "$code" ] && [ "$code" != "000" ]; then
if [ "$liveness" = "1" ]; then
# Any HTTP status = ALIVE for auth-gated/redirect endpoints
return 0
else
# Bare-200 or specific expected pattern
if echo "$code" | grep -qE "^(${expected})$"; then
return 0
else
kind="unexpected:$code"
LAST_KIND="$kind"
return 1
fi
fi
else
[ -z "$kind" ] && kind="refused"
LAST_KIND="$kind"
return 1
fi
}
# ── Main ────────────────────────────────────────────────────────────────────
FAILED=()
FAILED_KIND=()
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
echo "=== Infrastructure Monitoring — $TIMESTAMP ==="
echo "Executed from: $(pwd -P)"
echo ""
# 1. Grafana (CT 116 :3001 /api/health) — bare-200
if probe_http "$GRAFANA_HOST" "$GRAFANA_PORT" "$GRAFANA_PATH" "$GRAFANA_EXPECTED"; then
echo " ✅ Grafana: alive"
else
echo " 🔴 Grafana: probe-failed: ${GRAFANA_HOST}:${GRAFANA_PORT} (expected ${GRAFANA_EXPECTED}) (<${LAST_KIND}>)"
FAILED+=("grafana")
fi
# 2. Prometheus (CT 116 :9090 /-/healthy) — bare-200
if probe_http "$PROMETHEUS_HOST" "$PROMETHEUS_PORT" "$PROMETHEUS_PATH" "$PROMETHEUS_EXPECTED"; then
echo " ✅ Prometheus: alive"
else
echo " 🔴 Prometheus: probe-failed: ${PROMETHEUS_HOST}:${PROMETHEUS_PORT} (expected ${PROMETHEUS_EXPECTED}) (<${LAST_KIND}>)"
FAILED+=("prometheus")
fi
# 3. LiteLLM (CT 116 :80/litellm/health via nginx) — liveness (any HTTP = alive)
if probe_http "$LITELLM_HOST" "$LITELLM_PORT" "$LITELLM_PATH" "" "" "" "http" "$LITELLM_LIVENESS"; then
echo " ✅ LiteLLM: alive"
else
echo " 🔴 LiteLLM: probe-failed: ${LITELLM_HOST}:${LITELLM_PORT}${LITELLM_PATH} (any-HTTP liveness) (<${LAST_KIND}>)"
FAILED+=("litellm")
fi
# 4. PVE API (5 real nodes :8006 /api2/json/version, -k, liveness)
PVE_FAILED=()
for node in "${PVE_NODES[@]}"; do
if probe_http "$node" "$PVE_API_PORT" "$PVE_API_PATH" "" "$PVE_API_USE_K" "" "https" "$PVE_API_LIVENESS"; then
echo " ✅ PVE API ${node}: alive"
else
echo " 🔴 PVE API ${node}: probe-failed: ${node}:${PVE_API_PORT} (any-HTTP liveness, -k for self-signed) (<${LAST_KIND}>)"
PVE_FAILED+=("$node")
fi
done
if [ ${#PVE_FAILED[@]} -gt 0 ]; then
FAILED+=("pve-api: ${PVE_FAILED[*]}")
fi
# 5. GPU exporters (:9400/metrics) — bare-200
GPU_FAILED=()
for host in "${GPU_HOSTS[@]}"; do
if probe_http "$host" "$GPU_PORT" "$GPU_PATH" "$GPU_EXPECTED"; then
echo " ✅ GPU exporter ${host}: alive"
else
echo " 🔴 GPU exporter ${host}: probe-failed: ${host}:${GPU_PORT} (expected 200) (<${LAST_KIND}>)"
GPU_FAILED+=("$host")
fi
done
if [ ${#GPU_FAILED[@]} -gt 0 ]; then
FAILED+=("gpu-exporters: ${GPU_FAILED[*]}")
fi
# 6. Docker Stats (CT 116 :9324, 127.0.0.1 via SSH) — 200|404
if probe_http "127.0.0.1" "$DOCKER_STATS_PORT" "/" "$DOCKER_STATS_EXPECTED" "" "$CT116_SSH_HOST"; then
echo " ✅ Docker Stats: alive"
else
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:${DOCKER_STATS_PORT} (expected 200|404) (<${LAST_KIND}>)"
FAILED+=("docker-stats")
fi
# 7. PVE Exporter (CT 116 :9221, 127.0.0.1 via SSH) — 200|404
if probe_http "127.0.0.1" "$PVE_EXPORTER_PORT" "/" "$PVE_EXPORTER_EXPECTED" "" "$CT116_SSH_HOST"; then
echo " ✅ PVE Exporter: alive"
else
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:${PVE_EXPORTER_PORT} (expected 200|404) (<${LAST_KIND}>)"
FAILED+=("pve-exporter")
fi
# ── Summary ─────────────────────────────────────────────────────────────────
echo ""
if [ ${#FAILED[@]} -eq 0 ]; then
echo " ✅ All legs OK"
exit 0
else
for f in "${FAILED[@]}"; do
echo " 🔴 FAILED: $f"
done
exit 1
fi
+36 -156
View File
@@ -35,12 +35,7 @@ def run_command(cmd, timeout=15):
return 1, "", str(e)
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
"""Probe HTTP endpoint and return (status_code, failure_kind)
Returns:
(code, None) if successful or HTTP response received
(000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc.
"""
"""Probe HTTP endpoint and return status code"""
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
@@ -52,64 +47,15 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
cmd += " -L"
cmd += " '" + url + "'"
try:
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0:
# Check if this is a timeout from run_command (rc=1, stderr="TIMEOUT")
if rc == 1 and stderr == "TIMEOUT":
return (000, "timeout after " + str(timeout) + "s")
# Otherwise, determine failure kind from curl exit code
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
elif rc == 28:
return (000, "timeout after " + str(timeout) + "s")
elif rc == 7:
return (000, "connection refused")
elif rc == 6:
return (000, "dns failure")
elif rc == 35:
return (000, "ssl error")
elif rc == 52:
return (000, "empty response")
else:
return (000, "curl exit " + str(rc))
return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response")
except subprocess.TimeoutExpired:
return (000, "timeout after " + str(timeout) + "s")
def check_host_health(host_ip):
"""Check if the GPU host's llama-chat-api health endpoint is reachable
Returns: (healthy: bool, detail: str)
"""
code, _ = probe_http("http://" + host_ip + ":8080/health", timeout=10)
if code == 200:
return True, "host healthy (200)"
elif code == 000:
return False, "host unreachable (timeout or refused)"
else:
return False, "host unhealthy (HTTP " + str(code) + ")"
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
cmd = "curl -s -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
cmd += " '" + url + "'"
rc, stdout, stderr = run_command(cmd, timeout)
# Return first 200 chars, single line
body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else ""
return body
if rc != 0 and "TIMEOUT" not in stderr:
return 000 # Connection failed
return int(stdout) if stdout.isdigit() else 000
def check_liveliness():
"""Step 1: Liveliness probe"""
code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
def check_containers():
@@ -132,86 +78,33 @@ def check_model_probes():
results = []
# Host health mapping: model -> host IP
model_hosts = {
"gpu-dense": "192.168.68.8", # RTX 3090
"gpu-vision": "192.168.68.110", # RTX 5070
"strix-moe": "192.168.68.15" # Strix Halo
}
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
host_ip = model_hosts[model]
# Single-host aliases: 30s initial timeout, retry once at 90s on failure
# Worst-case prefill ~76s, so 90s retry ensures we cover it
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
# Single-host aliases: 30s timeout
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
first_kind = None
if code == 000 and failure_kind:
first_kind = failure_kind
time.sleep(1)
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=90)
if code == 000 and failure_kind:
# Both attempts failed - check host health to distinguish busy from down
host_healthy, host_detail = check_host_health(host_ip)
if host_healthy:
results.append((model, False, "busy (completion timed out after retry; " + host_detail + ")"))
else:
# Host unreachable - report both kinds
if first_kind:
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
else:
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
elif code == 200:
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
elif code in (401, 403):
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
alias = "monitor-20260813"
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
# Pool alias (syslog-auto): 60s timeout, retry once on 000
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
if code == 000:
# Retry once with same timeout
time.sleep(1)
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)"))
elif code in (401, 403):
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
alias = "monitor-20260813"
results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
return results
@@ -238,33 +131,33 @@ def check_admin_key_list():
# Try to parse the response
try:
data = json.loads(stdout)
# Response is a dict with "keys" (paginated list) and "total_count" fields
# Response is a dict with "keys" field
if isinstance(data, dict) and "keys" in data:
key_count = data.get("total_count", len(data["keys"]))
key_count = len(data["keys"])
elif isinstance(data, list):
key_count = len(data)
else:
key_count = 0
if key_count == 0:
return "Admin Key List", False, "admin-call-failed (empty response)"
return "Admin Key List", True, str(key_count) + " total (" + str(len(data.get("keys", []) if isinstance(data, dict) else data)) + " on page 1)" if isinstance(data, dict) else str(key_count) + " total"
return "Admin Key List", True, str(key_count) + " keys"
except Exception as e:
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
def check_github_status():
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
code, _ = probe_http("https://status.github.com/api/status.json", timeout=15)
code = probe_http("https://status.github.com/api/status.json", timeout=15)
# GitHub status API returns 301 redirect, which is expected behavior
return "GitHub Status", code == 301, str(code)
def check_prometheus():
"""Step 4: Prometheus health"""
code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
def check_grafana():
"""Step 9: Grafana health"""
code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
def check_docker_stats():
@@ -288,7 +181,6 @@ def main():
print("")
all_pass = True
degraded = [] # Track degraded (busy) checks
# Run all checks
checks = [
@@ -308,15 +200,9 @@ def main():
# Model probes
model_results = check_model_probes()
for name, passed, detail in model_results:
# Check if this is a busy (degraded) verdict
if not passed and detail.startswith("busy "):
status = "⚠️"
degraded.append(name)
else:
status = "✅" if passed else "❌"
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
# Only set all_pass=False for real failures (not busy)
if not passed and not detail.startswith("busy "):
if not passed:
all_pass = False
# Admin key list
@@ -342,16 +228,10 @@ def main():
print("")
if all_pass:
if degraded:
print("✅ All checks passed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
else:
print("✅ All checks passed")
print("✅ All checks passed")
return 0
else:
if degraded:
print("❌ Some checks failed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
else:
print("❌ Some checks failed")
print("❌ Some checks failed")
return 1
if __name__ == "__main__":
+1 -1
View File
@@ -12,6 +12,7 @@ set -euo pipefail
declare -A CT_NODES=(
# amdpve (192.168.68.15)
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
[112]=amdpve # tanko
[113]=amdpve # baggy
[115]=amdpve # scottdenya
[120]=amdpve # adguard2 (added 2026-09-12)
@@ -20,7 +21,6 @@ declare -A CT_NODES=(
[102]=minipve # adguard (was acerpve)
[104]=minipve # authentik
[110]=minipve # gitea
[112]=minipve # tanko (was amdpve — migrated 2026-09-27; live per pvesh)
[116]=minipve # syslog-api
[119]=minipve # infisical-vault
# storepve (192.168.68.6)
+2 -68
View File
@@ -1,7 +1,7 @@
#!/bin/bash
# pm2-self-heal — hourly PM2 process check
# Part of the pm2-self-heal prose contract
# Alerts via Telegram (primary) and Zulip DM (secondary, abiba-zulip restored 2026-08-03)
# Alerts via Telegram (abiba-zulip decommissioned 2026-07-04)
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
@@ -46,75 +46,9 @@ if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
fi
fi
# Check abiba-zulip (live Zulip bridge, heartbeating)
ZULIP_LINE=$(echo "$STATUS" | grep "abiba-zulip")
ZULIP_STATUS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $10}' | xargs)
ZULIP_RESTARTS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $9}' | xargs)
if [ "$ZULIP_STATUS" != "online" ]; then
pm2 restart abiba-zulip > /dev/null 2>&1
sleep 3
ZULIP_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-zulip")
ZULIP_STATUS2=$(echo "$ZULIP_LINE2" | awk -F'│' '{print $10}' | xargs)
if [ "$ZULIP_STATUS2" = "online" ]; then
msg="⚠️ abiba-zulip was **$ZULIP_STATUS** → restarted to online"
ALERTS="${ALERTS}${msg}\n"
else
msg="🚨 abiba-zulip **failed restart** (was $ZULIP_STATUS, still $ZULIP_STATUS2)"
ALERTS="${ALERTS}${msg}\n"
fi
elif [ "$ZULIP_RESTARTS" -gt 5 ]; then
msg="⚠️ abiba-zulip has **$ZULIP_RESTARTS** restarts (high count)"
ALERTS="${ALERTS}${msg}\n"
fi
# Check gitea-runner
GITEA_LINE=$(echo "$STATUS" | grep "gitea-runner")
GITEA_STATUS=$(echo "$GITEA_LINE" | awk -F'│' '{print $10}' | xargs)
GITEA_RESTARTS=$(echo "$GITEA_LINE" | awk -F'│' '{print $9}' | xargs)
if [ "$GITEA_STATUS" != "online" ]; then
pm2 restart gitea-runner > /dev/null 2>&1
sleep 3
GITEA_LINE2=$(pm2 status --no-color 2>/dev/null | grep "gitea-runner")
GITEA_STATUS2=$(echo "$GITEA_LINE2" | awk -F'│' '{print $10}' | xargs)
if [ "$GITEA_STATUS2" = "online" ]; then
msg="⚠️ gitea-runner was **$GITEA_STATUS** → restarted to online"
ALERTS="${ALERTS}${msg}\n"
else
msg="🚨 gitea-runner **failed restart** (was $GITEA_STATUS, still $GITEA_STATUS2)"
ALERTS="${ALERTS}${msg}\n"
fi
elif [ "$GITEA_RESTARTS" -gt 5 ]; then
msg="⚠️ gitea-runner has **$GITEA_RESTARTS** restarts (high count)"
ALERTS="${ALERTS}${msg}\n"
fi
# Check zulip-watchdog
WATCHDOG_LINE=$(echo "$STATUS" | grep "zulip-watchdog")
WATCHDOG_STATUS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $10}' | xargs)
WATCHDOG_RESTARTS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $9}' | xargs)
if [ "$WATCHDOG_STATUS" != "online" ]; then
pm2 restart zulip-watchdog > /dev/null 2>&1
sleep 3
WATCHDOG_LINE2=$(pm2 status --no-color 2>/dev/null | grep "zulip-watchdog")
WATCHDOG_STATUS2=$(echo "$WATCHDOG_LINE2" | awk -F'│' '{print $10}' | xargs)
if [ "$WATCHDOG_STATUS2" = "online" ]; then
msg="⚠️ zulip-watchdog was **$WATCHDOG_STATUS** → restarted to online"
ALERTS="${ALERTS}${msg}\n"
else
msg="🚨 zulip-watchdog **failed restart** (was $WATCHDOG_STATUS, still $WATCHDOG_STATUS2)"
ALERTS="${ALERTS}${msg}\n"
fi
elif [ "$WATCHDOG_RESTARTS" -gt 5 ]; then
msg="⚠️ zulip-watchdog has **$WATCHDOG_RESTARTS** restarts (high count)"
ALERTS="${ALERTS}${msg}\n"
fi
# Log check
{
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS zulip=$ZULIP_STATUS gitea=$GITEA_STATUS watchdog=$WATCHDOG_STATUS alerts=${ALERTS:+yes}"
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS alerts=${ALERTS:+yes}"
[ -n "$ALERTS" ] && echo "$ALERTS"
} >> "$LOG"
+2 -2
View File
@@ -47,8 +47,8 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (5 nodes):**
- amdpve (192.168.68.15): kagentz, baggy, scottdenya, adguard2
- minipve (192.168.68.12): abiba, tanko, adguard, authentik, gitea, syslog-api, infisical-vault
- amdpve (192.168.68.15): kagentz, tanko, baggy, scottdenya, adguard2
- minipve (192.168.68.12): abiba, adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
- acerpve (192.168.68.9): llm-gpu
- ocupve (192.168.68.5): ocu-llm
+1 -16
View File
@@ -135,22 +135,7 @@ fi
echo " Cross-contract: $WARNINGS total warnings across all checks"
# ── 4. Committed-credential scan ──
# The 2026-09-17 purge removed six live credentials that had sat in .md prose
# and scripts for weeks. This step makes that class of commit FAIL the gate
# instead of printing a warning. Patterns: scripts/secret-patterns.tsv.
# Only deliberate synthetic examples may be listed in scripts/secret-allowlist.tsv,
# each with a reason. Run `bash scripts/secret-scan.sh --staged` before committing.
echo ""
echo "── 4. Secret scan (committed credentials) ──"
if bash scripts/secret-scan.sh; then
echo " ✅ No committed credentials"
else
echo " ❌ COMMITTED CREDENTIAL DETECTED"
FAILED=1
fi
# ── 5. Summary ──
# ── 4. Summary ──
echo ""
echo "═══════════════════════════════════"
if [ $FAILED -eq 1 ]; then
-166
View File
@@ -1,166 +0,0 @@
#!/bin/bash
# proxmox-monitor.sh — Proxmox Cluster + Docker Monitoring Health Check
# Implements proxmox-monitor.prose.md (check-health section)
#
# Legs: Prometheus, Grafana, Docker Stats Exporter, PVE Exporter, PBS GC
# All legs must return 200 for healthy status.
#
# Run: bash scripts/proxmox-monitor.sh
# Exits 0 if all probes pass, 1 if any fails.
set -uo pipefail
CT116_HOST="192.168.68.116"
FAILED=()
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
echo "=== Proxmox Monitor — $TIMESTAMP ==="
echo "Executed from: $(pwd -P)"
echo ""
# 1. Prometheus health (bound to 0.0.0.0:9090 on .116)
PROM_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:9090/-/healthy 2>/dev/null)
PROM_CODE=$(printf '%s' "$PROM_CODE" | tr -d '[:space:]')
[ -n "$PROM_CODE" ] || PROM_CODE="000"
if [ "$PROM_CODE" = "200" ]; then
echo " ✅ Prometheus: alive"
else
echo " 🔴 Prometheus: probe-failed: ${CT116_HOST}:9090 (expected 200, got ${PROM_CODE})"
FAILED+=("prometheus")
fi
# 2. Grafana health (bound to 0.0.0.0:3001 on .116)
GRAF_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:3001/api/health 2>/dev/null)
GRAF_CODE=$(printf '%s' "$GRAF_CODE" | tr -d '[:space:]')
[ -n "$GRAF_CODE" ] || GRAF_CODE="000"
if [ "$GRAF_CODE" = "200" ]; then
echo " ✅ Grafana: alive"
else
echo " 🔴 Grafana: probe-failed: ${CT116_HOST}:3001 (expected 200, got ${GRAF_CODE})"
FAILED+=("grafana")
fi
# 3. Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — probe from .116 localhost)
DOCKER_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics" 2>/dev/null)
DOCKER_CODE=$(printf '%s' "$DOCKER_CODE" | tr -d '[:space:]')
[ -n "$DOCKER_CODE" ] || DOCKER_CODE="000"
if [ "$DOCKER_CODE" = "200" ]; then
echo " ✅ Docker Stats: alive"
else
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:9324 (expected 200, got ${DOCKER_CODE})"
FAILED+=("docker-stats")
fi
# 4. PVE Exporter (bound to 127.0.0.1:9221 on .116 — probe from .116 localhost)
PVE_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics" 2>/dev/null)
PVE_CODE=$(printf '%s' "$PVE_CODE" | tr -d '[:space:]')
[ -n "$PVE_CODE" ] || PVE_CODE="000"
if [ "$PVE_CODE" = "200" ]; then
echo " ✅ PVE Exporter: alive"
else
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:9221 (expected 200, got ${PVE_CODE})"
FAILED+=("pve-exporter")
fi
# 5. PBS GC liveness (storepve-datastore GC must have run within 48h)
# Use absolute path for pct to avoid PATH issues in non-interactive ssh
PBS_GC_OUTPUT=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192.168.68.6 \
"/sbin/pct exec 107 -- /sbin/proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null)
PBS_GC_OUTPUT=$(printf '%s' "$PBS_GC_OUTPUT" | tr -d '[:space:]')
[ -n "$PBS_GC_OUTPUT" ] || PBS_GC_OUTPUT="000"
if [ "$PBS_GC_OUTPUT" = "000" ]; then
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)"
FAILED+=("pbs-gc")
else
# Parse the JSON to get storepve-datastore's state with four distinct outcomes:
# 1. probe-failed: non-zero ssh status / empty / unparseable JSON
# 2. running: collection in progress (last-run-endtime absent or 0, but upid present)
# 3. stale: no completed run within 48h
# 4. healthy: completed within 48h
PBS_GC_RESULT=$(echo "$PBS_GC_OUTPUT" | python3 -c "
import sys, json
try:
data = json.load(sys.stdin)
for store in data:
if store['store'] == 'storepve-datastore':
endtime = store.get('last-run-endtime')
upid = store.get('upid')
pending = store.get('pending-bytes', 0)
# State 2: Running (collection in progress) — last-run-endtime absent while run is in progress
if (endtime is None or endtime == 0) and upid is not None:
print(f'running|{pending}')
break
# State 3: No completed run (never-run or stale)
if endtime is None or endtime == 0:
print(f'no-completed-run|{pending}')
break
# States 3 & 4: Completed (has endtime)
print(f'completed|{endtime}|{pending}')
break
else:
print(f'absent|0')
except json.JSONDecodeError:
print(f'unparseable|0')
" 2>/dev/null)
# Parse the state|endtime|pending format
PBS_GC_STATE=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1)
if [ "$PBS_GC_STATE" = "unparseable" ]; then
# State 1: probe-failed (unparseable JSON)
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)"
FAILED+=("pbs-gc")
elif [ "$PBS_GC_STATE" = "absent" ]; then
# State 1: probe-failed (store not found)
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (storepve-datastore not found)"
FAILED+=("pbs-gc")
elif [ "$PBS_GC_STATE" = "running" ]; then
# State 2: collection in progress — do NOT fail
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
echo " ⏳ PBS GC: running (started: in-progress, pending-bytes: ${PENDING_BYTES} B)"
elif [ "$PBS_GC_STATE" = "no-completed-run" ]; then
# State 3: no completed run within 48h
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
echo " 🔴 PBS GC: no completed run within 48h (pending-bytes: ${PENDING_BYTES} B)"
FAILED+=("pbs-gc")
else
# States 3 & 4: completed (has endtime)
LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f3)
# Convert epoch to age in hours
NOW_EPOCH=$(date -u +%s)
AGE_HOURS=$(( (NOW_EPOCH - LAST_RUN_ENDTIME) / 3600 ))
if [ $AGE_HOURS -gt 48 ]; then
# State 3: stale (no completed run within 48h)
echo " 🔴 PBS GC: stale — last completed run was ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)"
FAILED+=("pbs-gc")
else
# State 4: healthy (completed within 48h)
echo " ✅ PBS GC: healthy — last completed run ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)"
fi
fi
fi
# ── Summary ─────────────────────────────────────────────────────────────────
echo ""
if [ ${#FAILED[@]} -eq 0 ]; then
echo " ✅ All legs OK"
exit 0
else
for f in "${FAILED[@]}"; do
echo " 🔴 FAILED: $f"
done
exit 1
fi
-209
View File
@@ -1,209 +0,0 @@
#!/usr/bin/env bash
# revision-preflight.sh — prove the copy a contract is about to execute is the
# copy that is merged.
#
# Usage:
# revision-preflight.sh [options] <script-path> <clone-path>
#
# Options:
# --ref <ref> Ref to compare against (default: origin/master)
# --no-fetch Do not refresh the ref first (see FRESHNESS)
# --fetch-timeout <secs> Bound the default fetch (default: 20; 0 = no bound)
# --quiet Print nothing on success
# -h, --help Show this help
#
# Exit codes:
# 0 the executing script byte-matches <ref>:<repo-relative-path>
# 1 the copy is NOT the merged one -> REASON=mismatch:<class>
# 2 the check could not be performed -> REASON=cannot-verify:<class>
#
# Every non-zero exit prints one machine-readable line
# REASON=<class>
# followed by the human explanation. The two top-level classes are deliberately
# distinct: "I could not check" is a different situation from "this copy is
# wrong", and an operator must never have to guess which they are looking at.
#
# cannot-verify:fetch-failed the remote could not be reached (or timed out)
# cannot-verify:ref-unresolvable <ref> does not exist in the clone
# mismatch:path-absent the script does not exist in <ref>
# mismatch:content the script differs from <ref>
# mismatch:detached-head the clone is on a detached HEAD
# mismatch:clone-ahead local HEAD is ahead of <ref> (mid-review?)
#
# FRESHNESS
# A guard is only as good as the ref it compares against. On 2026-09-25 a
# stale local origin/master made an ancestry check on this fleet report
# "unlanded work" for a branch that had in fact merged, and it would equally
# have passed a stale script as current. So by default this guard FETCHES the
# remote before comparing, bounded by --fetch-timeout so a hung remote cannot
# block a scheduled contract. With --no-fetch it compares against whatever the
# local ref points at and says so out loud; it never silently assumes
# freshness.
#
# FAIL CLOSED
# An unresolvable path or ref is a FAILURE, never a warning. "Cannot verify"
# is precisely the state a stale or hand-edited copy produces, so treating it
# as success would defeat the guard. The original draft did exactly that: it
# resolved the master revision with
# `git show origin/master:$(basename "$SCRIPT")`, which drops the scripts/
# prefix, queries the repo root, fails, and exited 0 — passing a script that
# exists in no revision at all.
#
# WHICH CLONE
# Pass the clone the contract is actually executing from. See
# docs/contract-execution-pinning.md for which clone each contract pins.
set -euo pipefail
REF="origin/master"
FETCH=1
QUIET=0
FETCH_TIMEOUT="${REVISION_PREFLIGHT_FETCH_TIMEOUT:-20}"
usage() {
sed -n '2,55p' "$0" | sed 's/^# \{0,1\}//'
}
while [[ $# -gt 0 ]]; do
case "$1" in
--ref)
[[ $# -ge 2 ]] || { echo "revision-preflight: --ref needs a value" >&2; exit 2; }
REF="$2"; shift 2 ;;
--no-fetch) FETCH=0; shift ;;
--fetch-timeout)
[[ $# -ge 2 ]] || { echo "revision-preflight: --fetch-timeout needs a value" >&2; exit 2; }
FETCH_TIMEOUT="$2"; shift 2 ;;
--quiet) QUIET=1; shift ;;
-h|--help) usage; exit 0 ;;
--) shift; break ;;
-*) echo "revision-preflight: unknown option: $1" >&2; exit 2 ;;
*) break ;;
esac
done
if [[ $# -lt 2 ]]; then
usage >&2
exit 2
fi
SCRIPT="$1"
CLONE="$2"
say() { [[ $QUIET -eq 1 ]] || echo "$@" >&2; }
# refuse <class> <explanation...> -> the copy is not the merged one
refuse() {
local class="$1"; shift
echo "REASON=mismatch:${class}" >&2
echo "❌ revision-preflight: MISMATCH (${class}) — refusing to report from this copy" >&2
for line in "$@"; do echo " $line" >&2; done
exit 1
}
# unverifiable <class> <explanation...> -> the check could not be performed
unverifiable() {
local class="$1"; shift
echo "REASON=cannot-verify:${class}" >&2
echo "❌ revision-preflight: CANNOT VERIFY (${class}) — refusing to report unverified" >&2
for line in "$@"; do echo " $line" >&2; done
exit 2
}
# ── 1. inputs must exist ──────────────────────────────────────────────────────
if [[ ! -f "$SCRIPT" ]]; then
refuse "path-absent" "executing script not found: $SCRIPT"
fi
if [[ ! -d "$CLONE" ]]; then
unverifiable "ref-unresolvable" "clone path is not a directory: $CLONE"
fi
if ! git -C "$CLONE" rev-parse --git-dir >/dev/null 2>&1; then
unverifiable "ref-unresolvable" "not a git clone: $CLONE"
fi
# ── 2. resolve the repo-relative path (the original defect) ───────────────────
CLONE_ABS=$(cd "$CLONE" && pwd)
SCRIPT_ABS=$(cd "$(dirname "$SCRIPT")" && pwd)/$(basename "$SCRIPT")
case "$SCRIPT_ABS" in
"$CLONE_ABS"/*) REL="${SCRIPT_ABS#"$CLONE_ABS"/}" ;;
*) refuse "content" "script is outside the clone: $SCRIPT_ABS is not under $CLONE_ABS" ;;
esac
# ── 3. refresh the ref, bounded, so a hung remote cannot block a contract ────
if [[ $FETCH -eq 1 ]]; then
REMOTE="${REF%%/*}"
[[ "$REMOTE" == "$REF" ]] && REMOTE="origin"
FETCH_CMD=(git -C "$CLONE" fetch --quiet "$REMOTE")
if [[ "$FETCH_TIMEOUT" != "0" ]]; then
if ! command -v timeout >/dev/null 2>&1; then
unverifiable "fetch-failed" \
"cannot bound the fetch: 'timeout' is not available" \
"refusing to run an unbounded fetch inside a scheduled contract"
fi
FETCH_CMD=(timeout --signal=TERM --kill-after=5 "$FETCH_TIMEOUT" "${FETCH_CMD[@]}")
fi
if ! "${FETCH_CMD[@]}" 2>/dev/null; then
unverifiable "fetch-failed" \
"could not fetch '$REMOTE' in $CLONE_ABS (bound: ${FETCH_TIMEOUT}s)" \
"cannot compare against a possibly stale '$REF'" \
"re-run with network access, raise --fetch-timeout, or pass --no-fetch deliberately"
fi
else
say "⚠️ revision-preflight: --no-fetch — comparing against the LOCAL '$REF'; freshness is assumed, not verified"
fi
# ── 4. resolve the merged revision; unresolvable is a failure ────────────────
if ! git -C "$CLONE" rev-parse --verify --quiet "$REF" >/dev/null; then
unverifiable "ref-unresolvable" \
"ref '$REF' does not resolve in $CLONE_ABS" \
"the clone may never have fetched, or the ref name may be wrong"
fi
REF_COMMIT=$(git -C "$CLONE" rev-parse --short "$REF")
TMPFILE=$(mktemp)
trap 'rm -f "$TMPFILE"' EXIT
if ! git -C "$CLONE" show "$REF:$REL" > "$TMPFILE" 2>/dev/null; then
refuse "path-absent" \
"'$REL' does not exist in $REF ($REF_COMMIT)" \
"a path absent from $REF can never be a merged copy" \
"script: $SCRIPT_ABS"
fi
# ── 5. compare ───────────────────────────────────────────────────────────────
EXEC_SHA=$(sha256sum "$SCRIPT" | cut -d' ' -f1)
MERGED_SHA=$(sha256sum "$TMPFILE" | cut -d' ' -f1)
if [[ "$EXEC_SHA" != "$MERGED_SHA" ]]; then
DETAIL=("script: $SCRIPT_ABS"
"clone: $CLONE_ABS"
"executed: $EXEC_SHA"
"merged: $MERGED_SHA ($REF:$REL @ $REF_COMMIT)")
# Name WHY it differs: a detached HEAD or a branch legitimately ahead of the
# ref is a much more benign situation than a hand-edited file, and the
# operator must be able to tell them apart.
if ! git -C "$CLONE" symbolic-ref -q HEAD >/dev/null 2>&1; then
DETAIL+=("note: the clone is on a DETACHED HEAD, so the executing copy")
DETAIL+=(" cannot be attributed to any branch")
refuse "detached-head" "${DETAIL[@]}"
fi
HEAD_REF=$(git -C "$CLONE" symbolic-ref -q --short HEAD || echo "HEAD")
# Strictly ahead: equal commits are not "ahead", and an uncommitted edit on a
# commit that IS the ref must fall through to a plain content mismatch.
REF_OID=$(git -C "$CLONE" rev-parse "$REF" 2>/dev/null || echo "")
HEAD_OID=$(git -C "$CLONE" rev-parse HEAD 2>/dev/null || echo "")
if [[ -n "$REF_OID" && "$REF_OID" != "$HEAD_OID" ]] \
&& git -C "$CLONE" merge-base --is-ancestor "$REF" HEAD 2>/dev/null; then
AHEAD=$(git -C "$CLONE" rev-list --count "$REF..HEAD" 2>/dev/null || echo "?")
DETAIL+=("note: '$HEAD_REF' is AHEAD of $REF by $AHEAD commit(s)")
DETAIL+=(" (a legitimate mid-review state, not a hand-edited file)")
refuse "clone-ahead" "${DETAIL[@]}"
fi
DETAIL+=("branch: $HEAD_REF")
refuse "content" "${DETAIL[@]}"
fi
say "✅ revision-preflight: $REL matches $REF @ $REF_COMMIT ($EXEC_SHA)"
exit 0
-395
View File
@@ -1,395 +0,0 @@
#!/usr/bin/env python3
"""Agent-consumption layer in front of SearXNG + Firecrawl.
Multi-engine aggregation returns results with no dedupe, no filtering and no
reranking. Measured 2026-09-26 that put bestbuy.com and merriam-webster.com into
"best practices agent context management", and put four SEO blogs ABOVE the real
Proxmox forum threads on a precise technical query. Identical queries also ranked
differently between runs, so the fix has to be deterministic rather than
dependent on engine mood.
This module turns the raw result list into something an agent can actually use:
1. DEDUPE the same page arriving from several engines
2. DROP clear non-answers (homepages, shopping, dictionaries, logins)
3. DEMOTE config-listed low-authority hosts; PROMOTE primary sources
4. STABLE SORT so ordering is reproducible run to run
5. EXTRACT page text for the top N under an explicit character budget,
so one call returns usable material instead of a snippet
6. EMIT stable JSON with engine provenance
Policy lives in config/search-ranking.yaml, not in this file.
Usage:
search-agent-consume.py "query text" # JSON to stdout
search-agent-consume.py --no-extract "query" # ranking only, no Firecrawl
search-agent-consume.py --explain "query" # include drop/demote reasons
Exit: 0 ok, 1 no results survived filtering, 2 the layer could not run.
"""
from __future__ import annotations
import json
import os
import sys
import time
import urllib.parse
import urllib.request
from pathlib import Path
SEARXNG_URL = os.environ.get("SEARXNG_URL", "http://192.168.68.7:8888").rstrip("/")
FIRECRAWL_URL = os.environ.get("FIRECRAWL_URL", "http://192.168.68.7:3002").rstrip("/")
CONFIG_PATH = os.environ.get(
"SEARCH_RANKING_CONFIG",
str(Path(__file__).resolve().parent.parent / "config" / "search-ranking.yaml"),
)
HTTP_TIMEOUT = float(os.environ.get("SEARCH_CONSUME_TIMEOUT", "25"))
def _load_config() -> dict:
"""Load the ranking policy.
PyYAML is used when present; otherwise a tiny built-in parser handles the
flat lists in this specific file, so the layer never hard-fails on a host
without PyYAML.
"""
text = Path(CONFIG_PATH).read_text()
try:
import yaml # type: ignore
return yaml.safe_load(text)
except ImportError:
return _parse_flat_yaml(text)
def _parse_flat_yaml(text: str) -> dict:
"""Minimal fallback parser: top-level keys, nested one level, flat lists."""
import re
out: dict = {}
stack: list[tuple[int, dict]] = [(-1, out)]
section: dict | None = None
for raw in text.splitlines():
line = raw.split("#", 1)[0].rstrip()
if not line.strip():
continue
indent = len(line) - len(line.lstrip())
body = line.strip()
if body.startswith("- "):
if section is not None:
section.setdefault("_list", []).append(
body[2:].strip().strip("'\"")
)
continue
if ":" in body:
key, _, val = body.partition(":")
key, val = key.strip(), val.strip()
if val:
# write to the INNERMOST open section, not the document root
stack[-1][1][key] = _scalar(val)
section = None
else:
while stack and indent <= stack[-1][0]:
stack.pop()
parent = stack[-1][1]
new: dict = {}
parent[key] = new
stack.append((indent, new))
section = new
# flatten "_list" holders back into their parent as plain lists
def fix(node):
if isinstance(node, dict):
if set(node.keys()) == {"_list"}:
return node["_list"]
return {k: fix(v) for k, v in node.items()}
return node
return fix(out)
def _scalar(v: str):
if v.lower() in ("true", "false"):
return v.lower() == "true"
try:
return int(v)
except ValueError:
pass
try:
return float(v)
except ValueError:
pass
return v.strip("'\"")
# ── filtering ────────────────────────────────────────────────────────────────
def _host(url: str) -> str:
return (urllib.parse.urlparse(url).netloc or "").lower().split(":")[0]
def _registrable(host: str) -> str:
"""Best-effort registrable domain so sub.forum.proxmox.com matches proxmox.com."""
parts = host.split(".")
if len(parts) <= 2:
return host
# handle common two-label public suffixes
two = ".".join(parts[-2:])
if parts[-2] in ("co", "com", "org", "net", "ac", "gov") and len(parts) >= 3:
return ".".join(parts[-3:])
return two
def _host_in(host: str, domains) -> bool:
if not domains:
return False
reg = _registrable(host)
for d in domains:
d = str(d).lower()
if host == d or host.endswith("." + d) or reg == d:
return True
return False
def _normalise_url(url: str) -> str:
"""Strip tracking params and fragments so the same page dedupes."""
p = urllib.parse.urlparse(url)
q = [
(k, v)
for k, v in urllib.parse.parse_qsl(p.query, keep_blank_values=True)
if not k.lower().startswith(("utm_", "fbclid", "gclid", "mc_", "ref"))
]
path = p.path.rstrip("/") or "/"
return urllib.parse.urlunparse(
(p.scheme.lower(), p.netloc.lower(), path, "", urllib.parse.urlencode(q), "")
)
def non_answer_reason(result: dict, cfg: dict) -> str | None:
"""Return why this result is a non-answer, or None if it may be returned."""
na = cfg.get("non_answer", {}) or {}
url = result.get("url", "")
p = urllib.parse.urlparse(url)
host = _host(url)
path = p.path or ""
if _host_in(host, na.get("hosts")):
return "shopping_or_dictionary_host"
if na.get("host_root", True) and path in ("", "/"):
# A preferred host's front door may legitimately be the answer
# (a repo, a docs site). Everything else is navigational.
if not _host_in(host, cfg.get("prefer_domains")):
return "navigational_host_root"
low = url.lower()
for pat in na.get("path_patterns", []) or []:
if pat.lower() in low:
return f"path_pattern:{pat}"
qkeys = {k.lower() for k in (na.get("query_keys") or [])}
if qkeys & {k.lower() for k, _ in urllib.parse.parse_qsl(p.query)}:
return "search_or_shopping_query"
return None
def source_type(url: str, cfg: dict) -> str:
host = _host(url)
if _host_in(host, ["github.com", "gitlab.com", "codeberg.org", "sourceforge.net"]):
return "code"
if _host_in(host, ["stackoverflow.com", "stackexchange.com", "superuser.com",
"serverfault.com", "askubuntu.com"]):
return "qa"
if _host_in(host, ["forum.proxmox.com", "forum.", "discourse"]) or "forum." in host:
return "forum"
if _host_in(host, ["news.ycombinator.com", "lobste.rs", "reddit.com"]):
return "discussion"
if _host_in(host, cfg.get("prefer_domains")):
return "official"
if _host_in(host, cfg.get("demote_domains")):
return "content-farm"
return "web"
def rank(results: list[dict], cfg: dict) -> tuple[list[dict], list[dict]]:
"""Dedupe, drop non-answers, demote/ promote, stable sort.
Returns (kept, dropped) where dropped carries the reason, because a filter
nobody can audit is a filter nobody should trust.
"""
rank_cfg = cfg.get("ranking", {}) or {}
demote_pen = float(rank_cfg.get("demote_penalty", 1000))
prefer_bonus = float(rank_cfg.get("prefer_bonus", 100))
multi_bonus = float(rank_cfg.get("multi_engine_bonus", 25))
seen: dict[str, dict] = {}
dropped: list[dict] = []
for pos, r in enumerate(results):
url = r.get("url")
if not url:
continue
key = _normalise_url(url)
engine = r.get("engine", "?")
# 1. dedupe: same normalised URL from several engines
if key in seen:
seen[key].setdefault("engines", []).append(engine)
seen[key]["duplicate_of"] = True
continue
reason = non_answer_reason(r, cfg)
if reason:
dropped.append({"url": url, "reason": reason, "position": pos + 1})
continue
seen[key] = {
"title": (r.get("title") or "").strip(),
"url": url,
"engines": [engine],
"position": pos,
"score": 0.0,
}
kept = []
for item in seen.values():
host = _host(item["url"])
score = -float(item["position"]) # original order is the base signal
if _host_in(host, cfg.get("demote_domains")):
score -= demote_pen
if _host_in(host, cfg.get("prefer_domains")):
score += prefer_bonus
if len(item["engines"]) > 1:
score += multi_bonus * (len(item["engines"]) - 1)
item["score"] = round(score, 2)
item["host"] = host
item["source_type"] = source_type(item["url"], cfg)
kept.append(item)
# stable: score desc, then original position asc => reproducible run to run
kept.sort(key=lambda i: (-i["score"], i["position"]))
return kept, dropped
# ── extraction ───────────────────────────────────────────────────────────────
def _post_json(url: str, payload: dict, timeout: float) -> dict:
req = urllib.request.Request(
url,
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=timeout) as resp:
return json.loads(resp.read().decode("utf-8", "replace"))
def extract(items: list[dict], cfg: dict) -> dict:
"""Fetch page text for the top N under a global character budget."""
ex = cfg.get("extraction", {}) or {}
top_n = int(ex.get("top_n", 5))
total_budget = int(ex.get("total_chars", 12000))
per_item = int(ex.get("per_item_chars", 4000))
timeout = float(ex.get("timeout_seconds", 45))
used = 0
failures = 0
t0 = time.time()
for item in items[:top_n]:
remaining = total_budget - used
if remaining <= 200:
item["excerpt"] = ""
item["extraction"] = "skipped_budget_exhausted"
continue
cap = min(per_item, remaining)
try:
data = _post_json(
f"{FIRECRAWL_URL}/v1/scrape",
{"url": item["url"], "formats": ["markdown"]},
timeout,
)
md = ((data.get("data") or {}).get("markdown") or "").strip()
if not md:
item["excerpt"] = ""
item["extraction"] = "empty"
failures += 1
continue
item["excerpt"] = md[:cap]
item["extraction"] = "ok" if len(md) <= cap else "truncated"
used += len(item["excerpt"])
except Exception as exc: # noqa: BLE001
item["excerpt"] = ""
item["extraction"] = f"failed:{type(exc).__name__}"
failures += 1
return {
"extracted": min(top_n, len(items)),
"chars_used": used,
"budget": total_budget,
"failures": failures,
"seconds": round(time.time() - t0, 2),
}
# ── entry point ──────────────────────────────────────────────────────────────
def consume(query: str, do_extract: bool = True, explain: bool = False) -> dict:
cfg = _load_config()
url = f"{SEARXNG_URL}/search?" + urllib.parse.urlencode(
{"q": query, "format": "json"}
)
with urllib.request.urlopen(url, timeout=HTTP_TIMEOUT) as resp:
raw = json.loads(resp.read().decode("utf-8", "replace"))
results = raw.get("results", [])
kept, dropped = rank(results, cfg)
extraction = extract(kept, cfg) if do_extract else None
out = {
"query": query,
"raw_result_count": len(results),
"returned_count": len(kept),
"dropped_count": len(dropped),
"engines": sorted({r.get("engine", "?") for r in results}),
"results": [
{
"rank": i + 1,
"title": it["title"],
"url": it["url"],
"host": it["host"],
"source_type": it["source_type"],
"engines": sorted(set(it["engines"])),
"score": it["score"],
"excerpt": it.get("excerpt", ""),
"extraction": it.get("extraction", "not_attempted"),
}
for i, it in enumerate(kept)
],
"extraction": extraction,
}
if explain:
out["dropped"] = dropped
return out
def main() -> int:
args = [a for a in sys.argv[1:] if not a.startswith("--")]
do_extract = "--no-extract" not in sys.argv
explain = "--explain" in sys.argv
if not args:
print(__doc__)
return 2
query = " ".join(args)
try:
out = consume(query, do_extract=do_extract, explain=explain)
except Exception as exc: # noqa: BLE001
print(f"LAYER FAILED: {type(exc).__name__}: {exc}", file=sys.stderr)
return 2
print(json.dumps(out, indent=2))
return 0 if out["returned_count"] else 1
if __name__ == "__main__":
sys.exit(main())
-313
View File
@@ -1,313 +0,0 @@
#!/usr/bin/env python3
"""Search-stack visibility check.
The fleet shares one SearXNG instance (search) plus one extraction service
(Firecrawl). A broken search stack used to fail silently: one engine answered
and nobody could tell that the other engines had stopped contributing, or that
an enabled engine was returning nothing at all without reporting an error.
This check makes those failures visible and non-zero:
* runs two fixed queries against SearXNG; FAILS when fewer than two engines
contribute to a query, printing the contributing engines and every
``unresponsive_engines`` entry;
* FAILS when a known page cannot be extracted to non-empty markdown through
Firecrawl;
* reports every *silent zero* engine explicitly -- an engine that is enabled,
is eligible for the query category, is not listed in
``unresponsive_engines``, and still contributed no results.
Exit code 0 = healthy, 1 = degraded, 2 = the check could not run at all.
Environment overrides (all optional):
SEARXNG_URL default http://192.168.68.7:8888
FIRECRAWL_URL default http://192.168.68.7:3002
SEARCH_CHECK_QUERIES comma-separated fixed queries
SEARCH_CHECK_MIN_ENGINES default 2
SEARCH_CHECK_TIMEOUT per-request timeout in seconds, default 25
SEARCH_CHECK_EXTRACT_URL page used for the extraction leg
SEARCH_CHECK_ENGINES comma-separated engine names the stack is expected to
run; a silent zero is reported for any of them that is
enabled but contributes nothing with no error
"""
from __future__ import annotations
import json
import os
import sys
import urllib.error
import urllib.parse
import urllib.request
SEARXNG_URL = os.environ.get("SEARXNG_URL", "http://192.168.68.7:8888").rstrip("/")
FIRECRAWL_URL = os.environ.get("FIRECRAWL_URL", "http://192.168.68.7:3002").rstrip("/")
QUERIES = [
q.strip()
for q in os.environ.get(
"SEARCH_CHECK_QUERIES", "proxmox backup server,python asyncio tutorial"
).split(",")
if q.strip()
]
MIN_ENGINES = int(os.environ.get("SEARCH_CHECK_MIN_ENGINES", "2"))
TIMEOUT = float(os.environ.get("SEARCH_CHECK_TIMEOUT", "25"))
EXTRACT_URL = os.environ.get(
"SEARCH_CHECK_EXTRACT_URL", "https://en.wikipedia.org/wiki/Proxmox_Virtual_Environment"
)
# The general web-search engines this stack intentionally runs. A general query
# is expected to draw on these; an enabled one that returns nothing without an
# error is the silent-zero failure this check exists to expose. Specialised
# engines (images, videos, translate, currency, arxiv, npm, ...) are excluded on
# purpose -- contributing nothing to a general query is correct for them.
DEFAULT_EXPECTED_ENGINES = [
"bing",
"brave",
"google cse",
"yandex",
"duckduckgo",
]
EXPECTED_ENGINES = [
e.strip()
for e in os.environ.get(
"SEARCH_CHECK_ENGINES", ",".join(DEFAULT_EXPECTED_ENGINES)
).split(",")
if e.strip()
]
def _get_json(url: str) -> dict:
req = urllib.request.Request(url, headers={"User-Agent": "search-stack-check/1.0"})
with urllib.request.urlopen(req, timeout=TIMEOUT) as resp:
return json.loads(resp.read().decode("utf-8", "replace"))
def _post_json(url: str, payload: dict) -> dict:
data = json.dumps(payload).encode("utf-8")
req = urllib.request.Request(
url,
data=data,
headers={
"Content-Type": "application/json",
"User-Agent": "search-stack-check/1.0",
},
)
with urllib.request.urlopen(req, timeout=TIMEOUT) as resp:
return json.loads(resp.read().decode("utf-8", "replace"))
def enabled_expected_engines() -> set[str]:
"""Expected engines that SearXNG reports as actually enabled."""
cfg = _get_json(f"{SEARXNG_URL}/config")
enabled = {e["name"] for e in cfg.get("engines", []) if e.get("enabled")}
return {name for name in EXPECTED_ENGINES if name in enabled}
def unresponsive_names(pairs: list) -> dict[str, str]:
"""``unresponsive_engines`` is a list of [name, reason] pairs (or strings)."""
out: dict[str, str] = {}
for item in pairs or []:
if isinstance(item, (list, tuple)) and len(item) >= 2:
out[str(item[0])] = str(item[1])
elif isinstance(item, str):
out[item] = "unresponsive"
return out
# ── QUALITY GUARD (search-agent-consumption) ─────────────────────────────────
# The agent-consumption layer applies a deterministic demote/drop policy. Without
# an assertion here it could silently rot back to raw engine ordering - the same
# way the endpoint colours silently rotted before 2026-09-26.
QUALITY_QUERIES = [
"best practices agent context management",
"proxmox thin pool metadata exhaustion recovery",
]
# A demoted (content-farm) host must never occupy the top 3 for these queries.
QUALITY_TOP_N = 3
# Non-answers that must never be returned for these queries at all.
QUALITY_BANNED_HOSTS = ["bestbuy.com", "merriam-webster.com"]
def _consumption_layer_path():
here = os.path.dirname(os.path.abspath(__file__))
return os.path.join(here, "search-agent-consume.py")
def check_ranking_quality() -> list[str]:
"""Return a list of quality failures; empty means healthy."""
import subprocess as _sp
layer = _consumption_layer_path()
if not os.path.exists(layer):
return [f"agent-consumption layer missing: {layer}"]
failures: list[str] = []
for query in QUALITY_QUERIES:
r = _sp.run([sys.executable, layer, "--no-extract", "--explain", query],
capture_output=True, text=True, timeout=120)
if r.returncode != 0:
failures.append(f"{query!r}: layer exited {r.returncode} ({r.stderr[:120]})")
continue
try:
data = json.loads(r.stdout)
except json.JSONDecodeError:
failures.append(f"{query!r}: layer returned unparseable JSON")
continue
results = data.get("results", [])
if len(results) < QUALITY_TOP_N:
failures.append(f"{query!r}: only {len(results)} results returned")
continue
# load the demote list from the SAME config the layer uses
cfg_path = os.path.join(os.path.dirname(layer), "..", "config", "search-ranking.yaml")
demoted: set[str] = set()
try:
sys.path.insert(0, os.path.dirname(layer))
import importlib.util as _iu
spec = _iu.spec_from_file_location("_sac_cfg", layer)
mod = _iu.module_from_spec(spec)
spec.loader.exec_module(mod)
demoted = set(mod._load_config().get("demote_domains", []) or [])
except Exception: # noqa: BLE001
failures.append(f"{query!r}: could not load demote_domains from config")
for item in results[:QUALITY_TOP_N]:
host = (item.get("host") or "")
for d in demoted:
if host == d or host.endswith("." + d):
failures.append(
f"{query!r}: demoted host {host} in top {QUALITY_TOP_N}"
)
for item in results:
host = (item.get("host") or "")
for b in QUALITY_BANNED_HOSTS:
if host == b or host.endswith("." + b):
failures.append(f"{query!r}: non-answer host {host} returned")
return failures
def main() -> int:
failures: list[str] = []
print(f"Search stack check -- {SEARXNG_URL}")
print(f"Queries: {QUERIES!r} min contributing engines: {MIN_ENGINES}")
print("=" * 72)
try:
eligible = enabled_expected_engines()
except Exception as exc: # noqa: BLE001 - report, do not traceback
print(f"FAIL: could not read /config from SearXNG: {exc!r}")
return 2
print(f"Expected engines, enabled ({len(eligible)}): {sorted(eligible)}")
missing = sorted(set(EXPECTED_ENGINES) - eligible)
if missing:
print(f"Expected engines NOT enabled: {missing}")
failures.append(f"expected engines not enabled in SearXNG: {missing}")
contributed: dict[str, int] = {name: 0 for name in eligible}
silent_zero_all: dict[str, list[str]] = {}
for query in QUERIES:
url = f"{SEARXNG_URL}/search?" + urllib.parse.urlencode(
{"q": query, "format": "json"}
)
print("-" * 72)
print(f"QUERY: {query!r}")
try:
data = _get_json(url)
except Exception as exc: # noqa: BLE001
print(f" FAIL: query request failed: {exc!r}")
failures.append(f"query {query!r} request failed: {exc!r}")
continue
results = data.get("results", [])
engines: dict[str, int] = {}
for r in results:
name = r.get("engine", "?")
engines[name] = engines.get(name, 0) + 1
unresponsive = unresponsive_names(data.get("unresponsive_engines", []))
print(f" results: {len(results)}")
print(f" contributing engines: {engines or '(none)'}")
print(f" unresponsive_engines: {unresponsive or '(none)'}")
for name in engines:
contributed[name] = contributed.get(name, 0) + engines[name]
if len(engines) < MIN_ENGINES:
msg = (
f"query {query!r} had only {len(engines)} contributing engine(s) "
f"({sorted(engines)}); need >= {MIN_ENGINES}"
)
print(f" FAIL: {msg}")
failures.append(msg)
silent = sorted(
n for n in eligible if n not in engines and n not in unresponsive
)
if silent:
silent_zero_all[query] = silent
print(
" SILENT ZERO (enabled, no error, no results -- reported, "
f"not fatal): {silent}"
)
print("=" * 72)
print("Engine contribution across all queries:")
for name in sorted(contributed):
status = "ZERO" if contributed[name] == 0 else "ok"
print(f" {name:<24} {contributed[name]:>4} {status}")
if silent_zero_all:
print("-" * 72)
print("SILENT-ZERO ENGINES REPORTED (no error raised, no results returned):")
for query, names in silent_zero_all.items():
print(f" {query!r}: {names}")
print(" NOTE: a silent zero is REPORTED, not counted as a failure. These")
print(" engines are expected to answer a general query, but contributing")
print(" nothing to one query can be legitimate (result de-duplication, or")
print(" an engine that only fires on certain query shapes). Only the")
print(f" <{MIN_ENGINES}-contributing-engine floor and the extraction leg fail the run.")
print("-" * 72)
print(f"EXTRACTION: scraping {EXTRACT_URL} via {FIRECRAWL_URL}/v1/scrape")
try:
payload = _post_json(
f"{FIRECRAWL_URL}/v1/scrape",
{"url": EXTRACT_URL, "formats": ["markdown"]},
)
markdown = ((payload.get("data") or {}).get("markdown") or "").strip()
if not markdown:
msg = "extraction returned empty markdown"
print(f" FAIL: {msg}")
failures.append(msg)
else:
print(f" ok: {len(markdown)} chars of markdown returned")
print(f" first line: {markdown.splitlines()[0][:120]!r}")
except Exception as exc: # noqa: BLE001
msg = f"extraction request failed: {exc!r}"
print(f" FAIL: {msg}")
failures.append(msg)
print("-" * 72)
print("RANKING QUALITY (agent-consumption layer)")
quality = check_ranking_quality()
if quality:
for q in quality:
print(f" FAIL: {q}")
failures.extend(quality)
else:
print(" ok: no demoted host in the top 3; no banned non-answer returned")
print("=" * 72)
if failures:
print("VERDICT: FAIL")
for f in failures:
print(f" - {f}")
return 1
print("VERDICT: PASS -- multiple engines contributing, extraction healthy")
return 0
if __name__ == "__main__":
sys.exit(main())
-55
View File
@@ -1,55 +0,0 @@
# secret-allowlist.tsv — exceptions for scripts/secret-scan.sh, every entry with a reason.
#
# Format: <rule-id|*><TAB><path-glob><TAB><literal-substring><TAB><reason>
# Blank lines and lines whose first field starts with '#' are ignored.
# A finding is suppressed only when ALL THREE of rule, path and literal match:
# * the rule id equals the finding's rule id, or is '*'
# * the finding's repo-relative path matches <path-glob> (bash glob)
# * the finding's line contains <literal-substring> verbatim
# An entry whose reason is empty is a hard error (exit 2) — no silent exceptions.
#
# RULE: never allowlist a live credential, and never broaden an entry (rule '*',
# a wide path glob, or a short generic literal) just to silence a finding.
# If the finding is real, remove the credential from the file.
#
# Entries are one per deliberate synthetic example, so the file reads as an
# audit trail of reviewed exceptions rather than a list of things to ignore.
# Rule '*' is used only where the same literal is matched by more than one rule.
#
# ── The 2026-09-17 purge placeholders ─────────────────────────────────────
# PR #112 replaced six live credentials with `«vault: <project>/<env> <SECRET>»`
# references. Those references are safe by construction (they name where the
# secret is read from), but they are listed here explicitly rather than being
# filtered by a general "vault" rule, so a new occurrence still needs a
# deliberate, reasoned entry.
secret-assign litellm-api-keys.prose.md MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY» 2026-09-17 purge: replaced the live Mumuni LiteLLM key with its vault reference; no literal credential.
secret-assign litellm-api-keys.prose.md MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY» 2026-09-17 purge: replaced the live Mumuni Zulip key with its vault reference; no literal credential.
* infrastructure-control.prose.md PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
cred-prose infrastructure-control.prose.md Admin credentials: 2026-09-17 purge: the Stirling admin user/password are two `«vault: ...»` references; no literal credential.
* scripts/daily-infra-report.py PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
secret-assign stirling-pdf-agent-access.prose.md «vault: infrastructure/production STIRLING_API_KEY» 2026-09-17 purge: Stirling PDF API key is read from the vault; the curl example only names the vault path.
bearer-token agent-zero-fix-summary.md «vault: agents/production OPENROUTER_API_KEY» 2026-09-17 purge: OpenRouter key is read from the vault; the example curl only names the vault path.
# ── Deliberate synthetic examples in contracts (not from the purge) ───────
# These exist to teach the rule they illustrate. They are listed here so the
# guard is never taught to skip the words "synthetic"/"example" — a fabricated
# example is always an explicit exception, never a pattern-level exemption.
* hermes-key-enforcement.prose.md sk-synthetic-external-example Rule 15 illustration of a hardcoded external key that is tolerated; fabricated, never a live key.
* hermes-key-enforcement.prose.md sk-synthetic-example-12345 Rule 15 illustration of a forbidden hardcoded key; fabricated, never a live key.
openai-key hermes-key-enforcement.prose.md sk-synthetic-litellm- Fabricated key name inside a `grep 'LITELLM_API_KEY=...'` example; not a live key.
secret-assign hermes-key-enforcement.prose.md sk-NEW_KEY Placeholder standing for the rotated key in an `infisical secrets set` command; not a literal key.
openrouter-key agent-zero-openrouter-key.prose.md sk-or-v1-synthetic Synthetic key prefix in the contract's example response; the real key is read from the vault.
openai-key litellm-api-keys.prose.md sk-synthetic-tanko-example Fabricated key name in migration history prose; not a live key.
openai-key litellm-self-heal.prose.md sk-syslog-local-master-key Deprecated local LiteLLM master key name documented as no-live-usage; kept for history, not a usable credential.
# ── Redacted evidence, not a credential ──────────────────────────────────
secret-assign docs/probe-drift-round2-evidence.md =sk-... Probe evidence records redacted key trailers (`sk-...x6uw`); the usable part of the key is not present.
# ── tests/test_secret_scan.sh fixtures ───────────────────────────────────
# The self-test plants these fabricated values into a TEMP tree, whose path no
# entry here covers, so each still fails the guard when planted (see the test's
# "... fails the guard" cases). They are listed only so the repo-wide scan of
# the test file itself stays quiet.
* tests/test_secret_scan.sh sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef Self-test fixture: fabricated OpenRouter-shaped key written to a temp tree; the guard must fail on it there.
bearer-token tests/test_secret_scan.sh Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb Self-test fixture: fabricated Bearer token written to a temp tree; the guard must fail on it there.
proxmox-token tests/test_secret_scan.sh PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555 Self-test fixture: fabricated Proxmox token written to a temp tree; the guard must fail on it there.
private-key tests/test_secret_scan.sh -----BEGIN OPENSSH PRIVATE KEY----- Self-test fixture: fabricated PEM banner written to a temp tree; the guard must fail on it there.
cred-prose tests/test_secret_scan.sh Admin credentials: Self-test fixture: fabricated prose credential line written to a temp tree; the guard must fail on it there.
secret-assign tests/test_secret_scan.sh DB_PASSWORD=correct-horse-battery-staple Self-test fixture: fabricated password assignment written to a temp tree; the guard must fail on it there.
Can't render this file because it contains an unexpected character in line 23 and column 25.
-22
View File
@@ -1,22 +0,0 @@
# secret-patterns.tsv — checked-in pattern list for scripts/secret-scan.sh
#
# Format: <rule-id><TAB><POSIX ERE><TAB><description><TAB><check>
# Blank lines and lines whose first field starts with '#' are ignored.
# <check> is optional; the only value today is "value", which tells the scanner
# to run the matched value through its inert-value classifier (see
# value_is_inert in secret-scan.sh) so bare identifiers, env refs and dotted
# code access are not reported as credentials. Omit the column to report every
# regex hit.
# Matching is case-insensitive, so `API_KEY` and `api_key` both count.
#
# Add a rule here, never inline in secret-scan.sh: this file is the single
# auditable list of what the guard considers credential-shaped.
openai-key \bsk-[A-Za-z0-9_-]{16,} OpenAI/LiteLLM-style "sk-" secret key (also hyphenated sk-proj- keys)
openrouter-key \bsk-or-v1-[A-Za-z0-9_-]{8,} OpenRouter API key
stripe-live-key \bsk_live_[A-Za-z0-9]{8,} Stripe live secret key
proxmox-token PVEAPIToken=[^[:space:]"']+ Proxmox API token literal
bearer-token bearer[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) literal Bearer token (http header or prose)
auth-header authorization:[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) Authorization header carrying a raw literal value
private-key -----BEGIN [A-Z ]*PRIVATE KEY----- PEM private key block
cred-prose credentials?[[:space:]]*[:=][[:space:]]*[^[:space:]] prose credential line carrying a value
secret-assign (api[_-]?key|apikey|passwd|password|secret|token)s?["']?[[:space:]]*[:=][[:space:]]*["']?(«.{3,}»|[A-Za-z0-9_./+=-]{8,}) credential assignment carrying a literal value value
Can't render this file because it contains an unexpected character in line 5 and column 48.
-269
View File
@@ -1,269 +0,0 @@
#!/usr/bin/env bash
# secret-scan.sh — commit-time secret guard. FAILS (exit 1) on a credential-shaped
# string, so a build cannot go green with a credential committed to it.
#
# Usage:
# scripts/secret-scan.sh # scan the whole git-tracked tree (default)
# scripts/secret-scan.sh --tree
# scripts/secret-scan.sh --path DIR # scan an arbitrary directory (git not required)
# scripts/secret-scan.sh --staged # scan added lines in the index (pre-commit)
# scripts/secret-scan.sh --diff REF # scan added lines since REF (e.g. origin/master)
# --quiet only print the verdict and findings, no per-mode banner
#
# Exit codes: 0 clean, 1 credential found, 2 usage/config error.
#
# Patterns live in scripts/secret-patterns.tsv
# Exceptions live in scripts/secret-allowlist.tsv (every entry carries a reason;
# a missing reason is a hard error, so the guard fails closed).
#
# Dependencies are deliberately bash + coreutils + grep + sed/awk + git. The
# Gitea Actions runner executes job steps INSIDE the runner container, which
# has no node and no python by default: keep this script free of both.
set -uo pipefail
SELF_DIR=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
ROOT=$(cd -- "$SELF_DIR/.." && pwd)
PATTERNS_FILE="$SELF_DIR/secret-patterns.tsv"
ALLOWLIST_FILE="$SELF_DIR/secret-allowlist.tsv"
# The guard's own definition files are not scannable content: the pattern list
# necessarily contains the pattern text, and the allowlist necessarily contains
# the allowed literals. Narrow, exact-path exclusion — not a wildcard.
SELF_FILES=(
"scripts/secret-scan.sh"
"scripts/secret-patterns.tsv"
"scripts/secret-allowlist.tsv"
)
MODE="tree"
PATH_DIR=""
DIFF_REF=""
QUIET=0
usage() {
sed -n '2,20p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'
exit 2
}
while [ $# -gt 0 ]; do
case "$1" in
--tree) MODE="tree" ;;
--path) MODE="path"; PATH_DIR="${2:-}"; shift ;;
--staged) MODE="staged" ;;
--diff) MODE="diff"; DIFF_REF="${2:-}"; shift ;;
--quiet) QUIET=1 ;;
-h|--help) usage ;;
*) echo "secret-scan: unknown argument '$1'" >&2; usage ;;
esac
shift
done
[ -f "$PATTERNS_FILE" ] || { echo "secret-scan: missing $PATTERNS_FILE" >&2; exit 2; }
[ -f "$ALLOWLIST_FILE" ] || { echo "secret-scan: missing $ALLOWLIST_FILE" >&2; exit 2; }
if [ "$MODE" = "path" ] && [ -z "$PATH_DIR" ]; then
echo "secret-scan: --path needs a directory" >&2; exit 2
fi
if [ "$MODE" = "diff" ] && [ -z "$DIFF_REF" ]; then
echo "secret-scan: --diff needs a base ref" >&2; exit 2
fi
# ── Load patterns ──────────────────────────────────────────────────────────
RULE_IDS=()
RULE_RES=()
RULE_DESCS=()
RULE_CHECKS=()
COMBINED=""
while IFS=$'\t' read -r id re desc check; do
case "$id" in ''|'#'*) continue ;; esac
[ -n "$re" ] || continue
RULE_IDS+=("$id"); RULE_RES+=("$re"); RULE_DESCS+=("$desc"); RULE_CHECKS+=("${check:-}")
if [ -z "$COMBINED" ]; then COMBINED="($re)"; else COMBINED="$COMBINED|($re)"; fi
done < "$PATTERNS_FILE"
if [ "${#RULE_IDS[@]}" -eq 0 ]; then
echo "secret-scan: no patterns loaded from $PATTERNS_FILE" >&2; exit 2
fi
# ── Load allowlist (fails closed on a missing reason) ──────────────────────
AL_RULES=()
AL_GLOBS=()
AL_LITS=()
AL_REASONS=()
AL_LINENO=0
while IFS=$'\t' read -r rule glob lit reason; do
AL_LINENO=$((AL_LINENO + 1))
case "$rule" in ''|'#'*) continue ;; esac
if [ -z "$glob" ] || [ -z "$lit" ] || [ -z "$reason" ]; then
echo "secret-scan: ❌ $ALLOWLIST_FILE:$AL_LINENO — allowlist entry needs <rule> <path-glob> <literal> <reason>; reason-based exceptions only, refusing to run" >&2
exit 2
fi
AL_RULES+=("$rule"); AL_GLOBS+=("$glob"); AL_LITS+=("$lit"); AL_REASONS+=("$reason")
done < "$ALLOWLIST_FILE"
# nocasematch is toggled only around the regex test; path globs must stay
# case-sensitive, so it is never left on.
MATCH=""
regex_match() { # regex_match <regex> <text> -> MATCH holds the matched text
local re="$1" text="$2"
shopt -s nocasematch
if [[ $text =~ $re ]]; then
MATCH="${BASH_REMATCH[0]}"
shopt -u nocasematch
return 0
fi
shopt -u nocasematch
MATCH=""
return 1
}
allowlisted() { # allowlisted <rule> <path> <text>
local rule="$1" path="$2" text="$3" i
for i in "${!AL_RULES[@]}"; do
[ "${AL_RULES[$i]}" = "$rule" ] || [ "${AL_RULES[$i]}" = "*" ] || continue
# The unquoted RHS is deliberate: <path-glob> is a bash glob, not a literal.
# shellcheck disable=SC2053
[[ $path == ${AL_GLOBS[$i]} ]] || continue
[[ $text == *"${AL_LITS[$i]}"* ]] || continue
return 0
done
return 1
}
mask_value() { # mask_value <text> <match> — never echo a credential to logs.
# Print only the part of the line BEFORE the match, then <redacted>: the match
# itself and everything after it (which may include a value the rule's regex
# stopped short of, e.g. `credentials:` followed by a backticked password) is
# never written to stdout.
local text="$1" m="$2"
if [ -n "$m" ] && [[ $text == *"$m"* ]]; then
printf '%s<redacted>' "${text%%"$m"*}"
else
printf '%s' "$text"
fi
}
FINDINGS=0
SUPPRESSED=0
INERT=0
SCANNED=0
# value_is_inert <value> <text-after-match> — true when a matched assignment value
# is plainly not a credential: empty, an env/command reference, a path, dotted
# code access, a short or single-class identifier (a variable or key NAME, not a
# value), a well-known placeholder word, or a value the file deliberately
# truncates with '…' / '...' (a redacted prefix is not a usable credential).
# Deliberately does NOT know the words "synthetic" or "example": a fabricated
# example must be an explicit allowlist entry.
value_is_inert() {
local v="$1" rest="$2"
case "$rest" in '…'*|'...'*) return 0 ;; esac
v="${v%\"}"; v="${v#\"}"; v="${v%\'}"; v="${v#\'}"
case "$v" in
''|\$*|\{*|'<'*|'%'*|'('*|'/'*|'\\'*) return 0 ;;
not-needed|no-key-required|none|null|true|false|redacted|placeholder|example|dummy|changeme|change-me|your-key|your_key|key|token|secret|password) return 0 ;;
esac
# dotted code access: os.environ.get / process.env.ZULIP_API_KEY / cfg.a
if [[ $v =~ ^[a-z_][a-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)+$ ]]; then return 0; fi
# bare identifier (no punctuation beyond _): a NAME, not a value. A real
# secret in this shape is long and mixes letters with digits.
if [[ $v =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
[ "${#v}" -lt 20 ] && return 0
[[ $v =~ [0-9] ]] || return 0
return 1
fi
return 1
}
report_finding() { # report_finding <path> <line> <text>
local path="$1" line="$2" text="$3" i val
for i in "${!RULE_IDS[@]}"; do
regex_match "${RULE_RES[$i]}" "$text" || continue
SCANNED=$((SCANNED + 1))
if [ "${RULE_CHECKS[$i]}" = "value" ]; then
val="${MATCH#*[:=]}"
val="${val# }"
if value_is_inert "$val" "${text#*"$MATCH"}"; then
INERT=$((INERT + 1))
continue
fi
fi
if allowlisted "${RULE_IDS[$i]}" "$path" "$text"; then
SUPPRESSED=$((SUPPRESSED + 1))
continue
fi
FINDINGS=$((FINDINGS + 1))
printf ' ❌ %s:%s [%s] %s\n' "$path" "$line" "${RULE_IDS[$i]}" "${RULE_DESCS[$i]}"
printf ' | %s\n' "$(mask_value "$text" "$MATCH")"
done
}
self_excluded() { # self_excluded <repo-relative-path>
local p="$1" s
for s in "${SELF_FILES[@]}"; do
[ "$p" = "$s" ] && return 0
done
return 1
}
# ── Collect candidate lines and scan them ─────────────────────────────────
if [ "$MODE" = "tree" ] || [ "$MODE" = "path" ]; then
if [ "$MODE" = "tree" ]; then
BASE="$ROOT"
git -C "$BASE" rev-parse --git-dir >/dev/null 2>&1 || { echo "secret-scan: --tree needs a git checkout (use --path DIR)" >&2; exit 2; }
mapfile -d '' candidate < <(git -C "$BASE" ls-files -z 2>/dev/null)
if [ "${#candidate[@]}" -eq 0 ]; then
echo "secret-scan: ❌ no tracked files — refusing to report clean" >&2; exit 2
fi
else
BASE=$(cd -- "$PATH_DIR" 2>/dev/null && pwd) || { echo "secret-scan: --path '$PATH_DIR' is not a directory" >&2; exit 2; }
mapfile -t candidate < <(cd -- "$BASE" && find . -type f -not -path './.git/*' | sed 's|^\./||')
if [ "${#candidate[@]}" -eq 0 ]; then
echo "secret-scan: ❌ no files under $BASE — refusing to report clean" >&2; exit 2
fi
fi
[ "$QUIET" -eq 1 ] || echo "── secret scan ($MODE): ${#candidate[@]} files under $BASE ──"
for rel in "${candidate[@]}"; do
[ -f "$BASE/$rel" ] || continue
self_excluded "$rel" && continue
while IFS= read -r hit; do
[ -n "$hit" ] || continue
report_finding "$rel" "${hit%%:*}" "${hit#*:}"
done < <(grep -nEIi -e "$COMBINED" "$BASE/$rel" 2>/dev/null || true)
done
else
# --staged / --diff: only ADDED lines, with the post-change line number.
if [ "$MODE" = "staged" ]; then
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines in the index ──"
DIFF_TEXT=$(git -C "$ROOT" diff --cached --unified=0 --no-color -- . 2>/dev/null)
else
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines since $DIFF_REF ──"
DIFF_TEXT=$(git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"...HEAD 2>/dev/null \
|| git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"..HEAD 2>/dev/null)
fi
if [ -z "$DIFF_TEXT" ]; then
[ "$QUIET" -eq 1 ] || echo " (no added lines)"
fi
while IFS=$'\t' read -r rel line text; do
[ -n "$rel" ] || continue
self_excluded "$rel" && continue
report_finding "$rel" "$line" "$text"
done < <(printf '%s\n' "$DIFF_TEXT" | awk '
/^\+\+\+ / { f=$2; sub(/^b\//,"",f); next }
/^@@ / { if (match($0, /\+[0-9]+/)) ln=substr($0, RSTART+1, RLENGTH-1)+0; next }
(/^\+/ && !/^\+\+\+/) { print f "\t" ln "\t" substr($0,2); ln++; next }
')
fi
# ── Verdict ────────────────────────────────────────────────────────────────
if [ "$FINDINGS" -gt 0 ]; then
echo ""
echo "❌ SECRET SCAN FAILED — $FINDINGS credential-shaped string(s) in ${MODE} content."
echo " Fix: remove the credential and read it from the vault/env."
echo " Only a deliberate synthetic example may be added to scripts/secret-allowlist.tsv,"
echo " one entry per file/rule/literal, with a reason. Never allowlist a live credential."
exit 1
fi
echo "✅ secret scan clean (${MODE}; ${SUPPRESSED} allowlisted exception(s), ${INERT} inert value(s) ignored)"
exit 0
+1 -1
View File
@@ -1,6 +1,6 @@
#!/bin/bash
# swap-gpu-dense-model.sh — Swap RTX 3090 from qwen3.6-27B-code to SmartCode-Fable-5
# Run when download completes: ssh llmuser@192.168.68.8 'sudo bash -s' < this script
# Run when download completes: ssh root@192.168.68.8 'bash -s' < this script
#
# Usage: bash swap-gpu-dense-model.sh
# Requires: new model at /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf
-228
View File
@@ -1,228 +0,0 @@
#!/bin/bash
# test_infra_monitoring.sh — Asserts that probe calls use the documented targets.
#
# Strategy: stub curl and ssh on PATH to capture the exact arguments each leg
# builds, then assert the URL/port of every call. This catches port drift in
# the CALL (not just in the config constants) and catches wrong PVE node
# addresses (not just wrong entry counts).
#
# Run: bash scripts/test_infra_monitoring.sh
# Exits 0 if all assertions pass, 1 otherwise.
set -uo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
SCRIPT="${SCRIPT_DIR}/infra-monitoring.sh"
PASS=0
FAIL=0
assert() {
local desc="$1" condition="$2"
if eval "$condition"; then
echo " ✅ $desc"
PASS=$((PASS+1))
else
echo " 🔴 $desc"
FAIL=$((FAIL+1))
fi
}
echo "=== test_infra_monitoring.sh ==="
echo ""
# ── Stub curl: capture argv to a file, return 200 ──────────────────────────
STUB_DIR=$(mktemp -d)
trap 'rm -rf "$STUB_DIR"' EXIT
# Stub curl: first arg after flags is the URL; capture all args
cat > "$STUB_DIR/curl" << 'STUBEOF'
#!/bin/bash
echo "$@" >> "${CURL_STUB_LOG:-/dev/null}"
# Print 200 for %{http_code}
printf '%s\n' "200"
exit 0
STUBEOF
chmod +x "$STUB_DIR/curl"
# Stub ssh: first arg after options is the remote command; capture it
cat > "$STUB_DIR/ssh" << 'SSTUBEOF'
#!/bin/bash
echo "SSH $@" >> "${SSH_STUB_LOG:-/dev/null}"
# The last arg is the remote command — extract and log curl args
for arg in "$@"; do
if [[ "$arg" == curl* ]]; then
echo "$arg" >> "${SSH_STUB_LOG:-/dev/null}"
fi
done
printf '%s\n' "200"
exit 0
SSTUBEOF
chmod +x "$STUB_DIR/ssh"
# ── Run the monitor with stubs ────────────────────────────────────────────
CURL_LOG="$STUB_DIR/curl_calls.log"
SSH_LOG="$STUB_DIR/ssh_calls.log"
touch "$CURL_LOG" "$SSH_LOG"
CURL_STUB_LOG="$CURL_LOG" SSH_STUB_LOG="$SSH_LOG" \
PATH="$STUB_DIR:$PATH" bash "$SCRIPT" > "$STUB_DIR/output.txt" 2>&1
# ── 1. Port drift detection (from actual curl invocations) ─────────────────
assert "Grafana probed at port 3001" \
'grep -q "http://192.168.68.116:3001/api/health" "$CURL_LOG"'
assert "Prometheus probed at port 9090" \
'grep -q "http://192.168.68.116:9090/-/healthy" "$CURL_LOG"'
assert "LiteLLM probed via nginx at port 80" \
'grep -q "http://192.168.68.116:80/litellm/health" "$CURL_LOG"'
assert "PVE API probed at port 8006" \
'grep -q ":8006/api2/json/version" "$CURL_LOG"'
assert "GPU exporter probed at port 9400" \
'grep -q ":9400/metrics" "$CURL_LOG"'
# ── 2. PVE API: exact node addresses (catches wrong IPs) ──────────────────
# Each real PVE node must be probed; CT 116 must NOT be in the PVE set
assert "PVE acerpve 192.168.68.9 probed" \
'grep -q "https://192.168.68.9:8006/api2/json/version" "$CURL_LOG"'
assert "PVE minipve 192.168.68.12 probed" \
'grep -q "https://192.168.68.12:8006/api2/json/version" "$CURL_LOG"'
assert "PVE storepve 192.168.68.6 probed" \
'grep -q "https://192.168.68.6:8006/api2/json/version" "$CURL_LOG"'
assert "PVE amdpve 192.168.68.15 probed" \
'grep -q "https://192.168.68.15:8006/api2/json/version" "$CURL_LOG"'
assert "PVE ocupve 192.168.68.5 probed" \
'grep -q "https://192.168.68.5:8006/api2/json/version" "$CURL_LOG"'
# CT 116 (.116) must NOT appear as a PVE API target
assert "CT 116 (.116) NOT probed as PVE API node" \
'! grep -q "https://192.168.68.116:8006" "$CURL_LOG"'
# ── 3. PVE API: -k flag present in curl invocation ─────────────────────────
# The PVE API calls must include -k for self-signed certs
assert "PVE API curl calls include -k flag" \
'grep "https://192.168.68.9:8006" "$CURL_LOG" | grep -q -- "-k"'
# ── 4. Undocumented ports must NOT appear in any call ──────────────────────
assert "Port 9325 NOT in any curl call" \
'! grep -q ":9325" "$CURL_LOG"'
assert "Port 9405 NOT in any curl call" \
'! grep -q ":9405" "$CURL_LOG"'
# ── 5. Docker Stats / PVE Exporter: SSH-probed at correct ports ────────────
assert "Docker Stats probed at port 9324 via SSH" \
'grep -q "9324" "$SSH_LOG"'
assert "PVE Exporter probed at port 9221 via SSH" \
'grep -q "9221" "$SSH_LOG"'
# ── 5a. Per-leg assertions (proves which leg owns which port) ──────────────
# The script source must show Docker Stats using $DOCKER_STATS_PORT and
# PVE Exporter using $PVE_EXPORTER_PORT in the correct leg sections
assert "Docker Stats leg uses DOCKER_STATS_PORT constant" \
'grep -A 3 "# 6. Docker Stats" "$SCRIPT" | grep -q "\$DOCKER_STATS_PORT"'
assert "PVE Exporter leg uses PVE_EXPORTER_PORT constant" \
'grep -A 3 "# 7. PVE Exporter" "$SCRIPT" | grep -q "\$PVE_EXPORTER_PORT"'
# Verify the constants themselves are set to the correct values
assert "DOCKER_STATS_PORT constant set to 9324" \
'grep -q "^DOCKER_STATS_PORT=\"9324\"" "$SCRIPT"'
assert "PVE_EXPORTER_PORT constant set to 9221" \
'grep -q "^PVE_EXPORTER_PORT=\"9221\"" "$SCRIPT"'
# ── 5b. Stale port 9323 (dockerd) NOT probed ──────────────────────────────
assert "Port 9323 (dockerd) NOT in SSH log" \
'! grep -q "9323" "$SSH_LOG"'
# ── 6. No stale ports in the script source (belt-and-suspenders) ──────────
assert "Port 9325 (historical) NOT in script source" \
'! grep -q "9325" "$SCRIPT"'
assert "Port 9405 (historical) NOT in script source" \
'! grep -q "9405" "$SCRIPT"'
# ── 7. Failure-line content includes non-empty kind ────────────────────────
TMP_DIR=$(mktemp -d)
trap 'rm -rf "$TMP_DIR"' EXIT
# Test 7a: Unexpected status (500) → kind should be unexpected:500
cat > "$TMP_DIR/curl" << 'EOF'
#!/bin/bash
# Stub: return 500 for Grafana port, 200 otherwise
for arg in "$@"; do
if [[ "$arg" == *":3001"* ]]; then
echo "500"
exit 0
fi
done
echo "200"
exit 0
EOF
chmod +x "$TMP_DIR/curl"
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
assert "Grafana failure line exists (unexpected status)" \
'[[ -n "$GRAFANA_FAIL" ]]'
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
assert "Grafana failure kind is non-empty (unexpected status)" \
'[[ -n "$KIND" ]]'
# Test 7b: TLS error (000 + exit 60) → kind should be tls
cat > "$TMP_DIR/curl" << 'EOF'
#!/bin/bash
# Stub: return 000 and exit 60 for Grafana port (TLS error) on ALL invocations
for arg in "$@"; do
if [[ "$arg" == *":3001"* ]]; then
echo "000"
exit 60
fi
done
echo "200"
exit 0
EOF
chmod +x "$TMP_DIR/curl"
# Also stub ssh to return 000 + exit 60 for the retry
cat > "$TMP_DIR/ssh" << 'EOF'
#!/bin/bash
for arg in "$@"; do
if [[ "$arg" == curl* ]]; then
echo "000"
exit 60
fi
done
echo "200"
exit 0
EOF
chmod +x "$TMP_DIR/ssh"
export PATH="$TMP_DIR:$PATH"
OUT=$(bash "$SCRIPT" 2>&1)
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
assert "Grafana failure line exists (TLS error)" \
'[[ -n "$GRAFANA_FAIL" ]]'
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
assert "Grafana failure kind is tls" \
'[[ "$KIND" == "tls" ]]'
# ── Summary ─────────────────────────────────────────────────────────────────
echo ""
echo "Results: ${PASS} passed, ${FAIL} failed"
if [ $FAIL -gt 0 ]; then
echo " 🔴 TESTS FAILED"
exit 1
else
echo " ✅ ALL TESTS PASSED"
exit 0
fi
-125
View File
@@ -1,125 +0,0 @@
#!/bin/bash
# test_proxmox_monitor.sh — Stub-driven tests for PBS GC liveness leg
set -uo pipefail
SCRIPT="$(cd "$(dirname "$0")" && pwd)/proxmox-monitor.sh"
PASS=0
FAIL=0
# ── Helpers ────────────────────────────────────────────────────────────────
assert() {
local desc="$1" cond="$2"
if eval "$cond" 2>/dev/null; then
echo " ✅ $desc"
PASS=$((PASS+1))
else
echo " 🔴 $desc"
FAIL=$((FAIL+1))
fi
}
# ── 1. Healthy: fresh GC, 0 B pending ─────────────────────────────────────
TMP_DIR=$(mktemp -d)
FRESH_ENDTIME=$(( $(date -u +%s) - (1 * 3600) )) # 1 hour ago
cat > "$TMP_DIR/ssh" << EOF
#!/bin/bash
# Stub: return valid JSON with fresh endtime
echo '[{"store":"storepve-datastore","last-run-endtime":$FRESH_ENDTIME,"pending-bytes":0}]'
exit 0
EOF
chmod +x "$TMP_DIR/ssh"
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
assert "Healthy: PBS line exists" '[[ -n "$PBS_LINE" ]]'
assert "Healthy: shows healthy verdict" '[[ "$PBS_LINE" == *"healthy"* ]]'
assert "Healthy: shows pending-bytes 0 B" '[[ "$PBS_LINE" == *"pending-bytes: 0 B"* ]]'
rm -rf "$TMP_DIR"
# ── 2. Stale: GC >48h old ─────────────────────────────────────────────────
TMP_DIR=$(mktemp -d)
OLD_ENDTIME=$(( $(date -u +%s) - (49 * 3600) )) # 49 hours ago
cat > "$TMP_DIR/ssh" << EOF
#!/bin/bash
echo '[{"store":"storepve-datastore","last-run-endtime":$OLD_ENDTIME,"pending-bytes":1048576}]'
exit 0
EOF
chmod +x "$TMP_DIR/ssh"
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
assert "Stale: PBS line exists" '[[ -n "$PBS_LINE" ]]'
assert "Stale: shows stale verdict" '[[ "$PBS_LINE" == *"stale"* ]]'
assert "Stale: shows pending-bytes 1048576 B" '[[ "$PBS_LINE" == *"pending-bytes: 1048576 B"* ]]'
rm -rf "$TMP_DIR"
# ── 3. Probe-failed: empty output ─────────────────────────────────────────
TMP_DIR=$(mktemp -d)
cat > "$TMP_DIR/ssh" << 'EOF'
#!/bin/bash
exit 0
EOF
chmod +x "$TMP_DIR/ssh"
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
assert "Probe-failed (empty): PBS line exists" '[[ -n "$PBS_LINE" ]]'
assert "Probe-failed (empty): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
rm -rf "$TMP_DIR"
# ── 4. Probe-failed: unparseable output ───────────────────────────────────
TMP_DIR=$(mktemp -d)
cat > "$TMP_DIR/ssh" << 'EOF'
#!/bin/bash
echo "proxmox-backup-manager: command not found"
exit 0
EOF
chmod +x "$TMP_DIR/ssh"
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
assert "Probe-failed (unparseable): PBS line exists" '[[ -n "$PBS_LINE" ]]'
assert "Probe-failed (unparseable): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
rm -rf "$TMP_DIR"
# ── 5. Null endtime: never-run ────────────────────────────────────────────
TMP_DIR=$(mktemp -d)
cat > "$TMP_DIR/ssh" << 'EOF'
#!/bin/bash
echo '[{"store":"storepve-datastore","last-run-endtime":null,"pending-bytes":0}]'
exit 0
EOF
chmod +x "$TMP_DIR/ssh"
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
assert "Null endtime: PBS line exists" '[[ -n "$PBS_LINE" ]]'
assert "Null endtime: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
rm -rf "$TMP_DIR"
# ── 6. Datastore absent ───────────────────────────────────────────────────
TMP_DIR=$(mktemp -d)
cat > "$TMP_DIR/ssh" << 'EOF'
#!/bin/bash
echo '[{"store":"s3-archive","last-run-endtime":null,"pending-bytes":0}]'
exit 0
EOF
chmod +x "$TMP_DIR/ssh"
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
assert "Datastore absent: PBS line exists" '[[ -n "$PBS_LINE" ]]'
assert "Datastore absent: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
rm -rf "$TMP_DIR"
# ── Summary ────────────────────────────────────────────────────────────────
echo ""
echo "Results: ${PASS} passed, ${FAIL} failed"
if [ $FAIL -gt 0 ]; then
echo " 🔴 TESTS FAILED"
exit 1
else
echo " ✅ ALL TESTS PASSED"
exit 0
fi
+27 -79
View File
@@ -7,22 +7,11 @@
# agent leg is retired — see the note after the Tanko leg.
set -euo pipefail
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
# Never fall back to a literal key.
# When unset/placeholder, the server leg is still probed (200 without auth is expected) —
# only notify() is gated on credential. The pi/Tanko/kagentz
# legs do not need the Zulip API key. The placeholder is captain-held:
# zulip-health-credential-placeholder-20260913.
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_SITE="https://chat.sysloggh.net"
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
OWNER_ZULIP_ID="9"
# Track whether the Zulip API credential is usable
ZULIP_CRED_OK=1
if [ -z "$ZULIP_API_KEY" ] || [[ "$ZULIP_API_KEY" == *"placeholder"* ]] || [[ "$ZULIP_API_KEY" == *"REDACTED"* ]]; then
ZULIP_CRED_OK=0
fi
LOG="/root/zulip-health-monitor.log"
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
@@ -33,32 +22,26 @@ notify() {
local severity="$1" msg="$2"
echo "[$severity] $msg"
# Zulip DM to owner (skip if no credential)
if [ "$ZULIP_CRED_OK" -eq 1 ]; then
local content="${severity} Zulip Monitor: ${msg}"
local form
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
-d "${form}" > /dev/null 2>&1 || true
# Zulip stream post to #agent-hub on topic 'zulip-health'
local stream_content="${severity} Zulip Monitor: ${msg}"
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
> /dev/null 2>&1 \
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)">> "$LOG"
else
echo " ALERT SUPPRESSED (no credential): ${severity} ${msg}" >> "$LOG"
fi
# Zulip DM to owner
local content="${severity} Zulip Monitor: ${msg}"
local form
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
-d "${form}" > /dev/null 2>&1 || true
# Zulip stream post to #agent-hub on topic 'zulip-health'
local stream_content="${severity} Zulip Monitor: ${msg}"
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
> /dev/null 2>&1 \
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
}
# ── Global: Zulip Server ──
# F3: Always probe server regardless of credential — 200 without auth is expected
# (verified live: server_settings returns 200 with no credential or wrong key).
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
https://chat.sysloggh.net/api/v1/server_settings \
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null) || SERVER_CODE="000"
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
if [ "$SERVER_CODE" != "200" ]; then
@@ -135,16 +118,16 @@ case "$PI_VERDICT" in
esac
# -- abiba-leg-end
# ── Platform B: Tanko (DSH dsh-web on minipve CT 112) ──
# ── Platform B: Tanko (DSH dsh-web on amdpve CT 112) ──
# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker
# key availability varies — so probes run from the minipve vantage via `pct exec`.
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on minipve
# (192.168.68.12). The gateway binds 127.0.0.1:3080 loopback-only by design — a
# key availability varies — so probes run from the amdpve vantage via `pct exec`.
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on amdpve
# (192.168.68.15). The gateway binds 127.0.0.1:3080 loopback-only by design — a
# remote :3080 probe is refused and is NOT a fault.
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.12 \
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
"pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true)
[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown"
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.12 \
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
"pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true)
[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000"
@@ -176,11 +159,6 @@ fi
# never contact her former host.
# ── Platform C: Agent Zero (kagentz) ──
# C1: A2A liveness (no credential needed) — probes the container's internal :80/a2a/
# C2: A2A response verification (needs LITELLM_KEY) — probes POST /a2a with auth
# C3: Public access path (no credential needed) — probes https://kagentz.sysloggh.net/
# C1: A2A liveness (container-internal probe)
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
@@ -189,53 +167,23 @@ AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
if [ "$AZ_A2A_CODE" = "000" ]; then
notify "🔴" "kagentz A2A server DOWN (connection failed)"
ISSUES=$((ISSUES + 1))
echo " kagentz C1: ❌ A2A down (HTTP 000)" >> "$LOG"
echo " kagentz: ❌ A2A down (HTTP 000)" >> "$LOG"
else
case "$AZ_A2A_CODE" in
200|401)
echo " kagentz C1: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
echo " kagentz: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
*)
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
ISSUES=$((ISSUES + 1))
echo " kagentz C1: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
esac
fi
# C3: Public access path (the captain's point of view)
# Probes the public URL that NetBird proxies to the container. 200/302/401 = alive,
# 502 = proxy's "upstream refused" page (incident), connection failed = incident.
# Never restarts anything — the contract forbids restarting the platform.
KAGENTZ_PUBLIC_CODE=$(curl -s -o /dev/null --connect-timeout 10 --max-time 15 \
-w '%{http_code}' https://kagentz.sysloggh.net/ 2>/dev/null) || KAGENTZ_PUBLIC_CODE="000"
KAGENTZ_PUBLIC_CODE=$(printf '%s' "$KAGENTZ_PUBLIC_CODE" | tr -d '[:space:]')
[ -n "$KAGENTZ_PUBLIC_CODE" ] || KAGENTZ_PUBLIC_CODE="000"
if [ "$KAGENTZ_PUBLIC_CODE" = "000" ]; then
notify "🔴" "kagentz public URL DOWN (connection failed)"
ISSUES=$((ISSUES + 1))
echo " kagentz C3: ❌ public URL down (HTTP 000)" >> "$LOG"
elif [ "$KAGENTZ_PUBLIC_CODE" = "502" ]; then
notify "🔴" "kagentz public URL 502 (upstream refused)"
ISSUES=$((ISSUES + 1))
echo " kagentz C3: ❌ public URL 502 (upstream refused)" >> "$LOG"
else
case "$KAGENTZ_PUBLIC_CODE" in
200|302|401)
echo " kagentz C3: ✅ public URL alive (HTTP $KAGENTZ_PUBLIC_CODE)" >> "$LOG" ;;
*)
notify "🟡" "kagentz public URL answered HTTP $KAGENTZ_PUBLIC_CODE — running, unexpected status"
ISSUES=$((ISSUES + 1))
echo " kagentz C3: 🟡 public URL unexpected http=$KAGENTZ_PUBLIC_CODE (running, warning)" >> "$LOG" ;;
echo " kagentz: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
esac
fi
# ── Summary ──
# The run verdict is non-optimistic: when ISSUES > 0, the run is an INCIDENT.
# The lane must quote this Result line verbatim in its status report.
if [ "$ISSUES" -eq 0 ]; then
echo " Result: ✅ 0 issues (all healthy)" >> "$LOG"
echo " Result: ✅ All healthy" >> "$LOG"
else
echo " Result: 🔴 INCIDENT — $ISSUES issue(s) found" >> "$LOG"
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
fi
-137
View File
@@ -1,137 +0,0 @@
---
kind: function
name: search-agent-consumption
description: >
Agent-consumption layer in front of SearXNG + Firecrawl. Raw multi-engine
aggregation returns results with no dedupe, no filtering and no reranking;
measured 2026-09-26 that put bestbuy.com and merriam-webster.com into "best
practices agent context management", and put four SEO blogs above the real
Proxmox forum threads on a precise technical query. Identical queries also
ranked DIFFERENTLY between runs, which is why the layer is deterministic
rather than dependent on engine behaviour.
Pipeline: dedupe -> drop non-answers -> demote content farms / promote primary
sources -> stable sort -> extract page text for the top N under an explicit
character budget -> stable JSON. Policy lives in config, not code.
Call it when an agent needs search RESULTS rather than links: it returns usable
page text in one call instead of a snippet plus a second fetch.
version: 1.0.0
---
## Where the policy lives
`config/search-ranking.yaml` — reviewable, no code change needed to adjust:
| key | effect |
| --- | --- |
| `non_answer.hosts` / `path_patterns` / `query_keys` / `host_root` | dropped outright |
| `demote_domains` | ranked below everything, never dropped |
| `prefer_domains` | promoted above default rank |
| `ranking.*` | `demote_penalty`, `prefer_bonus`, `multi_engine_bonus` |
| `extraction.*` | `top_n`, `total_chars`, `per_item_chars`, `timeout_seconds` |
**Demotion, not deletion, for content farms**: a genuinely useful hit is not lost,
it simply cannot outrank a primary source. Non-answers are dropped because they
cannot answer a question at all.
## Usage
```bash
python3 scripts/search-agent-consume.py "query text" # JSON
python3 scripts/search-agent-consume.py --no-extract "query" # ranking only
python3 scripts/search-agent-consume.py --explain "query" # + drop reasons
```
Exit `0` ok, `1` nothing survived filtering, `2` the layer could not run.
## Output shape
Stable JSON:
```json
{
"query": "...",
"raw_result_count": 46,
"returned_count": 44,
"dropped_count": 2,
"engines": ["bing", "brave", "duckduckgo", "yandex"],
"results": [
{"rank": 1, "title": "...", "url": "...", "host": "...",
"source_type": "official|code|qa|forum|discussion|web|content-farm",
"engines": ["bing"], "score": 100.0,
"excerpt": "...", "extraction": "ok|truncated|skipped_budget_exhausted|empty|failed:<Type>"}
],
"extraction": {"extracted": 5, "chars_used": 12000, "budget": 12000,
"failures": 0, "seconds": 5.28}
}
```
`--explain` adds `dropped: [{url, reason, position}]` so the filter is auditable
rather than magic.
## Measured before/after (2026-09-26)
Fixed query set. Relevance judged per query, not by impression.
**`best practices agent context management`**
| | before (raw SearXNG) | after (layer) |
| --- | --- | --- |
| 1-2 | anthropic, stackai | anthropic, langchain |
| 3-4 | aitechmonk, agentic-design | jetbrains, blog.jetbrains |
| 5-6 | mindstudio, sparkco | docs.langchain, reddit |
| 7-8 | langchain, medium | cursor, reddit |
| verdict | 4 relevant of 10; 4 content farms; medium.com twice | top 8 all primary/discussion; no content farm in the top 8 |
**`proxmox thin pool metadata exhaustion recovery`**
| | before | after |
| --- | --- | --- |
| 1-4 | vormox, linuxoperatingsystem, riparazioneserver, bigiron (all SEO/thin) | forum.proxmox.com, forum.proxmox.com, gist.github, github |
| 5-9 | forum.proxmox.com x2, voxfor, github, gist | forum.proxmox.com, serverfault, forum.proxmox.com, reddit |
The primary sources moved from positions 5-9 to 1-4.
**Rule proof** (`--explain`, and a direct check of the classifier):
```
DigitalOcean docs -> KEEP (a '/products/' path rule was REMOVED after the
before/after run caught it dropping this page)
Best Buy -> DROP shopping_or_dictionary_host
Merriam-Webster -> DROP shopping_or_dictionary_host
bare homepage -> DROP navigational_host_root
proxmox.com home -> KEEP (preferred host root: a repo/docs front door is
legitimately the answer)
github repo -> KEEP
```
**Extraction cost (criterion 4):**
```
extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s
whole run end-to-end: 6.4s wall
```
## Regression guard
`search-stack-visibility` asserts the layer still ranks correctly: for the fixed
query set, no `demote_domains` host may appear in the top 3, and the two known
non-answers must not be returned. Without it this layer could silently rot back
to raw ordering, which is exactly what happened to the endpoint colours.
## Reachability, and one honest gap
- **Hermes agents** reach it directly: it reads the same `SEARXNG_URL` and
`FIRECRAWL_URL` they already use.
- **pi agents (MCP search server)**: the MCP server's request/response shape is
**not ours to change**, so this layer is **NOT** wired into it. That is a real
gap, stated rather than claimed as coverage. Closing it would require a change
on the MCP side, which is outside this repo.
## Constraints
Does not touch the live SearXNG or Firecrawl service paths. Third-party
`google cse` is not a hard requirement of this layer — if it 429s, ranking still
works from the remaining engines. No credential is added or required.
-124
View File
@@ -1,124 +0,0 @@
---
kind: function
name: search-stack-visibility
description: >
Makes the shared search stack observable. Every agent reaches one SearXNG
instance (http://192.168.68.7:8888) and one extraction service (Firecrawl,
http://192.168.68.7:3002). Before this check the stack could degrade to a
single engine, or an enabled engine could return nothing at all, without any
error surfacing anywhere.
This contract runs scripts/search-stack-check.py, which:
* runs two fixed queries against SearXNG and FAILS when fewer than two
engines contribute, printing the contributing engines and every
unresponsive_engines entry;
* checks extraction by scraping a known page through Firecrawl and FAILS
when the returned markdown is empty or the request fails;
* reports every silent-zero engine explicitly (enabled, not in
unresponsive_engines, contributed no results).
Multi-engine state (2026-09-25): bing, google cse, brave and yandex
contribute on every query. duckduckgo is NOT working: the house egress IP
and the VPS fallback egress are both flagged by DuckDuckGo and it reports
CAPTCHA. It is left enabled as best-effort coverage so that a recovery shows
up as a contribution.
google cse is a third party's public search-engine id hardcoded in the
SearXNG build. Quota and availability are outside our control.
SCHEDULED: /etc/cron.d/contract-runner on CT 100 (abiba), hourly at :15,
via scripts/contract-run.sh search-stack-visibility. Logs land in
/var/log/contract-runs/. A failure also raises a firstmate inbox note.
version: 1.1.0
---
## Purpose
The fleet has exactly one search endpoint and one extraction endpoint. If
either degrades, every agent silently loses capability at the same moment.
The failure mode this contract exists to close is *silent* degradation: a
query that still returns a page of results while all but one engine have
stopped contributing, or an enabled engine that answers with zero results and
raises no error.
## Execution model
The contract is a host-scheduled check, not an agent workflow. It is driven by
`scripts/contract-run.sh search-stack-visibility` from
`/etc/cron.d/contract-runner` on CT 100. `contract-run.sh` resolves the
mapping to `scripts/search-stack-check.py`, runs it under a timeout, writes a
timestamped log to `/var/log/contract-runs/`, and on non-zero exit raises a
firstmate inbox note through `bin/fm-inbox.sh`.
## What passing looks like
```
$ bash scripts/contract-run.sh search-stack-visibility
Expected engines, enabled (5): ['bing', 'brave', 'duckduckgo', 'google cse', 'yandex']
queries: 'proxmox backup server' -> contributing: bing, brave, google cse, yandex
unresponsive: duckduckgo=CAPTCHA
'python asyncio tutorial' -> contributing: bing, brave, google cse, yandex
EXTRACTION: 71016 chars of markdown returned
VERDICT: PASS -- multiple engines contributing, extraction healthy
```
## What failing looks like
* A query whose results come from fewer than `SEARCH_CHECK_MIN_ENGINES`
engines (default 2) fails and names the engines that did contribute.
* An extraction request that errors or returns empty markdown fails.
## Silent zeros are reported, not fatal
An enabled, expected engine that contributed nothing **without reporting an
error** is printed under `SILENT-ZERO ENGINES REPORTED`, and each occurrence is
annotated `reported, not fatal`. This is deliberate:
* a general query can legitimately draw zero results from an engine that only
fires on certain query shapes, and results are de-duplicated across engines,
so a zero does not by itself prove the engine is broken;
* the run therefore fails only on the two conditions that do prove loss of
capability -- fewer than two contributing engines, and a broken extraction
leg.
A run can consequently print `VERDICT: PASS` while still listing a silent
zero. That is the intended relationship: the zero is *visible*, not *fatal*.
An engine that fails with an error (for example DuckDuckGo returning CAPTCHA)
appears in `unresponsive_engines` instead.
## Google coverage is third-party, not ours
The free Google-derived results come from the SearXNG build's built-in
`google cse` engine. It uses **a third party's public search-engine id
hardcoded in the build** (`google_cse.py`, `CX = "partner-pub-8993..."`,
blackle.com), not a key or id we own. Its quota and availability are outside
our control and it can be rate-limited or withdrawn without notice. No engine
in this build accepts our own Google Custom Search key; using our own free key
would require a small wrapper service, which is deliberately **not** built.
## Configuration
Environment overrides (see the script docstring for the full list):
| Variable | Default | Meaning |
| --- | --- | --- |
| `SEARXNG_URL` | `http://192.168.68.7:8888` | SearXNG base URL |
| `FIRECRAWL_URL` | `http://192.168.68.7:3002` | Firecrawl base URL |
| `SEARCH_CHECK_QUERIES` | `proxmox backup server,python asyncio tutorial` | fixed queries |
| `SEARCH_CHECK_MIN_ENGINES` | `2` | minimum contributing engines per query |
| `SEARCH_CHECK_ENGINES` | `bing,brave,google cse,yandex,duckduckgo` | engines a silent zero is reported for |
| `SEARCH_CHECK_EXTRACT_URL` | Wikipedia Proxmox article | page used for the extraction leg |
## Known residual risk
DuckDuckGo is **not** working. The house egress IP is CAPTCHA'd by
DuckDuckGo, and a forward proxy on the VPS (`10.10.10.1:3128`, WireGuard) was
built as a second egress -- but DuckDuckGo has since flagged the VPS address
too (HTTP 202 with challenge markers), so DuckDuckGo now reports CAPTCHA on
both paths. It is left enabled as best-effort coverage: if DuckDuckGo
unflags either address it will show up as a contribution, and until then it is
visible in `unresponsive_engines` every run. It is never a required engine.
The VPS forward proxy remains a real service (`/opt/fwd-proxy`,
`restart: unless-stopped`, healthy healthcheck, Docker enabled at boot) so the
second egress path is available for any engine that benefits from it in future.
+2 -2
View File
@@ -27,7 +27,7 @@ Agent (Hermes/pi) ──curl + X-API-Key──► Stirling-PDF (:8989) ──►
| Base URL | `http://192.168.68.7:8989` |
| Auth Method | API Key (header) |
| Header Name | `X-API-Key` |
| API Key | `«vault: infrastructure/production STIRLING_API_KEY»` |
| API Key | `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88` |
| Key Source | `SECURITY_CUSTOMGLOBALAPIKEY` in `/opt/home_stack/docker-compose.yml` |
| Swagger | `http://192.168.68.7:8989/swagger-ui.html` |
| Health | `http://192.168.68.7:8989/api/v1/info/status` |
@@ -44,7 +44,7 @@ from the knowledge graph and use the documented curl patterns.
Direct bash invocations:
```bash
curl -X POST "http://192.168.68.7:8989/api/v1/split-pdf" \
-H "X-API-Key: «vault: infrastructure/production STIRLING_API_KEY»" \
-H "X-API-Key: adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88" \
-F "fileInput=@/path/to/file.pdf" \
-F "pageNumbers=1,2,3" \
-o /tmp/output.zip
-40
View File
@@ -1,40 +0,0 @@
#!/usr/bin/env bash
# Revision preflight guard: verify the script being executed matches origin/master
# Usage: revision-preflight.sh <script-path> <clone-path>
# Returns 0 if match, 1 if mismatch (prints both revisions)
set -euo pipefail
SCRIPT="${1:?Usage: revision-preflight.sh <script-path> <clone-path>}"
CLONE="${2:?Usage: revision-preflight.sh <script-path> <clone-path>}"
# Compute sha256 of the script being executed
EXEC_SHA=$(sha256sum "$SCRIPT" | cut -d' ' -f1)
# Compute sha256 of the merged origin/master version
# Extract to a temp file to avoid pipe issues
TMPFILE=$(mktemp)
trap 'rm -f "$TMPFILE"' EXIT
# Try to extract the file from origin/master
if git -C "$CLONE" show "origin/master:$(basename "$SCRIPT")" > "$TMPFILE" 2>/dev/null; then
MASTER_SHA=$(sha256sum "$TMPFILE" | cut -d' ' -f1)
else
echo "⚠️ revision-preflight: could not resolve origin/master revision for $(basename "$SCRIPT")" >&2
exit 0 # Warn but don't block if git show fails
fi
if [[ -z "$MASTER_SHA" || "$MASTER_SHA" == "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" ]]; then
echo "⚠️ revision-preflight: could not resolve origin/master revision for $(basename "$SCRIPT")" >&2
exit 0 # Warn but don't block if git show fails
fi
if [[ "$EXEC_SHA" != "$MASTER_SHA" ]]; then
echo "⚠️ revision-preflight: MISMATCH detected" >&2
echo " Executed: $EXEC_SHA ($(basename "$SCRIPT"))" >&2
echo " Merged: $MASTER_SHA (origin/master:$(basename "$SCRIPT"))" >&2
exit 1
else
echo "✅ revision-preflight: $SCRIPT matches origin/master ($EXEC_SHA)" >&2
exit 0
fi
+2 -72
View File
@@ -24,7 +24,7 @@ BASE = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
base_url: http://192.168.68.116/v1
max_tokens: 4096
default: syslog-auto
provider: harness
@@ -52,7 +52,7 @@ delegation:
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
base_url: http://192.168.68.116/v1
"""
@@ -189,73 +189,3 @@ def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
assert code == 1, out
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
def test_canonical_internal_path_passes(tmp_path):
"""Rule 5 must accept the canonical internal base from hermes-key-enforcement.prose.md."""
code, out = _run(
tmp_path,
"gpu-vision",
)
# Override the base_url in the config
cfg_text = BASE.format(alias="gpu-vision").replace(
" base_url: http://192.168.68.116/v1",
" base_url: http://192.168.68.116/litellm/v1",
)
code, out = _run_config(tmp_path, "canonical-internal.yaml", cfg_text)
assert code == 0, out
assert "RESULT: PASS" in out
# Verify the correct message is shown
assert "model.base_url is canonical" in out
def test_wrong_base_url_fails(tmp_path):
"""Rule 5 must reject paths outside the allowed list."""
cfg_text = BASE.format(alias="gpu-vision").replace(
" base_url: http://192.168.68.116/litellm/v1",
" base_url: http://192.168.68.116/litellm/v1/responses",
)
code, out = _run_config(tmp_path, "wrong-base.yaml", cfg_text)
assert code == 1, out
assert "RESULT: FAIL" in out
assert "model.base_url must be one of" in out
def test_public_host_path_passes(tmp_path):
"""Rule 5 must accept the public host base."""
cfg_text = BASE.format(alias="gpu-vision").replace(
" base_url: http://192.168.68.116/v1",
" base_url: https://litellm.sysloggh.net/v1",
)
code, out = _run_config(tmp_path, "public-host.yaml", cfg_text)
assert code == 0, out
assert "RESULT: PASS" in out
def test_old_rule5_check_would_fail_canonical(tmp_path):
"""
Proof that the OLD Rule 5 check would fail the canonical internal path.
This proves the bug existed before the fix.
"""
# OLD check expected /v1, so the canonical /litellm/v1 would have failed
canonical_cfg = BASE.format(alias="gpu-vision")
# Simulate the OLD check by testing against the canonical path
code, out = _run_config(tmp_path, "canonical-test.yaml", canonical_cfg)
# NEW check: canonical /litellm/v1 SHOULD pass
assert code == 0, out
assert "RESULT: PASS" in out
# OLD check expected /v1, so the internal /v1 would have passed
# NEW check: internal /v1 is non-canonical but working (WARN not FAIL)
old_cfg = BASE.format(alias="gpu-vision").replace(
" base_url: http://192.168.68.116/litellm/v1",
" base_url: http://192.168.68.116/v1",
)
code, out = _run_config(tmp_path, "old-check-test.yaml", old_cfg)
assert code == 0, out
assert "RESULT: PASS" in out
# NEW check should pass
new_cfg = BASE.format(alias="gpu-vision")
code, out = _run_config(tmp_path, "new-check-test.yaml", new_cfg)
assert code == 0, out
assert "RESULT: PASS" in out
@@ -1,236 +0,0 @@
"""Regression test for the fallback_providers list-shape crash in audit-hermes-config.py.
WHY THIS FILE EXISTS: audit-hermes-config.py assumed `fallback_providers` was always a dict
(single provider). Two live agents (koby, koonimo) carry it as a LIST of dicts (one entry per
fallback), so the script crashed with:
File "audit-hermes-config.py", line 211, in audit
fb.get("provider") == "deepseek",
AttributeError: 'list' object has no attribute 'get'
Both are REAL agent configs, so this is not a malformed-input case — the script simply could not
audit two of the four agents it exists to audit. Until fixed, the key-hygiene check had no
coverage for half the fleet while appearing to run.
These tests execute the real CLI (`python3 audit-hermes-config.py <config>`) and assert:
1. A config whose `fallback_providers` is a LIST of valid dicts does NOT crash (exit code is 0 or 1,
never a traceback/AttributeError).
2. A config whose `fallback_providers` contains a MALFORMED entry (a list element that is not a
mapping) reports a VIOLATION naming the offending entry, NOT an uncaught exception.
3. The dict shape still works (existing tests must stay green).
No network, vault, or SSH access is required.
"""
from __future__ import annotations
import pathlib
import subprocess
import sys
ROOT = pathlib.Path(__file__).resolve().parent.parent
AUDIT = ROOT / "audit-hermes-config.py"
# A valid config where fallback_providers is a LIST of dicts (the real koby/koonimo shape).
# One entry, well-formed: provider=deepseek, model=deepseek-v4-flash, api_key_env=DEEPSEEK_API_KEY.
# This must produce a real verdict (PASS or FAIL) without crashing.
LIST_SHAPE_VALID = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
- provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: gpu-vision
provider: harness
web_extract:
model: gpu-vision
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
"""
# A valid config where fallback_providers is a LIST with TWO entries (multiple fallbacks).
# Both entries well-formed. Must not crash and should produce a real verdict.
LIST_SHAPE_MULTI = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
- provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
- provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: gpu-vision
provider: harness
web_extract:
model: gpu-vision
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
"""
# A config where fallback_providers is a LIST containing a MALFORMED entry:
# one element is a plain string, not a mapping. The checker must report a VIOLATION
# naming the offending entry (fallback_providers[1]) and NOT crash.
LIST_SHAPE_MALFORMED = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
- provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
- "not-a-mapping"
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: gpu-vision
provider: harness
web_extract:
model: gpu-vision
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
"""
# The original DICT shape (single provider) must still work — existing behaviour preserved.
DICT_SHAPE_VALID = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: gpu-vision
provider: harness
web_extract:
model: gpu-vision
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
"""
def _run_config(tmp_path, name, text):
cfg = tmp_path / name
cfg.write_text(text)
proc = subprocess.run(
[sys.executable, str(AUDIT), str(cfg)],
capture_output=True, text=True,
)
return proc.returncode, proc.stdout, proc.stderr
def test_list_shape_single_entry_does_not_crash(tmp_path):
"""A LIST with one valid dict must not raise AttributeError; exit 0 (PASS)."""
code, out, err = _run_config(tmp_path, "list-single.yaml", LIST_SHAPE_VALID)
# Must NOT be a crash (traceback). A clean run exits 0 (PASS) or 1 (FAIL), never 2+ (exception).
assert code in (0, 1), f"Expected clean exit 0 or 1, got {code}\nSTDOUT:\n{out}\nSTDERR:\n{err}"
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
# The valid single-entry list should PASS (all rules satisfied).
assert code == 0, f"Expected PASS but got {code}\n{out}"
assert "RESULT: PASS" in out
def test_list_shape_multiple_entries_does_not_crash(tmp_path):
"""A LIST with two valid dicts must not raise AttributeError; exit 0 (PASS)."""
code, out, err = _run_config(tmp_path, "list-multi.yaml", LIST_SHAPE_MULTI)
assert code in (0, 1), f"Expected clean exit 0 or 1, got {code}\nSTDOUT:\n{out}\nSTDERR:\n{err}"
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
assert code == 0, f"Expected PASS but got {code}\n{out}"
assert "RESULT: PASS" in out
def test_list_shape_malformed_entry_reports_violation_not_crash(tmp_path):
"""A LIST containing a non-mapping element must be a reported VIOLATION, not a crash."""
code, out, err = _run_config(tmp_path, "list-malformed.yaml", LIST_SHAPE_MALFORMED)
# Must NOT be a crash.
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
# Should be a FAIL (exit 1) because the malformed entry is a violation.
assert code == 1, f"Expected FAIL (exit 1) but got {code}\n{out}"
assert "RESULT: FAIL" in out
# The violation must name the offending entry (fallback_providers[1]).
assert "fallback_providers[1]" in out, f"Violation did not name the offending entry:\n{out}"
def test_dict_shape_still_passes(tmp_path):
"""The original DICT shape (single provider) must still PASS — existing behaviour preserved."""
code, out, err = _run_config(tmp_path, "dict-valid.yaml", DICT_SHAPE_VALID)
assert code == 0, f"Expected PASS but got {code}\n{out}\nSTDERR:\n{err}"
assert "RESULT: PASS" in out
-80
View File
@@ -1,80 +0,0 @@
#!/bin/bash
# test_contract_run.sh — Tests for contract-run.sh
#
# Proves:
# 1. A passing contract exits 0 and does NOT send an alert
# 2. A failing contract exits non-zero and DOES send an alert
# 3. Log files are created in /var/log/contract-runs/
set -uo pipefail
TEST_DIR="$(cd "$(dirname "$0")" && pwd)"
SCRIPTS_DIR="$(dirname "$TEST_DIR")/scripts"
CONTRACT_RUN="${SCRIPTS_DIR}/contract-run.sh"
LOG_DIR="/var/log/contract-runs"
PASS=0
FAIL=0
# Test 1: Passing contract should exit 0
echo "=== Test 1: Passing contract ==="
# Use a simple passing contract (proxmox-monitor should pass if services are up)
bash "$CONTRACT_RUN" "proxmox-monitor"
EXIT_CODE=$?
if [ $EXIT_CODE -eq 0 ]; then
echo "✅ Test 1 PASSED: contract passed with exit code 0"
PASS=$((PASS + 1))
else
echo "🔴 Test 1 FAILED: expected exit code 0, got $EXIT_CODE"
FAIL=$((FAIL + 1))
fi
# Check log file was created
LATEST_LOG=$(ls -t "$LOG_DIR"/proxmox-monitor-*.log 2>/dev/null | head -1)
if [ -n "$LATEST_LOG" ] && [ -f "$LATEST_LOG" ]; then
echo "✅ Log file created: $LATEST_LOG"
PASS=$((PASS + 1))
else
echo "🔴 Log file not found"
FAIL=$((FAIL + 1))
fi
# Test 2: Failing contract should exit non-zero
echo ""
echo "=== Test 2: Failing contract ==="
# Create a temporary failing contract
TEMP_SCRIPT="${SCRIPTS_DIR}/test-failing-contract.sh"
cat > "$TEMP_SCRIPT" << 'EOF'
#!/bin/bash
echo "This is a test failure"
exit 1
EOF
chmod +x "$TEMP_SCRIPT"
# Temporarily modify contract-run.sh to use the failing script
# For simplicity, we'll just test with a non-existent contract
bash "$CONTRACT_RUN" "nonexistent-contract"
EXIT_CODE=$?
if [ $EXIT_CODE -ne 0 ]; then
echo "✅ Test 2 PASSED: failing contract exited with code $EXIT_CODE"
PASS=$((PASS + 1))
else
echo "🔴 Test 2 FAILED: expected non-zero exit, got 0"
FAIL=$((FAIL + 1))
fi
# Cleanup
rm -f "$TEMP_SCRIPT"
echo ""
echo "=== Summary ==="
echo "Passed: $PASS"
echo "Failed: $FAIL"
if [ $FAIL -eq 0 ]; then
echo "✅ All tests passed"
exit 0
else
echo "🔴 Some tests failed"
exit 1
fi
-76
View File
@@ -6,7 +6,6 @@ Tests:
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
"""
import json
import os
import subprocess
import sys
from pathlib import Path
@@ -25,53 +24,6 @@ def load_script():
return module
def test_pve_token_is_read_from_the_environment():
"""The PVE token must come from the injected environment, never a literal.
Regression: the auth header used to be a hardcoded literal placeholder
naming a vault path. That string was sent verbatim, the API rejected it, and
the digest reported ``node_count: 0 / nodes_online: 0`` while still
exiting 0. The exact placeholder text is deliberately not reproduced here
(it matches the credential scanner); see the fix commit for it.
"""
mod = load_script()
assert hasattr(mod, "pve_auth"), "pve_auth() must exist to resolve the token at call time"
sample = "unit-test-sample-value"
with patch.dict("os.environ", {"PVE_TOKEN": sample}, clear=False):
assert mod.pve_auth() == f"Authorization: {mod.PVE_AUTH_HEADER}{sample}"
assert mod.pve_auth().endswith(sample)
def test_missing_pve_token_is_degraded_not_a_placeholder():
"""With no PVE_TOKEN, pve_get must return None (probe failure), not send a placeholder."""
mod = load_script()
env = {k: v for k, v in os.environ.items() if k != "PVE_TOKEN"}
with patch.dict("os.environ", env, clear=True):
assert mod.pve_get("/api2/json/nodes") is None, (
"a missing PVE_TOKEN must degrade to None so the caller records a probe failure"
)
def test_unreachable_probe_is_recorded_as_a_failure():
"""An unreachable probe must be recorded, so the run cannot pass silently."""
mod = load_script()
assert hasattr(mod, "PROBE_FAILURES"), "PROBE_FAILURES must exist"
mod.PROBE_FAILURES.clear()
with patch.object(mod, "pve_get", return_value=None):
report = mod.collect()
assert report["pve_probe_status"] == "unreachable"
assert any("unreachable" in f for f in mod.PROBE_FAILURES), (
f"unreachable probe must be recorded in PROBE_FAILURES, got {mod.PROBE_FAILURES}"
)
def test_pve_token_placeholder_is_gone():
"""The literal placeholder must no longer appear anywhere in the script."""
src = (Path(__file__).parent.parent / "scripts" / "daily-infra-report.py").read_text()
assert "«vault:" not in src, "the unresolved vault placeholder must not remain in the script"
assert "AUTH = \"Authorization" not in src, "the hardcoded AUTH literal must be gone"
def test_nested_zulip_read_feeds_agent_card():
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
# Mock the http_get_body response with nested structure
@@ -183,32 +135,4 @@ if __name__ == "__main__":
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
sys.exit(1)
try:
test_pve_token_is_read_from_the_environment()
print("✓ test_pve_token_is_read_from_the_environment passed")
except AssertionError as e:
print(f"✗ test_pve_token_is_read_from_the_environment failed: {e}")
sys.exit(1)
try:
test_missing_pve_token_is_degraded_not_a_placeholder()
print("✓ test_missing_pve_token_is_degraded_not_a_placeholder passed")
except AssertionError as e:
print(f"✗ test_missing_pve_token_is_degraded_not_a_placeholder failed: {e}")
sys.exit(1)
try:
test_unreachable_probe_is_recorded_as_a_failure()
print("✓ test_unreachable_probe_is_recorded_as_a_failure passed")
except AssertionError as e:
print(f"✗ test_unreachable_probe_is_recorded_as_a_failure failed: {e}")
sys.exit(1)
try:
test_pve_token_placeholder_is_gone()
print("✓ test_pve_token_placeholder_is_gone passed")
except AssertionError as e:
print(f"✗ test_pve_token_placeholder_is_gone failed: {e}")
sys.exit(1)
print("All tests passed!")
+15 -20
View File
@@ -49,7 +49,7 @@ HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
TANKO_VANTAGE = "192.168.68.12" # minipve — Tanko CT 112 via pct exec
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
@@ -82,7 +82,7 @@ done
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
cmd="${*: -1}"
case "$host" in
192.168.68.12)
192.168.68.15)
case "$cmd" in
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
*curl*) printf '%s' "$TANKO_HTTP" ;;
@@ -98,8 +98,8 @@ exit 0
"""
CURL_STUB = r"""#!/usr/bin/env bash
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
# kagentz C3 public URL, and record every call (including notify) payloads.
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
# record every call (including notify) payloads.
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
case "$*" in
*:9200/health*)
@@ -109,8 +109,6 @@ case "$*" in
esac ;;
*server_settings*)
printf '%s' "$SERVER_HTTP" ;;
*kagentz.sysloggh.net*)
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
esac
exit 0
"""
@@ -123,8 +121,7 @@ def _write_exec(path: pathlib.Path, body: str) -> None:
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
az_a2a_code="401", az_a2a_exit=0,
kagentz_public_code="302"):
az_a2a_code="401", az_a2a_exit=0):
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
Only the LOG constant is rewritten (to keep the run inside the worktree).
@@ -157,7 +154,6 @@ def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
"PI_HTTP": "200",
"PI_BODY": CONNECTED_FIXTURE.read_text(),
"SERVER_HTTP": "200",
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
})
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
capture_output=True, text=True)
@@ -173,9 +169,8 @@ def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
assert "Server: ✅ HTTP 200" in log
assert "Abiba: ✅ Connected" in log
assert "Tanko: ✅ service=active http=200" in log
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
assert "Result: ✅ 0 issues (all healthy)" in log
assert "kagentz: ✅ A2A alive (HTTP 401)" in log
assert "Result: ✅ All healthy" in log
# A healthy run emits no notify at all — and certainly no Mumuni one.
assert proc.stdout == ""
@@ -212,8 +207,8 @@ def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
# The rest of the monitor still ran alongside the failing Tanko leg.
log = log_path.read_text()
assert "Abiba: ✅ Connected" in log
assert "kagentz C1: ✅ A2A alive" in log
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
assert "kagentz: ✅ A2A alive" in log
assert "Result: 🔴 1 issue(s) found" in log
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
@@ -221,9 +216,9 @@ def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
assert "kagentz C1: 🟡 A2A unexpected http=500 (running, warning)" in log
assert "kagentz C1: ✅ A2A alive" not in log
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
assert "kagentz: 🟡 A2A unexpected http=500 (running, warning)" in log
assert "kagentz: ✅ A2A alive" not in log
assert "Result: 🔴 1 issue(s) found" in log
assert "kagentz A2A server answered HTTP 500" in proc.stdout
@@ -235,10 +230,10 @@ def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
assert "kagentz C1: ✅ A2A alive" not in log
assert "kagentz: ❌ A2A down (HTTP 000)" in log
assert "kagentz: ✅ A2A alive" not in log
assert "unexpected" not in log
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
assert "Result: 🔴 1 issue(s) found" in log
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
-108
View File
@@ -1,108 +0,0 @@
#!/bin/bash
# test_pbs_gc_states.sh — Tests for PBS GC four-state logic
# Self-contained: inlines the SSH replacement logic
set -uo pipefail
TEST_DIR="$(cd "$(dirname "$0")" && pwd)"
SCRIPTS_DIR="$(dirname "$TEST_DIR")/scripts"
PROXMOX_MONITOR="${1:-${SCRIPTS_DIR}/proxmox-monitor.sh}"
PASS=0
FAIL=0
# Create the Python replacement script
REPLACE_SCRIPT=$(mktemp /tmp/replace_ssh_XXXXXX.py)
cat > "$REPLACE_SCRIPT" << 'PYEOF'
import sys
import re
wrapper = sys.argv[1]
monitor = sys.argv[2]
with open(monitor) as f:
c = f.read()
pattern = r'PBS_GC_OUTPUT=\$\(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192\.168\.68\.6 \\\n "/sbin/pct exec 107 -- /sbin/proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null\)'
replacement = 'PBS_GC_OUTPUT=$(cat "' + wrapper + '")'
if re.search(pattern, c):
c = re.sub(pattern, replacement, c)
with open(monitor, 'w') as f:
f.write(c)
PYEOF
run_test() {
local name="$1"
local json="$2"
local expected_behavior="$3"
local expected_pattern="$4"
local wrapper monitor
wrapper=$(mktemp /tmp/pbs-test-wrapper.XXXXXX)
monitor=$(mktemp /tmp/pbs-test-monitor.XXXXXX)
printf '%s\n' "$json" > "$wrapper"
cp "$PROXMOX_MONITOR" "$monitor"
# Replace the SSH call with cat "$wrapper"
python3 "$REPLACE_SCRIPT" "$wrapper" "$monitor"
local output exit_code
output=$(bash "$monitor" 2>&1)
exit_code=$?
local ok=true
if [ "$expected_behavior" = "fail" ]; then
# Should fail with PBS GC error
if ! echo "$output" | grep -q "🔴 PBS GC"; then ok=false; fi
if [ $exit_code -eq 0 ]; then ok=false; fi
else
# Should pass with expected pattern
if ! echo "$output" | grep -q "$expected_pattern"; then ok=false; fi
if echo "$output" | grep -q "🔴 PBS GC"; then ok=false; fi
fi
if $ok; then
echo " ✅ $name"
PASS=$((PASS + 1))
else
echo " 🔴 $name FAILED (exit=$exit_code)"
echo "$output" | grep "PBS GC" | sed 's/^/ /'
FAIL=$((FAIL + 1))
fi
rm -f "$wrapper" "$monitor"
}
echo "=== PBS GC Four-State Tests ==="
echo "Script: $PROXMOX_MONITOR"
echo ""
echo "1. probe-failed (unparseable JSON)"
run_test "unparseable-json" 'NOT JSON {{{' "fail" "probe-failed"
echo "2. probe-failed (store not found)"
run_test "store-missing" '[{"store": "other", "last-run-endtime": 1000}]' "fail" "probe-failed"
echo "3. probe-failed (empty body)"
run_test "empty-body" "" "fail" "probe-failed"
echo "4. running (in progress - upid set, no last-run-endtime)"
run_test "running" '[{"store": "storepve-datastore", "upid": "UPID:123:1:456:gc:root@pam", "pending-bytes": 100}]' "pass" "⏳ PBS GC: running"
echo "5. stale (last run >48h)"
STALE=$(date -u -d "50 hours ago" +%s)
run_test "stale" "[{\"store\": \"storepve-datastore\", \"last-run-endtime\": $STALE, \"pending-bytes\": 200}]" "fail" "stale"
echo "6. healthy (completed <48h)"
HEALTHY=$(date -u -d "1 hour ago" +%s)
run_test "healthy" "[{\"store\": \"storepve-datastore\", \"last-run-endtime\": $HEALTHY, \"pending-bytes\": 0}]" "pass" "✅ PBS GC: healthy"
echo ""
echo "=== Results: $PASS passed, $FAIL failed ==="
[ $FAIL -eq 0 ] && echo "✅ All passed" || echo "🔴 Some failed"
rm -f "$REPLACE_SCRIPT"
[ $FAIL -eq 0 ] && exit 0 || exit 1
-53
View File
@@ -104,59 +104,6 @@ def test_koby_ct111_is_on_storepve(ahc):
assert ahc.AGENTS["koby"]["pve"] == "storepve"
def test_tanko_ct112_is_probed_on_minipve(ahc, monkeypatch, capsys):
# CT 112 (tanko) was live-migrated to minipve (.12) on 2026-09-27; the
# amdpve mapping made `pct status 112` fail and read as ct-unreachable.
# Execute the probe and assert the host the script actually contacts.
probes = []
monkeypatch.setattr(
ahc, "ssh",
lambda host, cmd, user="root": probes.append((host, cmd)) or "status: running",
)
ahc.FAIL.clear()
ahc.REPORT_ONLY.clear()
try:
ahc.check_ct_liveness()
tanko_hosts = [h for h, cmd in probes if cmd == "pct status 112 2>/dev/null"]
assert tanko_hosts == ["192.168.68.12"]
finally:
ahc.FAIL.clear()
ahc.REPORT_ONLY.clear()
def test_gpu_rtx3090_probe_uses_llmuser_not_root(ahc, monkeypatch, capsys):
# 2026-09-28: root SSH to .8 was lost when the guest was rebuilt; llmuser
# owns llama-server and can read systemctl status and the :8080 pid. A root
# probe reads as UNREACHABLE for a healthy host (the reported bug). Execute
# check_gpu_ports() against an SSH boundary that only accepts llmuser@.8 and
# assert the .8 leg does not produce the false UNREACHABLE failure.
seen = []
def fake_ssh(host, cmd, user="root"):
seen.append((host, user))
if host == "192.168.68.8" and user != "llmuser":
return None # root SSH denied -> baseline false UNREACHABLE
if cmd.startswith("systemctl is-active"):
return "active"
if cmd.startswith("ss -tlnp"):
return "48351"
if cmd.startswith("curl"):
return '{"status":"ok"}'
return None
monkeypatch.setattr(ahc, "ssh", fake_ssh)
ahc.FAIL.clear()
try:
ahc.check_gpu_ports()
out = capsys.readouterr().out
assert "gpu-unreachable:192.168.68.8" not in ahc.FAIL
assert "\u2705 gpu-rtx3090 (.8): healthy" in out
assert ("192.168.68.8", "llmuser") in seen
assert not any(host == "192.168.68.8" and user == "root" for host, user in seen)
finally:
ahc.FAIL.clear()
def test_report_only_legs_never_count_as_failures(ahc):
for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)):
ahc.FAIL.clear()
-282
View File
@@ -1,282 +0,0 @@
#!/usr/bin/env bash
# Behavioural tests for scripts/revision-preflight.sh
#
# Every case builds a throwaway clone with a real bare remote, so origin/master
# is genuine and the guard's fetch path is exercised. Nothing outside mktemp is
# touched.
#
# The pre-fix draft is kept at tests/fixtures/revision-preflight.prefix.sh and
# is run against the SAME cases, to prove these tests bite: the pre-fix guard
# exits 0 where the fixed guard exits 1.
set -uo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/.." && pwd)"
GUARD="$REPO/scripts/revision-preflight.sh"
PREFIX_GUARD="$HERE/fixtures/revision-preflight.prefix.sh"
PASS=0
FAIL=0
FAILED_CASES=()
pass() { printf ' ✓ %s\n' "$1"; PASS=$((PASS + 1)); }
fail() { printf ' ✗ %s\n' "$1"; FAIL=$((FAIL + 1)); FAILED_CASES+=("$1"); }
# Build a clone with a real remote; echo the clone path.
make_clone() {
local tmp
tmp="$(mktemp -d)"
git init --bare -q "$tmp/remote.git"
git init -q "$tmp/clone"
(
cd "$tmp/clone" || exit 1
git config user.email test@example.invalid
git config user.name test
mkdir -p scripts
printf '#!/bin/bash\necho hello\n' > scripts/demo.sh
chmod +x scripts/demo.sh
git add -A
git commit -qm init
git branch -M master
git remote add origin "$tmp/remote.git"
git push -q origin master
git fetch -q origin
)
echo "$tmp/clone"
}
echo "== revision-preflight behavioural tests =="
# ── 1. match → exit 0 ────────────────────────────────────────────────────────
echo "1. matching copy"
C=$(make_clone)
if out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); then
pass "matching copy exits 0"
else
fail "matching copy should exit 0 (got $?, output: $out)"
fi
if [[ -z "$( "$GUARD" --quiet "$C/scripts/demo.sh" "$C" 2>&1 )" ]]; then
pass "--quiet prints nothing on a match"
else
fail "--quiet should print nothing on a match"
fi
rm -rf "$(dirname "$C")"
# ── 2. mismatch → exit 1 and names both hashes ───────────────────────────────
echo "2. mismatched copy"
C=$(make_clone)
printf '#!/bin/bash\necho TAMPERED\n' > "$C/scripts/demo.sh"
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 1 ]]; then pass "mismatch exits 1"; else fail "mismatch should exit 1 (got $rc)"; fi
if [[ "$out" == *"MISMATCH"* ]]; then pass "mismatch says MISMATCH"; else fail "mismatch should say MISMATCH"; fi
if [[ "$out" == *"executed:"* && "$out" == *"merged:"* ]]; then
pass "mismatch prints both revisions"
else
fail "mismatch should print both revisions"
fi
rm -rf "$(dirname "$C")"
# ── 3. paths under scripts/ resolve (the original basename defect) ───────────
echo "3. repo-relative path resolution"
C=$(make_clone)
if "$GUARD" --quiet "$C/scripts/demo.sh" "$C" >/dev/null 2>&1; then
pass "script under scripts/ resolves against origin/master"
else
fail "script under scripts/ must resolve (basename defect)"
fi
rm -rf "$(dirname "$C")"
# ── 4. script that exists in NO revision → must fail ─────────────────────────
echo "4. untracked script present in no revision"
C=$(make_clone)
printf '#!/bin/bash\necho never committed\n' > "$C/scripts/ghost.sh"
out=$("$GUARD" "$C/scripts/ghost.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 1 ]]; then pass "ghost script exits 1"; else fail "ghost script must exit 1 (got $rc)"; fi
if [[ "$out" == *"does not exist in"* ]]; then
pass "ghost script says it is absent from the ref"
else
fail "ghost script should say it is absent from the ref"
fi
rm -rf "$(dirname "$C")"
reason_of() { grep -m1 '^REASON=' <<<"$1" | cut -d= -f2-; }
# ── 5. unresolvable ref → CANNOT VERIFY (exit 2) ─────────────────────────────
echo "5. unresolvable ref"
C=$(make_clone)
out=$("$GUARD" --no-fetch --ref origin/nope "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 2 ]]; then pass "unresolvable ref exits 2 (cannot verify)"; else fail "unresolvable ref must exit 2 (got $rc)"; fi
if [[ "$(reason_of "$out")" == "cannot-verify:ref-unresolvable" ]]; then
pass "unresolvable ref names cannot-verify:ref-unresolvable"
else
fail "unresolvable ref should name its class (got: $(reason_of "$out"))"
fi
rm -rf "$(dirname "$C")"
# ── 5b. fetch failure → CANNOT VERIFY, and it names that ─────────────────────
echo "5b. fetch failed"
C=$(make_clone)
( cd "$C" && git remote set-url origin /nonexistent/definitely-not-a-repo )
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 2 ]]; then pass "fetch failure exits 2 (cannot verify)"; else fail "fetch failure must exit 2 (got $rc)"; fi
if [[ "$(reason_of "$out")" == "cannot-verify:fetch-failed" ]]; then
pass "fetch failure names cannot-verify:fetch-failed"
else
fail "fetch failure should name its class (got: $(reason_of "$out"))"
fi
if [[ "$out" == *"bound:"* ]]; then pass "fetch failure reports the bound"; else fail "fetch failure should report the bound"; fi
rm -rf "$(dirname "$C")"
# ── 5c. fetch timeout → CANNOT VERIFY, bounded (never hangs) ─────────────────
echo "5c. fetch timeout is bounded"
C=$(make_clone)
# a remote that will never answer: a fifo-backed git daemon is overkill, so use
# a black-hole address with a 1s bound and assert we return promptly.
( cd "$C" && git remote set-url origin http://10.255.255.1:9/never.git )
start=$(date +%s)
out=$("$GUARD" --fetch-timeout 1 "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
elapsed=$(( $(date +%s) - start ))
if [[ $rc -eq 2 ]]; then pass "timeout exits 2 (cannot verify)"; else fail "timeout must exit 2 (got $rc)"; fi
if [[ "$(reason_of "$out")" == "cannot-verify:fetch-failed" ]]; then
pass "timeout names cannot-verify:fetch-failed"
else
fail "timeout should name its class (got: $(reason_of "$out"))"
fi
if [[ $elapsed -le 10 ]]; then pass "timeout returned promptly (${elapsed}s, bound 1s)"; else fail "timeout did not bound the fetch (${elapsed}s)"; fi
rm -rf "$(dirname "$C")"
# ── 6. missing script → MISMATCH:path-absent (exit 1) ────────────────────────
echo "6. missing script"
C=$(make_clone)
out=$("$GUARD" "$C/scripts/nope.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 1 ]]; then pass "missing script exits 1"; else fail "missing script must exit 1 (got $rc)"; fi
if [[ "$(reason_of "$out")" == "mismatch:path-absent" ]]; then
pass "missing script names mismatch:path-absent"
else
fail "missing script should name its class (got: $(reason_of "$out"))"
fi
rm -rf "$(dirname "$C")"
# ── 7. script outside the clone → MISMATCH (exit 1) ──────────────────────────
echo "7. script outside the clone"
C=$(make_clone)
OUTSIDE=$(mktemp)
printf '#!/bin/bash\necho outside\n' > "$OUTSIDE"
out=$("$GUARD" "$OUTSIDE" "$C" 2>&1); rc=$?
if [[ $rc -eq 1 ]]; then pass "outside script exits 1"; else fail "outside script must exit 1 (got $rc)"; fi
rm -f "$OUTSIDE"; rm -rf "$(dirname "$C")"
# ── 7b. detached HEAD is named, not reported as a raw content mismatch ───────
echo "7b. detached HEAD"
C=$(make_clone)
( cd "$C" && printf '#!/bin/bash\necho TAMPERED\n' > scripts/demo.sh \
&& git add scripts/demo.sh && git commit -qm tamper && git checkout -q --detach HEAD )
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 1 ]]; then pass "detached HEAD exits 1"; else fail "detached HEAD must exit 1 (got $rc)"; fi
if [[ "$(reason_of "$out")" == "mismatch:detached-head" ]]; then
pass "detached HEAD names mismatch:detached-head"
else
fail "detached HEAD should name its class (got: $(reason_of "$out"))"
fi
rm -rf "$(dirname "$C")"
# ── 7c. a branch ahead of the ref is named as such, not as a raw mismatch ────
echo "7c. clone ahead of the ref"
C=$(make_clone)
( cd "$C" && git checkout -q -b feature \
&& printf '#!/bin/bash\necho FEATURE\n' > scripts/demo.sh \
&& git add scripts/demo.sh && git commit -qm feature )
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 1 ]]; then pass "clone ahead exits 1"; else fail "clone ahead must exit 1 (got $rc)"; fi
if [[ "$(reason_of "$out")" == "mismatch:clone-ahead" ]]; then
pass "clone ahead names mismatch:clone-ahead"
else
fail "clone ahead should name its class (got: $(reason_of "$out"))"
fi
if [[ "$out" == *"mid-review"* ]]; then pass "clone ahead explains it is a legitimate state"; else fail "clone ahead should explain the state"; fi
rm -rf "$(dirname "$C")"
# ── 7d. a genuine content mismatch is named as content ──────────────────────
echo "7d. genuine content mismatch"
C=$(make_clone)
printf '#!/bin/bash\necho TAMPERED\n' > "$C/scripts/demo.sh"
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
if [[ "$(reason_of "$out")" == "mismatch:content" ]]; then
pass "hand-edit names mismatch:content"
else
fail "hand-edit should name mismatch:content (got: $(reason_of "$out"))"
fi
rm -rf "$(dirname "$C")"
# ── 8. --no-fetch states the freshness assumption ────────────────────────────
echo "8. --no-fetch states its assumption"
C=$(make_clone)
out=$("$GUARD" --no-fetch "$C/scripts/demo.sh" "$C" 2>&1)
if [[ "$out" == *"freshness is assumed"* ]]; then
pass "--no-fetch states the freshness assumption"
else
fail "--no-fetch should state the freshness assumption"
fi
rm -rf "$(dirname "$C")"
# ── 8b. contract-run.sh defaults to warn, not enforce ───────────────────────
echo "8b. contract-run.sh default mode"
DEFAULT=$(grep -m1 'CONTRACT_REVISION_PREFLIGHT:-' "$REPO/scripts/contract-run.sh" | sed 's/.*:-//; s/}.*//')
if [[ "$DEFAULT" == "warn" ]]; then
pass "contract-run.sh defaults to warn"
else
fail "contract-run.sh default must be warn (found: '$DEFAULT')"
fi
if grep -q 'CONTRACT_REVISION_PREFLIGHT=enforce' "$REPO/scripts/contract-run.sh"; then
pass "enforce remains available and documented"
else
fail "enforce must remain documented"
fi
# ── 9. the pre-fix guard must FAIL these same cases (proves the tests bite) ──
echo "9. pre-fix draft fails the same cases (bite proof)"
if [[ ! -f "$PREFIX_GUARD" ]]; then
fail "pre-fix fixture missing: $PREFIX_GUARD"
else
# 9a. repo-relative path: pre-fix drops scripts/ and cannot resolve
C=$(make_clone)
out=$("$PREFIX_GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 0 && "$out" == *"could not resolve"* ]]; then
pass "pre-fix: exits 0 and cannot resolve scripts/demo.sh (defect confirmed)"
else
fail "pre-fix should exit 0 with 'could not resolve' (got rc=$rc)"
fi
rm -rf "$(dirname "$C")"
# 9b. ghost script: pre-fix passes a script that exists in no revision
C=$(make_clone)
printf '#!/bin/bash\necho never committed\n' > "$C/scripts/ghost.sh"
out=$("$PREFIX_GUARD" "$C/scripts/ghost.sh" "$C" 2>&1); rc=$?
if [[ $rc -eq 0 ]]; then
pass "pre-fix: PASSES a ghost script that exists in no revision (defect confirmed)"
else
fail "pre-fix was expected to wrongly pass the ghost script (got rc=$rc)"
fi
rm -rf "$(dirname "$C")"
# 9c. a file that DOES exist at the repo root still works pre-fix, showing
# the defect is specific to nested paths
C=$(make_clone)
printf '#!/bin/bash\necho root\n' > "$C/rootlevel.sh"
( cd "$C" && git add rootlevel.sh && git commit -qm root && git push -q origin master && git fetch -q origin )
if "$PREFIX_GUARD" "$C/rootlevel.sh" "$C" >/dev/null 2>&1; then
pass "pre-fix: root-level path resolves (so the defect is the basename, not git)"
else
fail "pre-fix should resolve a root-level tracked file"
fi
rm -rf "$(dirname "$C")"
fi
echo
echo " passed: $PASS failed: $FAIL"
if [[ $FAIL -gt 0 ]]; then
printf ' FAILED: %s\n' "${FAILED_CASES[@]}"
exit 1
fi
echo "All revision-preflight tests passed."
-153
View File
@@ -1,153 +0,0 @@
#!/usr/bin/env bash
# test_secret_scan.sh — self-test for the commit-time secret guard.
#
# Run: bash tests/test_secret_scan.sh
# Exit: 0 all cases passed, 1 a case failed.
#
# WHY THIS FILE EXISTS: a scanner that is never observed to fail is not a guard.
# Every fixture below is fabricated and pattern-shaped; the test writes it to a
# temp tree (a path no allowlist entry covers) and asserts the guard FAILS. The
# same fixtures are deliberately listed in scripts/secret-allowlist.tsv, so the
# repo-wide tree scan stays quiet while a planted copy still bites — that is the
# difference between an explicit, reasoned exception and a guard trained to
# ignore a word.
#
# Only bash + coreutils + grep. No python/node: the Gitea runner executes job
# steps inside the runner container, which has neither.
set -uo pipefail
HERE=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
ROOT=$(cd -- "$HERE/.." && pwd)
SCAN="$ROOT/scripts/secret-scan.sh"
PASS=0
FAIL=0
LAST_OUT=""
ok() { PASS=$((PASS + 1)); echo " ✅ $1"; }
bad() { FAIL=$((FAIL + 1)); echo " ❌ $1"; }
expect_exit() { # expect_exit <want-code> <label> <cmd...>
local want="$1" label="$2"; shift 2
local rc
LAST_OUT=$("$@" 2>&1); rc=$?
if [ "$rc" -eq "$want" ]; then ok "$label (exit $rc)"; else
bad "$label (wanted exit $want, got $rc)"
printf '%s\n' "$LAST_OUT" | sed 's/^/ /' | head -8
fi
}
expect_contains() { # expect_contains <label> <needle>
if printf '%s' "$LAST_OUT" | grep -qF -- "$2"; then ok "$1"; else
bad "$1 (output did not mention: $2)"
fi
}
TMPROOT=$(mktemp -d)
trap 'rm -rf "$TMPROOT"' EXIT
echo "── secret-scan self-test ──"
# ── 1. Guard syntax ───────────────────────────────────────────────────────
expect_exit 0 "scanner parses with bash -n" bash -n "$SCAN"
# ── 2. Guard FAILS on planted, pattern-matching fixtures ──────────────────
mkdir -p "$TMPROOT/planted"
cat > "$TMPROOT/planted/ops.env" <<'EOF'
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
EOF
expect_exit 1 "planted sk-or-v1 key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted sk-or-v1 key names the openrouter-key rule" "[openrouter-key]"
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/curl.sh" <<'EOF'
curl -s -H "Authorization: Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb" http://example.invalid/
EOF
expect_exit 1 "planted literal Bearer token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted Bearer token names the bearer-token rule" "[bearer-token]"
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/pve.sh" <<'EOF'
AUTH="Authorization: PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555"
EOF
expect_exit 1 "planted Proxmox token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted Proxmox token names the proxmox-token rule" "[proxmox-token]"
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/deploy-key.pem" <<'EOF'
-----BEGIN OPENSSH PRIVATE KEY-----
b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW
-----END OPENSSH PRIVATE KEY-----
EOF
expect_exit 1 "planted PEM private key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted PEM key names the private-key rule" "[private-key]"
# Prose is scanned exactly like code — the original exposures were in .md files.
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/handover.md" <<'EOF'
- Admin credentials: `admin` / `correct-horse-battery-staple`
EOF
expect_exit 1 "planted prose credential line fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted prose line names the cred-prose rule" "[cred-prose]"
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/config.env" <<'EOF'
DB_PASSWORD=correct-horse-battery-staple
EOF
expect_exit 1 "planted password assignment fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted password assignment names the secret-assign rule" "[secret-assign]"
# ── 3. Guard stays QUIET on inert values and on the real tree ─────────────
mkdir -p "$TMPROOT/inert"
cat > "$TMPROOT/inert/config.yaml" <<'EOF'
api_key: not-needed
bearer_token=monitor_key
api_key: $LITELLM_API_KEY
EOF
expect_exit 0 "env refs, sentinels and variable names are not credentials" bash "$SCAN" --path "$TMPROOT/inert" --quiet
expect_exit 0 "current repo tree passes the guard" bash "$SCAN" --tree
expect_contains "tree run reports the allowlisted exceptions it applied" "allowlisted exception(s)"
# ── 4. Allowlist entries are path-explicit, not word-based ────────────────
# This exact line is allowlisted in infrastructure-control.prose.md; the same
# text at an unlisted path must still fail, proving the exception is per-file
# and reviewed, not a blanket "ignore the word vault".
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/unlisted.md" <<'EOF'
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
EOF
expect_exit 1 "allowlisted text at an unlisted path still fails" bash "$SCAN" --path "$TMPROOT/planted" --quiet
# ── 5. Commit-time mode: the guard blocks a STAGED credential ─────────────
# A throwaway git repo with its own copy of the scanner, so this exercises the
# real pre-commit path (--staged) without touching this repo's index.
mkdir -p "$TMPROOT/repo/scripts"
cp "$SCAN" "$TMPROOT/repo/scripts/secret-scan.sh"
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/repo/scripts/secret-patterns.tsv"
cp "$ROOT/scripts/secret-allowlist.tsv" "$TMPROOT/repo/scripts/secret-allowlist.tsv"
git -C "$TMPROOT/repo" init -q
git -C "$TMPROOT/repo" -c user.email=t@example.invalid -c user.name=test commit -q --allow-empty -m base
cat > "$TMPROOT/repo/planted.env" <<'EOF'
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
EOF
git -C "$TMPROOT/repo" add planted.env
expect_exit 1 "staged credential fails at commit time (--staged)" bash "$TMPROOT/repo/scripts/secret-scan.sh" --staged --quiet
expect_contains "staged credential names the openrouter-key rule" "[openrouter-key]"
# ── 6. Fail closed: an allowlist entry without a reason is a hard error ───
mkdir -p "$TMPROOT/scanner" "$TMPROOT/clean"
cp "$SCAN" "$TMPROOT/scanner/secret-scan.sh"
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/scanner/secret-patterns.tsv"
printf '*\t*.md\twhatever\n' > "$TMPROOT/scanner/secret-allowlist.tsv"
echo "placeholder" > "$TMPROOT/clean/ok.md"
expect_exit 2 "allowlist entry with no reason fails closed" bash "$TMPROOT/scanner/secret-scan.sh" --path "$TMPROOT/clean" --quiet
# ── Verdict ───────────────────────────────────────────────────────────────
echo ""
if [ "$FAIL" -gt 0 ]; then
echo "❌ secret-scan self-test FAILED — $PASS passed, $FAIL failed"
exit 1
fi
echo "✅ secret-scan self-test passed ($PASS cases)"
-209
View File
@@ -1,209 +0,0 @@
#!/usr/bin/env python3
"""Behavioural tests for zulip-monitor.sh kagentz C1/C3 legs and the Result verdict.
WHY THIS FILE EXISTS: the 2026-09-19 kagentz A2A outage was correctly detected
by the monitor (C1 returned 000, the run logged an issue) but the lane's own
summarization in ops.status wrote "OK" with the note "A2A server DOWN — expected
(no credentials configured)". The optimistic verdict came from the lane, not the
script. The fix adds a C3 public-access-path leg and makes the Result line say
"INCIDENT" when issues are found, so the lane can quote it verbatim.
CONTRACT UNDER TEST:
* C1 (A2A liveness, no credential): 000 → INCIDENT.
* C3 (public access path, no credential): 502 → INCIDENT, 000 → INCIDENT,
200/302/401 → alive.
* Result verdict: when ISSUES > 0, the log's Result line says "INCIDENT",
not just "issues found".
* Healthy control: C1 401 + C3 302 → 0 issues, "all healthy".
HOW: behavioural execution using the sandbox pattern already in this repo
(tests/test_mumuni_monitor_removal.py). The sandbox copies the shipped monitor
verbatim, rewrites only its LOG constant, and runs it with stub ssh/curl on
PATH. Each test asserts from the run's own log/verdict, not from file text.
Usage: python3 -m pytest tests/test_zulip_kagentz_legs.py
"""
from __future__ import annotations
import os
import pathlib
import stat
import subprocess
ROOT = pathlib.Path(__file__).resolve().parents[1]
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
TANKO_VANTAGE = "192.168.68.12" # minipve — Tanko CT 112 via pct exec
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
# ── Stub ssh: answers Tanko and Agent Zero probes by env vars ──────────
SSH_STUB = r"""#!/usr/bin/env bash
# Stub ssh: record the target host, then answer by host + remote command.
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
host=""
for a in "$@"; do
case "$a" in
*@192.168.*) host="${a##*@}" ;;
esac
done
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
cmd="${*: -1}"
case "$host" in
192.168.68.12)
case "$cmd" in
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
*curl*) printf '%s' "$TANKO_HTTP" ;;
esac ;;
192.168.68.14)
case "$cmd" in
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "${AZ_A2A_EXIT:-0}" ;;
esac ;;
*)
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
esac
exit 0
"""
# ── Stub curl: serves Zulip server, Abiba health, and C3 public URL ───
CURL_STUB = r"""#!/usr/bin/env bash
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
# C3 public URL probe (https://kagentz.sysloggh.net/). Record every call.
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
case "$*" in
*:9200/health*)
case " $* " in
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
*) printf '%s' "$PI_BODY" ;; # body probe
esac ;;
*server_settings*)
printf '%s' "$SERVER_HTTP" ;;
*kagentz.sysloggh.net*)
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
esac
exit 0
"""
def _write_exec(path: pathlib.Path, body: str) -> None:
path.write_text(body)
path.chmod(path.stat().st_mode
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
az_a2a_code="401", az_a2a_exit=0,
kagentz_public_code="302"):
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
Only the LOG constant is rewritten (to keep the run inside the worktree).
Everything else — legs, labels, notify logic — is the shipped script.
"""
sandbox = tmp_path / "sandbox"
bindir = sandbox / "bin"
record = sandbox / "record"
bindir.mkdir(parents=True)
record.mkdir()
_write_exec(bindir / "ssh", SSH_STUB)
_write_exec(bindir / "curl", CURL_STUB)
source = ZULIP_MONITOR.read_text()
log_line = 'LOG="/root/zulip-health-monitor.log"'
assert log_line in source, "LOG constant moved — update the sandbox harness"
log_path = sandbox / "zulip-health-monitor.log"
script = sandbox / "zulip-monitor.sh"
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
env = dict(os.environ)
env.update({
"PATH": f"{bindir}:{env['PATH']}",
"RECORD_DIR": str(record),
"TANKO_SVC": tanko_svc,
"TANKO_HTTP": tanko_http,
"AZ_A2A_CODE": az_a2a_code,
"AZ_A2A_EXIT": str(az_a2a_exit),
"PI_HTTP": "200",
"PI_BODY": CONNECTED_FIXTURE.read_text(),
"SERVER_HTTP": "200",
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
})
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
capture_output=True, text=True)
return proc, record, log_path
# ── Required behavioural cases ─────────────────────────────────────────
def test_c3_502_is_incident(tmp_path):
"""C3 public leg returns 502 → the run's verdict is an INCIDENT,
the C3 line names the 502, and the run is not summarised as healthy."""
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="502")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# The C3 line names the 502.
assert "kagentz C3: ❌ public URL 502 (upstream refused)" in log
# The verdict is an INCIDENT, not healthy.
assert "Result: 🔴 INCIDENT" in log
assert "all healthy" not in log
# The run is not summarised as healthy.
assert "✅ 0 issues" not in log
def test_c3_000_is_incident(tmp_path):
"""C3 public leg returns 000 → INCIDENT."""
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="000")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# The C3 line reports the connection failure.
assert "kagentz C3: ❌ public URL down (HTTP 000)" in log
# The verdict is an INCIDENT.
assert "Result: 🔴 INCIDENT" in log
assert "all healthy" not in log
assert "✅ 0 issues" not in log
def test_healthy_control_c1_401_c3_302(tmp_path):
"""Healthy control: C1 401 plus C3 302 → 0 issues and a healthy verdict,
proving the new leg cannot cry wolf."""
proc, record, log_path = _run_monitor(tmp_path,
az_a2a_code="401",
kagentz_public_code="302")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# Both legs report alive.
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
# Zero issues, healthy verdict.
assert "Result: ✅ 0 issues (all healthy)" in log
# No INCIDENT.
assert "INCIDENT" not in log
# No notify fired for kagentz.
assert "kagentz public URL" not in proc.stdout
assert "kagentz A2A server" not in proc.stdout
def test_c1_000_is_incident(tmp_path):
"""C1 returns 000 → INCIDENT, keeping the leg that actually caught this
outage covered behaviourally."""
proc, record, log_path = _run_monitor(tmp_path,
az_a2a_code="000",
az_a2a_exit=7,
kagentz_public_code="302")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# The C1 line reports the A2A down.
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
# The verdict is an INCIDENT (even though C3 is healthy).
assert "Result: 🔴 INCIDENT" in log
assert "all healthy" not in log
# The notify fired for the A2A down.
assert "kagentz A2A server DOWN" in proc.stdout
+43 -73
View File
@@ -1,9 +1,9 @@
---
kind: responsibility
name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.4.0
version: 3.3.0
runtime_contract: 2
agent: abiba
report_only_agents:
@@ -28,9 +28,9 @@ session start.
## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to minipve (192.168.68.12) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200`
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
- **Relay access** via RA-H OS MCP for alert delivery
@@ -126,20 +126,13 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
- **On critical alert**: Escalate to relay message immediately, don't wait for schedule
## Execution
**Execution model**: This contract is executed by a host-scheduled cron job (see
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
### Liveness rule (scoped)
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN.
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above).
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
@@ -228,20 +221,20 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| Crash loop >10/h | Alert user |
### Step 3: Platform B — Tanko (DSH on minipve CT 112)
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
Mumuni is out of scope for this host (see the note above): she runs on her own
container and is monitored on her side.
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **minipve**
PVE host (**192.168.68.12**). Direct SSH to 192.168.68.122 is not a dependency
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **amdpve**
PVE host (**192.168.68.15**). Direct SSH to 192.168.68.122 is not a dependency
of this contract — per-worker key availability varies — so CT 112 probes run
from the minipve vantage via `pct exec`:
from the amdpve vantage via `pct exec`:
```bash
ssh root@192.168.68.12 "pct exec 112 -- <command>"
ssh root@192.168.68.15 "pct exec 112 -- <command>"
```
> **By design (verified 2026-09-08):** the `dsh-web` gateway binds
@@ -253,7 +246,7 @@ ssh root@192.168.68.12 "pct exec 112 -- <command>"
**B1: Gateway Service State (Tanko)**
```bash
ssh root@192.168.68.12 "pct exec 112 -- systemctl is-active dsh-web"
ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web"
```
Expected: `active`. Anything else → gateway service down → apply the Tanko heal
@@ -262,7 +255,7 @@ Expected: `active`. Anything else → gateway service down → apply the Tanko h
**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)**
```bash
ssh root@192.168.68.12 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
```
Alive = **ANY** HTTP status response from the endpoint — the expected set is
@@ -273,13 +266,13 @@ process answering `503` is running and self-heal must NOT restart-loop it.
Down = connection refused (`000`) or timeout only. Statuses outside the
expected set are logged/reported as a warning — reported, never healed on.
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to minipve)**
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)**
```bash
curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/
```
Fallback only — used when the monitoring node has no pct/SSH path to minipve.
Fallback only — used when the monitoring node has no pct/SSH path to amdpve.
Alive = **ANY** HTTP status response from the endpoint — healthy signals are
`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other
status, including `404`/`5xx`, also counts alive: the endpoint is up and
@@ -404,29 +397,29 @@ ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
4. Every later request through `/` presents that cookie; the token is not needed
again until the cookie expires or a new browser is used.
**Verification** (minipve vantage):
**Verification** (amdpve vantage):
```bash
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
# Expected: 302
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
ssh root@192.168.68.12 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
-w '%{http_code}\n' http://192.168.68.122:8081/"
# Expected: 000
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- cat /etc/dsh-web/launch-token")
ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the minted dsh-auth-... cookie (authority
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
# 4. Token refresh is non-disruptive and idempotent.
ssh root@192.168.68.12 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
# Expected: "token unchanged; nginx not reloaded" when nothing changed
```
@@ -437,32 +430,32 @@ fresh cookie. Both verified live 2026-09-11.
```bash
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
ssh root@192.168.68.12 "pct exec 112 -- systemctl restart dsh-web"
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
# until the socket answers (any status but 000) before asserting the cookie.
for i in $(seq 1 60); do
UP=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
[ "$UP" != "000" ] && break
sleep 2
done
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the pre-restart cookie is still accepted.
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
# manual run may no-op on the flock, so poll until the include carries a token
# the running process accepts (bounded wait) before the mint+reuse check.
for i in $(seq 1 60); do
TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
CODE=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
[ "$CODE" = "303" ] && break
sleep 2
done
# Expected: 303 — the include now holds the token the running process accepts.
ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the refreshed token minted a fresh cookie.
```
@@ -476,10 +469,9 @@ ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
> monitor issued a restart for something that could not start, posting a false
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
> liveness/response and public-path access only, and a probe must never restart
> a platform.
> liveness only, and a probe must never restart a platform.
**C1: A2A Server Health (no credential needed)**
**C1: A2A Server Health**
```bash
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
@@ -487,9 +479,9 @@ ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
```
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down (INCIDENT). Any other status → running but unexpected: log/report it, never restart.
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
**C2: A2A Response Verification (requires LITELLM_KEY)**
**C2: A2A Response Verification**
```bash
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
@@ -499,30 +491,15 @@ ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
```
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set (this is a credential issue, NOT a server-down incident).
**C3: Public Access Path (no credential needed)**
```bash
# Probes the public URL that NetBird proxies to the agent-zero container.
# This is the captain's point of view: if the captain can't reach it, it's down.
# 200/302/401 = alive, 502 = proxy's "upstream refused" page (INCIDENT),
# connection failed (000) = INCIDENT. Never restarts anything.
curl -s -o /dev/null --connect-timeout 10 --max-time 15 -w '%{http_code}' https://kagentz.sysloggh.net/
```
Expected: `200` (Agent Zero login page), `302` (redirect), or `401` (auth-gated) = alive. `502` = NetBird proxy's "upstream refused" page (the container's port 80 is not listening) = INCIDENT. `000` (connection failed) = INCIDENT. Any other status → running but unexpected: log/report it, never restart.
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
**Platform C Actions**
| Condition | Action |
|-----------|--------|
| C1 A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
| C1 A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
| C2 A2A returns 401 | Check `LITELLM_KEY` is set (credential issue, NOT server-down) |
| C3 public URL returns `502` (upstream refused) | Alert only — the container's port 80 is not listening; investigate the agent-zero container |
| C3 public URL returns `000` (connection failed) | Alert only — the public path is down; investigate the NetBird proxy or the container |
| C3 public URL returns a status other than `200`/`302`/`401` | Log/report as a warning — reported, never healed on |
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
### Step 5: Global Checks
@@ -536,19 +513,12 @@ If any bot processes >50 bot-originated messages in 15min → warning.
### Step 6: Compile and Report
1. Run `scripts/zulip-monitor.sh` and take its final `Result:` line as the
authoritative run verdict. The verdict line is either
`Result: ✅ 0 issues (all healthy)` or
`Result: 🔴 INCIDENT — N issue(s) found`.
2. Quote that `Result:` line verbatim in the status report. When it says
`INCIDENT`, the run MUST be reported as an incident — never summarised as
OK/healthy and never annotated as "expected".
3. Compile all platform checks and severity
4. Determine `overall_severity` from worst per-agent severity
5. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
6. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
7. If any agent critical or >2 degraded: send relay message to user
8. Update `last_check` timestamp in `### Maintains` snapshot
1. Compile all platform checks and severity
2. Determine `overall_severity` from worst per-agent severity
3. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
4. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
5. If any agent critical or >2 degraded: send relay message to user
6. Update `last_check` timestamp in `### Maintains` snapshot
### Restart Debounce