Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3b74ca28d1 | ||
|
|
af9397d672 | ||
|
|
97dc2d772f | ||
|
|
2238777a2f | ||
|
|
6b2ba1bba5 | ||
|
|
a48b947242 | ||
|
|
ba9d29b4b9 | ||
|
|
d6376e5142 | ||
|
|
2ea6b4fc17 | ||
|
|
c6fd8eece3 | ||
|
|
4c715526ef | ||
|
|
308265e7ce | ||
|
|
88e9243ce4 | ||
|
|
13ac189365 | ||
|
|
2851a0cfc8 | ||
|
|
e0c9852de8 | ||
|
|
bf3a1ba523 | ||
|
|
f230812e3a | ||
|
|
96769a103f | ||
|
|
d697baa7b6 | ||
|
|
fb7f351a2b | ||
|
|
e42b970dec | ||
|
|
eadb927ec1 | ||
|
|
d4e238047d | ||
|
|
73d5097555 | ||
|
|
c460ef905c | ||
|
|
d99b552448 | ||
|
|
fc0cd7a032 | ||
|
|
ba38efcd75 | ||
|
|
0d30091f62 | ||
|
|
0a41a2d584 | ||
|
|
de32f54337 | ||
|
|
042fd3ccc6 | ||
|
|
226f2ad55e | ||
|
|
552c776c0c | ||
|
|
778424acd4 | ||
|
|
c5a6dcd42a | ||
|
|
9faffe4f6b | ||
|
|
cd00cd0475 | ||
|
|
819cd53bce | ||
|
|
574cb99d76 | ||
|
|
9d64b0bd66 | ||
|
|
cdc7ad2c79 | ||
|
|
5409dfd73a | ||
|
|
8b2eba4f7a | ||
|
|
040fecef3e | ||
|
|
483a66b7b4 | ||
|
|
c0d04a2c02 | ||
|
|
fda6c844ff | ||
|
|
748ea389be | ||
|
|
f16a890d0e | ||
|
|
c666d3e15c | ||
|
|
c07aa5e382 | ||
|
|
87205e6fdb | ||
|
|
8f1e5eebc4 | ||
|
|
bb1b65340e | ||
|
|
9e87927444 | ||
|
|
b0683e9566 | ||
|
|
dd11c8f14f | ||
|
|
0732eed329 | ||
|
|
59ed7cdbf7 | ||
|
|
5d70bbf25b | ||
|
|
ccc916d1ec | ||
|
|
48fb263d4b | ||
|
|
64790ebd19 | ||
|
|
6fa0a255df | ||
|
|
c72436b406 | ||
|
|
a17379676d | ||
|
|
63990b84f7 | ||
|
|
7a5ddb46a9 | ||
|
|
568fec2efa | ||
|
|
641a52c6da | ||
|
|
58033f39c5 | ||
|
|
6c59988e7f | ||
|
|
9a2ee6faec | ||
|
|
e71ded3c8c | ||
|
|
a12abbeb14 | ||
|
|
0b9aebca37 | ||
|
|
fa458afa26 | ||
|
|
aa3da83af5 | ||
|
|
aee2de25ac | ||
|
|
76653381ec | ||
|
|
3b32cc9658 | ||
|
|
d07c4494b5 | ||
|
|
66ad5ac89d | ||
|
|
b10fd6fc98 | ||
|
|
8ff13d38f3 | ||
|
|
a13457bcd6 | ||
|
|
f59d1a2159 | ||
|
|
da8f5f43c9 | ||
|
|
8ae4b59150 | ||
|
|
ef7f90ef5a | ||
|
|
43e891e679 | ||
|
|
933cfd223b | ||
|
|
f77d6ca1d1 | ||
|
|
f4f8a4cab8 | ||
|
|
7e257ce512 | ||
|
|
c0454811bb | ||
|
|
3c7f5d7d65 | ||
|
|
03be9b13d0 | ||
|
|
c295322c85 | ||
|
|
385f7e0623 | ||
|
|
7efbfffe44 | ||
|
|
315fcbae23 | ||
|
|
93f15709d1 | ||
|
|
1137dd4582 | ||
|
|
7400dfd833 | ||
|
|
c65f5219e1 | ||
|
|
077972fa2b | ||
|
|
d2bca5405a | ||
|
|
8a5cba8515 | ||
|
|
20f882412f | ||
|
|
30b2fe3fdc | ||
|
|
5112c566c8 | ||
|
|
83307eb9b2 | ||
|
|
8245716286 | ||
|
|
cfb6c03572 | ||
|
|
85f70f65bc | ||
|
|
a820b3f7dd | ||
|
|
0b92ab17b1 | ||
|
|
9edefe036e | ||
|
|
dd6e1e8b22 | ||
|
|
c712d4faf0 | ||
|
|
7f62f19c24 | ||
|
|
57bfe7e06a | ||
|
|
9100ea3326 | ||
|
|
0f26119859 | ||
|
|
5c1c8d7c19 | ||
|
|
8a2ea2d0d7 | ||
|
|
dae8d14880 | ||
|
|
39209c7ac9 | ||
|
|
8ed3b9c606 | ||
|
|
6c616a9e58 | ||
|
|
b9b1712ac6 | ||
|
|
6fb411613e | ||
|
|
4bdd88613b | ||
|
|
bc7a55122f | ||
|
|
713b9ce80c | ||
|
|
b451d6f81a | ||
|
|
fe6eb35291 | ||
|
|
40cf057370 | ||
|
|
5053a33e2c | ||
|
|
b4b5321011 | ||
|
|
69940bc9eb | ||
|
|
65eaffe1c6 | ||
|
|
b386bd0c19 | ||
|
|
8a4dd08b05 | ||
|
|
88b6decb31 | ||
|
|
7bf9f78fc6 | ||
|
|
dd9e68329e | ||
|
|
cd9ec6a0df | ||
|
|
0c7d7be5ad | ||
|
|
9312c1a806 |
@@ -1,19 +1,18 @@
|
||||
name: PR Pipeline — Authorize → Validate → Review → Merge
|
||||
# TRIGGER IS INTENTIONALLY UNFILTERED — DO NOT RE-ADD A `paths:` FILTER.
|
||||
#
|
||||
# This workflow previously carried `paths: ['**.prose.md', 'scripts/**.sh',
|
||||
# '**.yaml', '**.yml']` on both `push` and `pull_request`. Any PR whose diff
|
||||
# touched none of those patterns (for example a `deliverables/`-only PR, or a
|
||||
# `scripts/*.py` / `bin/*` change) therefore produced NO Gitea Actions run at
|
||||
# all: validation, lint, ai-review and the merge gate were silently skipped.
|
||||
# Validation must run for every pull request and every push to master, so the
|
||||
# trigger is deliberately unconditional.
|
||||
on:
|
||||
push:
|
||||
branches: [master]
|
||||
paths:
|
||||
- '**.prose.md'
|
||||
- 'scripts/**.sh'
|
||||
- '**.yaml'
|
||||
- '**.yml'
|
||||
pull_request:
|
||||
types: [opened, synchronize, reopened]
|
||||
paths:
|
||||
- '**.prose.md'
|
||||
- 'scripts/**.sh'
|
||||
- '**.yaml'
|
||||
- '**.yml'
|
||||
|
||||
jobs:
|
||||
auth:
|
||||
@@ -72,6 +71,18 @@ jobs:
|
||||
git fetch origin "${{ gitea.ref }}" --depth=50
|
||||
git checkout "${{ gitea.sha }}"
|
||||
|
||||
- name: Committed-credential scan (secret guard)
|
||||
run: |
|
||||
# Fails the build on a credential-shaped string in the tree. Patterns
|
||||
# live in scripts/secret-patterns.tsv; the only tolerated literal
|
||||
# examples are in scripts/secret-allowlist.tsv, each with a reason.
|
||||
# Do not turn this into a warning: a warning in a stream nobody reads
|
||||
# is how six live credentials sat in this repo for weeks.
|
||||
bash scripts/secret-scan.sh
|
||||
|
||||
- name: Secret guard self-test
|
||||
run: bash tests/test_secret_scan.sh
|
||||
|
||||
- name: Structure + regression + consistency lint
|
||||
run: bash scripts/prose-lint.sh
|
||||
|
||||
|
||||
@@ -1 +1,2 @@
|
||||
__pycache__/
|
||||
state/host-disk-bands.json
|
||||
|
||||
@@ -51,6 +51,12 @@ Two incidents taught us this:
|
||||
- `/grafana/` nginx route — was reverted Jul 2, must not reappear
|
||||
- `CT 122` or `CT 123` as CT ID labels — don't exist in the cluster
|
||||
- These rules are hardcoded in `scripts/prose-lint.sh`
|
||||
- **Committed-credential guard:** `scripts/secret-scan.sh` FAILS the build on
|
||||
credential-shaped strings (patterns in `scripts/secret-patterns.tsv`, prose
|
||||
included). Tolerated literals are listed one-per-example with a reason in
|
||||
`scripts/secret-allowlist.tsv`; never allowlist a live credential. It runs in
|
||||
the CI lint job, in `scripts/prose-lint.sh`, and via
|
||||
`bash scripts/secret-scan.sh --staged` before committing.
|
||||
|
||||
### Stage 3 — AI Review
|
||||
- Diff is sent to `syslog-auto` model via LiteLLM
|
||||
|
||||
@@ -0,0 +1,142 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: agent-health-check
|
||||
description: >
|
||||
Consolidated agent health verification for the LiteLLM + GPU + Zulip +
|
||||
gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours
|
||||
via cron and on-demand via "run contract: agent-health-check". Verifies:
|
||||
LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway
|
||||
liveness, CT liveness, gateway log health, config YAML integrity,
|
||||
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
|
||||
title: Agent Health Check — Consolidated
|
||||
version: 1.0.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
---
|
||||
|
||||
# Agent Health Check
|
||||
|
||||
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
|
||||
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only.
|
||||
|
||||
**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
|
||||
|
||||
## Requires
|
||||
|
||||
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
|
||||
- **SSH access** to GPU hosts — `llmuser` on .8 (owns `llama-server`), `root` on .110 and .15 — and agent CTs (.122, .129, .114, .24)
|
||||
- **Python 3** for script execution
|
||||
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
|
||||
|
||||
## Maintains
|
||||
|
||||
- last_check: timestamp — When the last full diagnostic ran
|
||||
- overall_severity: "healthy" | "degraded" | "critical"
|
||||
- liteLLM_keys: map of agent → key validity
|
||||
- gpu_ports: map of host → port conflict status
|
||||
- agents: map of agent → streaming health + gateway liveness
|
||||
- ct_liveness: map of CT → active status
|
||||
- config_integrity: map of config file → valid/invalid
|
||||
|
||||
## Execution
|
||||
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||
execution — it only proves the agent read the result and reported it. The actual
|
||||
monitoring work happens in the host cron job.
|
||||
|
||||
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool
|
||||
calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Run the consolidated health check script
|
||||
python3 /root/scripts/agent-health-check.py --json
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the script
|
||||
executed from** so a stale-consumer report is distinguishable from a real fault
|
||||
at read time. Summarize actual results from each check. Apply the standing probe
|
||||
rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
|
||||
|
||||
**Expected output**: JSON with `overall_severity` field. If `healthy`, report
|
||||
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
||||
failures and their severity.
|
||||
|
||||
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
|
||||
MUST include one clause per check leg, in every state: healthy, degraded/warn,
|
||||
skipped, or failed. A missing leg must never look the same as a healthy leg.
|
||||
Required legs and their templates in every state:
|
||||
|
||||
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
|
||||
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
|
||||
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
|
||||
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
|
||||
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
|
||||
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
|
||||
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
|
||||
- The GPU leg has six non-healthy states the code can produce:
|
||||
(i) `gpu-unreachable:{host}` — SSH probe failed;
|
||||
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
|
||||
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
|
||||
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
|
||||
(v) unit active, /health body contains "error" — error response;
|
||||
(vi) unit active, /health body unrecognised — unknown health.
|
||||
In every case the failing host and reason must be named.
|
||||
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
|
||||
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
|
||||
- skipped: `CTs: SKIPPED (SSH access unavailable)`
|
||||
- `Vault secrets: 3/3 present`
|
||||
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
|
||||
- skipped: `Vault secrets: SKIPPED (vault not configured)`
|
||||
|
||||
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
|
||||
long as the host is identifiable from context; the full `probe-failed: <target>
|
||||
<kind>` form is required when a leg reports a failure in the detail section.
|
||||
|
||||
### Probe Shape (per standing rules from 1150.msg)
|
||||
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||
service answered — report the code, never "down". A redirect is not a failure.
|
||||
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
2. **A failed probe is never a service verdict.** Print
|
||||
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
||||
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
||||
then report.
|
||||
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
||||
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
||||
(retried at 25s: also timeout)" is actionable.
|
||||
|
||||
## Strategies
|
||||
|
||||
### When LiteLLM keys are invalid
|
||||
Report the specific agent + key name. Do not attempt to fix — credential
|
||||
rotation is a separate operation.
|
||||
|
||||
### When GPU port conflicts are detected
|
||||
Report the conflicting ports and processes. Do not kill processes — that's a
|
||||
destructive action requiring captain approval.
|
||||
|
||||
### When gateway liveness is degraded
|
||||
Report the specific CT + gateway status. Do not restart unless the restart
|
||||
debounce window has passed.
|
||||
|
||||
### When CT liveness is down
|
||||
Report the specific CT. Do not restart — that's a destructive action.
|
||||
|
||||
### When config YAML is invalid
|
||||
Report the specific file + parse error. Do not fix — that's a config change.
|
||||
|
||||
### When gateway log health is degraded
|
||||
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
|
||||
|
||||
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
|
||||
|
||||
## Continuity
|
||||
|
||||
- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
|
||||
- **On `agent-health` command**: Run on-demand and report to user
|
||||
- **On critical alert**: Escalate to relay message immediately
|
||||
@@ -16,8 +16,8 @@ litellm.exceptions.AuthenticationError: OpenrouterException -
|
||||
```
|
||||
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
||||
|
||||
**Old Key**: `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`
|
||||
**New Key**: `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`
|
||||
**Old Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
**New Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
|
||||
### 2. Telegram Bot Conflict (CRITICAL)
|
||||
@@ -48,14 +48,14 @@ McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 sec
|
||||
```bash
|
||||
# Container .env update
|
||||
sudo docker exec agent-zero bash -c '
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»|" /a0/usr/.env
|
||||
'
|
||||
```
|
||||
|
||||
**Verification**:
|
||||
```bash
|
||||
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
|
||||
-H "Authorization: Bearer «vault: agents/production OPENROUTER_API_KEY»" | python3 -m json.tool
|
||||
```
|
||||
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
||||
|
||||
@@ -140,7 +140,7 @@ Added section:
|
||||
|
||||
| Component | Status | Details |
|
||||
|-----------|--------|---------|
|
||||
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
|
||||
| **OpenRouter Key** | ✅ Valid | `«vault: agents/production OPENROUTER_API_KEY»` user verified, |
|
||||
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
||||
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
||||
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
||||
|
||||
@@ -54,7 +54,7 @@ description: >
|
||||
```
|
||||
|
||||
4. **Return status**
|
||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-synthetic...", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
||||
- If vault secret is missing: `{ vault_synced: false }`
|
||||
|
||||
@@ -89,8 +89,8 @@ description: >
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| **Key Prefix** | `sk-or-v1-0af3f3` |
|
||||
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
|
||||
| **Key Prefix** | `«vault: agents/production OPENROUTER_API_KEY»` |
|
||||
| **Full Key** | `«vault: agents/production OPENROUTER_API_KEY»` (in vault + /a0/usr/.env) |
|
||||
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
||||
| **Free Tier** | No |
|
||||
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
||||
@@ -101,7 +101,7 @@ description: >
|
||||
|
||||
| Date | Action | Notes |
|
||||
|------|--------|-------|
|
||||
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||
| 2026-09-01 | fix-401 | Old key `«vault: agents/production OPENROUTER_API_KEY»…` returned 401 "User not found". Replaced with new key `«vault: agents/production OPENROUTER_API_KEY»…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||
|
||||
## Infrastructure References
|
||||
|
||||
|
||||
+90
-23
@@ -94,7 +94,16 @@ def audit(path):
|
||||
cfg = yaml.safe_load(f)
|
||||
|
||||
model = cfg.get("model", {})
|
||||
fb = cfg.get("fallback_providers", {})
|
||||
fb_raw = cfg.get("fallback_providers", {})
|
||||
# Normalize: fallback_providers may be a dict (single provider) or a list of dicts
|
||||
# (one entry per fallback). Both shapes are valid; we must handle both without crashing.
|
||||
if isinstance(fb_raw, dict):
|
||||
fb_entries = [fb_raw]
|
||||
elif isinstance(fb_raw, list):
|
||||
fb_entries = fb_raw
|
||||
else:
|
||||
fb_entries = [fb_raw] # Let it fail the check below as malformed
|
||||
fb = fb_entries[0] if fb_entries else {}
|
||||
comp = cfg.get("compression", {})
|
||||
aux = cfg.get("auxiliary", {})
|
||||
deleg = cfg.get("delegation", {})
|
||||
@@ -114,12 +123,22 @@ def audit(path):
|
||||
)
|
||||
|
||||
# --- Rule 5: Main Config Base URL ---
|
||||
expected_base = "http://192.168.68.116/v1"
|
||||
check(
|
||||
model.get("base_url") == expected_base,
|
||||
"Rule 5",
|
||||
f"model.base_url must be {expected_base} (got {model.get('base_url')!r}) — /v1 not /litellm/v1",
|
||||
)
|
||||
# Canonical internal base (hermes-key-enforcement.prose.md:19/38/57/84/91/96)
|
||||
# and public base (serves /v1 only, per 2026-09-19 probe from CT 116).
|
||||
# The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only.
|
||||
# FAIL anything else (do not widen to accept any path ending in /v1).
|
||||
# Internal /v1 is non-canonical but working (authenticated via nginx), so WARN not FAIL.
|
||||
canonical_internal = "http://192.168.68.116/litellm/v1"
|
||||
public_host = "https://litellm.sysloggh.net/v1"
|
||||
non_canonical_internal = "http://192.168.68.116/v1"
|
||||
allowed_bases = (canonical_internal, public_host)
|
||||
actual_base = model.get("base_url")
|
||||
if actual_base in allowed_bases:
|
||||
check(True, "Rule 5", f"model.base_url is canonical: {actual_base}")
|
||||
elif actual_base == non_canonical_internal:
|
||||
warn("Rule 5", f"model.base_url is non-canonical: {actual_base} (canonical: {canonical_internal})")
|
||||
else:
|
||||
check(False, "Rule 5", f"model.base_url must be one of {allowed_bases} (got {actual_base!r})")
|
||||
|
||||
# --- Rule 6: max_tokens Is Required ---
|
||||
check(
|
||||
@@ -197,22 +216,32 @@ def audit(path):
|
||||
"Rule 14",
|
||||
f"delegation.provider must be 'harness' (got {deleg.get('provider')!r})",
|
||||
)
|
||||
check(
|
||||
fb.get("provider") == "deepseek",
|
||||
"Rule 14",
|
||||
f"fallback_providers.provider must be 'deepseek' (got {fb.get('provider')!r}) — "
|
||||
f"true fallback diversity, not same endpoint as primary",
|
||||
)
|
||||
check(
|
||||
fb.get("model") == "deepseek-v4-flash",
|
||||
"Rule 14",
|
||||
f"fallback_providers.model must be 'deepseek-v4-flash' (got {fb.get('model')!r})",
|
||||
)
|
||||
check(
|
||||
fb.get("api_key_env") == "DEEPSEEK_API_KEY",
|
||||
"Rule 14",
|
||||
f"fallback_providers.api_key_env must be DEEPSEEK_API_KEY (got {fb.get('api_key_env')!r})",
|
||||
)
|
||||
# Check each fallback entry. A malformed entry (not a mapping) is a VIOLATION, not a crash.
|
||||
for idx, entry in enumerate(fb_entries):
|
||||
prefix = f"fallback_providers[{idx}]"
|
||||
if not isinstance(entry, dict):
|
||||
check(
|
||||
False,
|
||||
"Rule 14",
|
||||
f"{prefix} must be a mapping (got {type(entry).__name__})",
|
||||
)
|
||||
continue
|
||||
check(
|
||||
entry.get("provider") == "deepseek",
|
||||
"Rule 14",
|
||||
f"{prefix}.provider must be 'deepseek' (got {entry.get('provider')!r}) — "
|
||||
f"true fallback diversity, not same endpoint as primary",
|
||||
)
|
||||
check(
|
||||
entry.get("model") == "deepseek-v4-flash",
|
||||
"Rule 14",
|
||||
f"{prefix}.model must be 'deepseek-v4-flash' (got {entry.get('model')!r})",
|
||||
)
|
||||
check(
|
||||
entry.get("api_key_env") == "DEEPSEEK_API_KEY",
|
||||
"Rule 14",
|
||||
f"{prefix}.api_key_env must be DEEPSEEK_API_KEY (got {entry.get('api_key_env')!r})",
|
||||
)
|
||||
|
||||
# --- custom_providers sanity ---
|
||||
check(
|
||||
@@ -262,6 +291,44 @@ def audit(path):
|
||||
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
||||
)
|
||||
|
||||
# --- MCP Server Checks (Rule 15) ---
|
||||
# Valid MCP server endpoints
|
||||
VALID_MCP_ENDPOINTS = {
|
||||
'ra-h-os': 'http://192.168.68.65:3100/mcp',
|
||||
'litellm': 'https://litellm.sysloggh.net/mcp',
|
||||
}
|
||||
|
||||
# Check MCP servers if they exist
|
||||
mcp_servers = cfg.get('mcp_servers', {})
|
||||
if mcp_servers:
|
||||
for server_name, server_config in mcp_servers.items():
|
||||
url = server_config.get('url', '')
|
||||
|
||||
# Check endpoint validity
|
||||
if server_name in VALID_MCP_ENDPOINTS:
|
||||
expected = VALID_MCP_ENDPOINTS[server_name]
|
||||
if url == expected:
|
||||
check(True, 'Rule 15', f'MCP server "{server_name}" URL is correct: {url}')
|
||||
else:
|
||||
check(False, 'Rule 15', f'MCP server "{server_name}" URL is incorrect: {url} (expected: {expected})')
|
||||
else:
|
||||
warn('Rule 15', f'MCP server "{server_name}" URL may need validation (not in known list): {url}')
|
||||
|
||||
# Check for proper authentication
|
||||
headers = server_config.get('headers', {})
|
||||
has_auth = False
|
||||
for key, value in headers.items():
|
||||
if 'key' in key.lower() or 'auth' in key.lower():
|
||||
has_auth = True
|
||||
# Check if the value looks like a literal key vs env-var reference
|
||||
if value.startswith('Bearer ') and value[7:].startswith('sk-'):
|
||||
check(True, 'Rule 15', f'MCP server "{server_name}" has valid auth header: {key}')
|
||||
else:
|
||||
warn('Rule 15', f'MCP server "{server_name}" header may use env-var instead of literal key: {key} = {value}')
|
||||
break
|
||||
if not has_auth:
|
||||
warn('Rule 15', f'MCP server "{server_name}" has no authentication header')
|
||||
|
||||
# --- Report ---
|
||||
print(f"{'=' * 60}")
|
||||
print(f"Hermes Config Audit: {path}")
|
||||
|
||||
@@ -0,0 +1,173 @@
|
||||
# Search ranking policy for the agent-consumption layer.
|
||||
#
|
||||
# Everything here is CONFIG, not code, so it is reviewable and changeable without
|
||||
# touching the module. Read by scripts/search-agent-consume.py.
|
||||
#
|
||||
# Why this file exists: multi-engine aggregation returns results with no
|
||||
# filtering, no dedupe and no reranking. On 2026-09-26 that put a shopping page
|
||||
# and a dictionary definition into "best practices agent context management",
|
||||
# and put four SEO blogs ABOVE the actual Proxmox forum threads on a precise
|
||||
# technical query. Identical queries also ranked differently between runs, which
|
||||
# is the strongest argument for a deterministic layer rather than hoping the
|
||||
# engines behave.
|
||||
|
||||
version: 1
|
||||
|
||||
# ── Non-answers: dropped outright, never returned ────────────────────────────
|
||||
# These are pages that cannot answer a question: navigational homepages,
|
||||
# shopping/product pages, dictionary definitions, and login walls.
|
||||
non_answer:
|
||||
# URL path is empty -> it is a site's front door, not an answer. Still allowed
|
||||
# when the host is explicitly preferred (see prefer_domains), because some
|
||||
# docs/repo front doors ARE the answer.
|
||||
host_root: true
|
||||
path_patterns:
|
||||
- '/dictionary/'
|
||||
- '/dictionary?'
|
||||
- '/wiki/Wiktionary:'
|
||||
- '/search?'
|
||||
- '/cart'
|
||||
- '/checkout'
|
||||
- '/login'
|
||||
- '/signin'
|
||||
- '/sign-in'
|
||||
- '/account/login'
|
||||
- '/shop/'
|
||||
- '/store/'
|
||||
- '/dp/' # Amazon-style product URL
|
||||
- '/gp/product/'
|
||||
- '/add-to-cart'
|
||||
- '/checkout'
|
||||
# NOTE: '/products/' and '/product/' were REMOVED as path patterns. They fired
|
||||
# on docs.digitalocean.com/products/inference/... — a legitimate documentation
|
||||
# page — which the 2026-09-26 before/after run caught. Shopping is caught by
|
||||
# the shopping HOST list instead, which does not have that false positive.
|
||||
# Query strings that betray a search/shopping surface rather than an article.
|
||||
query_keys:
|
||||
- 'q'
|
||||
- 'query'
|
||||
- 's'
|
||||
- 'search'
|
||||
- 'add-to-cart'
|
||||
# Hosts that are shopping/retail and never answer a technical question.
|
||||
hosts:
|
||||
- bestbuy.com
|
||||
- amazon.com
|
||||
- ebay.com
|
||||
- walmart.com
|
||||
- etsy.com
|
||||
- aliexpress.com
|
||||
- merriam-webster.com
|
||||
- dictionary.com
|
||||
- thesaurus.com
|
||||
- vocabulary.com
|
||||
- collinsdictionary.com
|
||||
|
||||
# ── Demotion: ranked below everything else, never dropped ────────────────────
|
||||
# Low-authority content farms / SEO aggregators. Demoted rather than dropped so
|
||||
# a genuinely useful hit is not lost, but it can never outrank a primary source.
|
||||
# Reviewable: add or remove hosts here, no code change required.
|
||||
demote_domains:
|
||||
- medium.com
|
||||
- sparkco.ai
|
||||
- mindstudio.ai
|
||||
- aitechmonk.com
|
||||
- stackai.com
|
||||
- agentic-design.ai
|
||||
- voxfor.com
|
||||
- bigiron.cc
|
||||
- linuxoperatingsystem.net
|
||||
- riparazioneserver.com
|
||||
- rossmanngroup.com
|
||||
- dev.to
|
||||
- hashnode.dev
|
||||
- substack.com
|
||||
- towardsdatascience.com
|
||||
- analyticsvidhya.com
|
||||
- geeksforgeeks.org
|
||||
- tutorialspoint.com
|
||||
- javatpoint.com
|
||||
- w3schools.com
|
||||
- scaler.com
|
||||
- simplilearn.com
|
||||
- udemy.com
|
||||
- coursera.org
|
||||
|
||||
# ── Preference: promoted above the default rank ──────────────────────────────
|
||||
# Primary sources: upstream repositories, official docs, Q&A, vendor
|
||||
# engineering blogs. These are what an agent should be reading.
|
||||
prefer_domains:
|
||||
# upstream repositories and code hosting
|
||||
- github.com
|
||||
- gitlab.com
|
||||
- codeberg.org
|
||||
- sourceforge.net
|
||||
- kernel.org
|
||||
- git.kernel.org
|
||||
# Q&A
|
||||
- stackoverflow.com
|
||||
- stackexchange.com
|
||||
- superuser.com
|
||||
- serverfault.com
|
||||
- askubuntu.com
|
||||
- discourse.org
|
||||
# vendor / project documentation and forums
|
||||
- proxmox.com
|
||||
- forum.proxmox.com
|
||||
- pve.proxmox.com
|
||||
- docs.python.org
|
||||
- developer.mozilla.org
|
||||
- kernelnewbies.org
|
||||
- man7.org
|
||||
- gnu.org
|
||||
- debian.org
|
||||
- ubuntu.com
|
||||
- redhat.com
|
||||
- kernel.dk # io_uring / Jens Axboe
|
||||
- github.io # project pages (docs, papers) — promoted, not authoritative by itself
|
||||
# vendor engineering blogs
|
||||
- anthropic.com
|
||||
- openai.com
|
||||
- googleblog.com
|
||||
- developers.googleblog.com
|
||||
- engineering.fb.com
|
||||
- netflixtechblog.com
|
||||
- aws.amazon.com
|
||||
- cloud.google.com
|
||||
- microsoft.com
|
||||
- learn.microsoft.com
|
||||
- apple.com
|
||||
- nvidia.com
|
||||
- intel.com
|
||||
- amd.com
|
||||
- redislabs.com
|
||||
- cloudflare.com
|
||||
- langchain.com
|
||||
- jetbrains.com
|
||||
- cursor.com
|
||||
# community discussion with high signal
|
||||
- news.ycombinator.com
|
||||
- lobste.rs
|
||||
- reddit.com
|
||||
|
||||
# ── Ranking weights ──────────────────────────────────────────────────────────
|
||||
# Final score = engine_score - demote_penalty + prefer_bonus, then a stable
|
||||
# tiebreak on original position so ordering is reproducible run to run.
|
||||
ranking:
|
||||
demote_penalty: 1000
|
||||
prefer_bonus: 100
|
||||
# Results that several engines independently returned are more likely real.
|
||||
multi_engine_bonus: 25
|
||||
# Shallow paths (e.g. /blog/x) are slightly less likely to be primary docs.
|
||||
host_root_allowed_when_preferred: true
|
||||
|
||||
# ── Extraction budget (criterion 4) ──────────────────────────────────────────
|
||||
# Return CONTENT, not just links, so an agent gets usable material in ONE call.
|
||||
extraction:
|
||||
top_n: 5 # how many results get page text extracted
|
||||
total_chars: 12000 # global budget across all extracted items
|
||||
per_item_chars: 4000 # cap for any single item, so one page cannot eat the budget
|
||||
timeout_seconds: 45 # per scrape
|
||||
# If extraction fails, the result is still returned with an empty excerpt —
|
||||
# a link is better than nothing, but the failure is recorded in the output.
|
||||
on_failure: keep_with_empty_excerpt
|
||||
+118
-2
@@ -38,6 +38,7 @@ index:
|
||||
by_category:
|
||||
compliance:
|
||||
- hermes-key-enforcement
|
||||
- litellm-api-keys
|
||||
- hermes-config-template
|
||||
- hermes-agent-baseline
|
||||
monitoring:
|
||||
@@ -46,6 +47,7 @@ index:
|
||||
- infrastructure-monitoring
|
||||
- zulip-health
|
||||
- litellm-health
|
||||
- daily-health-digest
|
||||
remediation:
|
||||
- litellm-self-heal
|
||||
- pm2-self-heal
|
||||
@@ -95,12 +97,14 @@ index:
|
||||
- infrastructure-maintenance
|
||||
- pm2-self-heal
|
||||
- disk-gc-threat-response
|
||||
- daily-health-digest
|
||||
gpu:
|
||||
- gpu-monitor
|
||||
- gpu-fleet
|
||||
proxmox:
|
||||
- proxmox-monitor
|
||||
litellm:
|
||||
- litellm-api-keys
|
||||
- litellm-health
|
||||
- litellm-self-heal
|
||||
memory:
|
||||
@@ -628,7 +632,7 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 3.3.0
|
||||
version: 3.4.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/15 * * * *'
|
||||
@@ -639,7 +643,7 @@ contracts:
|
||||
timeout: 120
|
||||
requires:
|
||||
- Zulip API key for abiba-bot@chat.sysloggh.net
|
||||
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
|
||||
- SSH access to minipve (192.168.68.12) for Tanko (CT 112) and the Agent Zero Docker host (.14)
|
||||
verification:
|
||||
postconditions:
|
||||
- check: bot registration active
|
||||
@@ -1869,6 +1873,118 @@ contracts:
|
||||
drift_alerts: []
|
||||
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||
- name: litellm-api-keys
|
||||
file: litellm-api-keys.prose.md
|
||||
kind: function
|
||||
category: compliance
|
||||
sensitivity: critical
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: on_demand
|
||||
cadence: null
|
||||
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 120
|
||||
requires: []
|
||||
protocol:
|
||||
- Load contract from prose-contracts/main
|
||||
- Retrieve master key from Infisical (project=infrastructure env=production)
|
||||
- Read live key-scoped model roster from CT 116 /v1/models
|
||||
- Create/rotate/verify the requested agent key with an EXPLICIT models list
|
||||
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
|
||||
verification:
|
||||
postconditions:
|
||||
- check: standard agent key is local-only
|
||||
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
|
||||
expect: 0 cloud models
|
||||
- check: key exists with correct alias
|
||||
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
|
||||
expect: 200 with matching alias
|
||||
artifact: key creation/rotation report
|
||||
receipt:
|
||||
format: json
|
||||
storage: ~/.hermes/runs/litellm-api-keys/
|
||||
graph_node: true
|
||||
escalation:
|
||||
info:
|
||||
action: log_to_receipt
|
||||
notify: []
|
||||
warning:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
critical:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
- ops
|
||||
- name: daily-health-digest
|
||||
file: daily-health-digest.prose.md
|
||||
kind: function
|
||||
category: monitoring
|
||||
sensitivity: normal
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 1.0.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: 30 10 * * *
|
||||
description: Daily at 10:30 UTC, dispatched on CT 100 as a firstmate message
|
||||
to the ops lane, which executes the pinned producer
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 300
|
||||
requires:
|
||||
- infisical (vault credentials injected at run time)
|
||||
protocol:
|
||||
- 'Execute from the PINNED clone only: /root/abiba-workspace/projects/prose-contracts'
|
||||
- cd /root/abiba-workspace/projects/prose-contracts
|
||||
- infisical run --env=prod -- python3 scripts/daily-infra-report.py
|
||||
- 'Never execute from a per-agent working copy (treehouse) - it drifts onto feature branches'
|
||||
- 'On failure: do not treat a 0/0 Proxmox section as evidence about the estate - it means could not look'
|
||||
verification:
|
||||
postconditions:
|
||||
- check: every Proxmox probe is reachable
|
||||
verify: >-
|
||||
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
|
||||
| grep -c 'pve_probe_status'
|
||||
expect: '1'
|
||||
- check: all nodes reported online
|
||||
verify: >-
|
||||
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
|
||||
| grep 'nodes_online'
|
||||
expect: nodes_online == node_count
|
||||
artifact: timestamped HTML dashboard emailed to jerome@sysloggh.com
|
||||
verify_commands:
|
||||
- infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
|
||||
- python3 -m pytest tests/test_daily_infra_report.py -q
|
||||
delivery:
|
||||
transport: zulip-dm-attachment
|
||||
recipient_user_id: 9
|
||||
sender: abiba-bot@chat.sysloggh.net
|
||||
key_source: abiba-bot Zulip key already on the execution host, read from the
|
||||
600-mode env file /root/.pi/agent/extensions/zulip/.env
|
||||
key_policy: do NOT add a vault entry - that is a captain decision under the auth-keys charter
|
||||
body: short Markdown pointer; the HTML attachment IS the report
|
||||
artifact: /var/log/daily-infra-report/infra-report-<UTCstamp>.html
|
||||
note: Replaced SMTP/mail on 2026-09-26 by captain decision. Removes the Google
|
||||
dependency entirely; closes daily-digest-mail-transport-20260921.
|
||||
exit_semantics:
|
||||
'1': missing PVE_TOKEN, unreachable Proxmox probe, missing/rejected Zulip
|
||||
credential, or a failed upload/post - raises an alert
|
||||
'0': healthy delivery only - there is no degraded delivery leg any more
|
||||
depends_on: []
|
||||
last_run: null
|
||||
last_status: null
|
||||
drift_alerts: []
|
||||
|
||||
koby_report_only: true
|
||||
koby_host: "CT 111 (tdunna)"
|
||||
koby_ip: ".129"
|
||||
|
||||
@@ -0,0 +1,203 @@
|
||||
---
|
||||
kind: function
|
||||
name: daily-health-digest
|
||||
description: >
|
||||
Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox
|
||||
nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails
|
||||
it as an HTML report.
|
||||
|
||||
Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to
|
||||
the ops lane, which executes the producer below. Until 2026-09-25 this ran
|
||||
with NO contract file at all, which is why the choice of execution copy was
|
||||
silently the operator's rather than the contract's.
|
||||
|
||||
EXECUTION IS PINNED. The producer must be run from the clone named under
|
||||
"Execution pinning" — not from an agent working copy.
|
||||
|
||||
Exit-code semantics (as they actually behave, verified 2026-09-25):
|
||||
* missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert
|
||||
* missing or rejected Zulip credential -> exit 1 (delivery is the only
|
||||
output path, so it is a real failure, not a degraded leg)
|
||||
* delivery failure -> exit 1, and the report body is printed AND persisted
|
||||
so the content is never swallowed
|
||||
|
||||
version: 2.0.0
|
||||
---
|
||||
|
||||
## Purpose
|
||||
|
||||
Give one daily, machine-collected picture of the estate so drift and outages
|
||||
are seen the day they happen rather than when something breaks. It is a
|
||||
*report*, not a repair: it changes nothing.
|
||||
|
||||
## Execution pinning
|
||||
|
||||
**Pinned execution path:**
|
||||
|
||||
```
|
||||
/root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py
|
||||
```
|
||||
|
||||
**Pinned clone:** `/root/abiba-workspace/projects/prose-contracts`
|
||||
|
||||
That is the cron's `FM_HOME` clone and the only stable, non-ephemeral copy.
|
||||
The treehouse clone (`/root/.treehouse/agent-workspace-*/…/projects/prose-contracts`)
|
||||
is a **per-agent working copy and must NOT be pinned or executed from** — it
|
||||
drifts onto feature branches, which is exactly how a stale producer reported a
|
||||
stale picture and nobody noticed.
|
||||
|
||||
See `docs/contract-execution-pinning.md`. Schedule and alerting live in
|
||||
`/etc/cron.d/contract-runner` on CT 100:
|
||||
|
||||
```
|
||||
30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1
|
||||
```
|
||||
|
||||
Invocation (credentials come from the vault; never inline them):
|
||||
|
||||
```bash
|
||||
cd /root/abiba-workspace/projects/prose-contracts
|
||||
infisical run --env=prod -- python3 scripts/daily-infra-report.py
|
||||
```
|
||||
|
||||
## Output shape
|
||||
|
||||
Modes:
|
||||
|
||||
| invocation | effect |
|
||||
| --- | --- |
|
||||
| *(none)* | collect, build the HTML dashboard, email it |
|
||||
| `--test-email` | same but with a `🧪 TEST —` subject prefix |
|
||||
| `--json` | print the collected data as JSON to stdout and **send no email** |
|
||||
|
||||
`--json` emits a single object with these top-level keys (observed on a live
|
||||
run 2026-09-25):
|
||||
|
||||
| key | type | meaning |
|
||||
| --- | --- | --- |
|
||||
| `nodes` | object (5) | per-node cpu/ram/disk/uptime/status |
|
||||
| `node_count`, `nodes_online` | int | Proxmox node totals |
|
||||
| `pve_probe_status`, `resources_probe_status` | `ok`\|`unreachable` | probe outcome |
|
||||
| `total_vms`, `running_vms`, `stopped_vms`, `vms_by_node` | — | guest inventory |
|
||||
| `storage`, `nfs` | array | datastore and mount usage |
|
||||
| `litellm` | object | inference checks |
|
||||
| `zulip_ext` | object | Zulip queue/serving state |
|
||||
| `agents` | object | per-agent health |
|
||||
| `docker_vm`, `docker_syslog`, `docker_netbird`, `endpoints` | object/array | Docker hosts and probed endpoints |
|
||||
|
||||
## What a healthy run looks like
|
||||
|
||||
```
|
||||
$ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
|
||||
"nodes_online": 5, "node_count": 5, "pve_probe_status": "ok",
|
||||
"resources_probe_status": "ok", "running_vms": 22, "total_vms": 22
|
||||
EXIT=0
|
||||
```
|
||||
|
||||
and in delivery mode:
|
||||
|
||||
```
|
||||
report ready: 16208 chars of HTML (delivered as a file attachment)
|
||||
Sending to the captain's Zulip DM...
|
||||
✅ Delivered to Zulip DM (user 9), message id 86221, attachment 16208 bytes
|
||||
at /user_uploads/2/45/m1cQesBFV78BGeNY2lN8xkN5/infra-report-20260926-153406.html
|
||||
```
|
||||
|
||||
Healthy means: every probe reports `ok`, `nodes_online == node_count`, and the
|
||||
email leg reports a successful send.
|
||||
|
||||
## Exit-code semantics — as they actually behave
|
||||
|
||||
Verified on 2026-09-25 by running each case deliberately.
|
||||
|
||||
| condition | exit | alert | notes |
|
||||
| --- | --- | --- | --- |
|
||||
| all probes reachable, email sent | 0 | — | healthy |
|
||||
| **missing `PVE_TOKEN`** | **1** | yes | `PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down)`, and `cluster resources unreachable` |
|
||||
| **Proxmox probe unreachable** | **1** | yes | same path as above; `pve_probe_status: unreachable` |
|
||||
| **missing/rejected Zulip credential** | **1** | yes | delivery is the only output path; report printed and persisted |
|
||||
| **upload or message post fails** | **1** | yes | report printed and persisted; message names which step failed |
|
||||
|
||||
The distinction is deliberate and must not be flattened:
|
||||
|
||||
* A **missing PVE token or an unreachable probe is a real failure** — the report
|
||||
would otherwise claim zero nodes and still look successful. That was the
|
||||
2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1.
|
||||
* A **missing email credential is survivable** — the report is still produced
|
||||
and is still useful. It is a `DEGRADED` leg and exits 0 by design.
|
||||
|
||||
`PROBE_FAILURES` and `DEGRADED_LEGS` are separate lists for exactly this
|
||||
reason. Do not merge them.
|
||||
|
||||
## Delivery: Zulip DM carrying the report as an HTML ATTACHMENT
|
||||
|
||||
Captain's decision 2026-09-26, clarified the same day: the digest is delivered to
|
||||
his **Zulip DM (user id 9)** from `abiba-bot@chat.sysloggh.net`, as an **HTML
|
||||
FILE** — an attachment, not HTML rendered in the message body and not a Markdown
|
||||
translation of it.
|
||||
|
||||
* the styled dashboard is built exactly as before and written to
|
||||
`/var/log/daily-infra-report/infra-report-<UTCstamp>.html`;
|
||||
* it is uploaded through `POST /api/v1/user_uploads`;
|
||||
* the **message body stays short Markdown** — subject line, top-line status
|
||||
(nodes online, guests running, any degraded legs), and a link to the
|
||||
attachment. The attachment IS the report; the body does not reproduce it.
|
||||
|
||||
This removes the Google dependency entirely: **no SMTP, no `EMAIL_PASSWORD`, no
|
||||
app password, nothing to rotate.** `daily-digest-mail-transport-20260921` is
|
||||
closed under this option.
|
||||
|
||||
The **10,000-character message cap does not apply** — it bounds message TEXT
|
||||
only, and the report travels as a file. Do not shrink the report to fit it.
|
||||
|
||||
The credential is abiba-bot's Zulip key already on the execution host at
|
||||
`/root/.pi/agent/extensions/zulip/.env` (`ABIBA_ZULIP_API_KEY`, mode 600,
|
||||
root-readable). **Do not place a new credential in the vault** — under the
|
||||
auth-keys charter that is a captain decision.
|
||||
|
||||
## What counts as a failure
|
||||
|
||||
A run FAILS (exit 1) when the report cannot be trusted or delivered:
|
||||
|
||||
* any probe is unreachable, so a section would silently be empty;
|
||||
* `PVE_TOKEN` is missing;
|
||||
* the Zulip credential is missing or rejected, or the upload/post fails.
|
||||
|
||||
There is **no degraded delivery leg any more**. Delivery is the only output
|
||||
path, so a missing credential is a failure rather than a survivable degradation —
|
||||
the previous "missing `EMAIL_PASSWORD` still exits 0" rule is retired with the
|
||||
mail transport.
|
||||
|
||||
**A delivery failure must never swallow the report.** On failure the script
|
||||
prints the report body to stdout *and* leaves the HTML artifact on disk, so the
|
||||
content is always recoverable from the run log. That closes the queued defect
|
||||
where a failed send printed only the transport error and the report never
|
||||
surfaced.
|
||||
|
||||
## Failure behaviour
|
||||
|
||||
* Non-zero exit with the alert text above; on the scheduled path the dispatch is
|
||||
a firstmate message, so the ops lane sees it and reports it.
|
||||
* On a probe failure the report must **not** be treated as evidence about the
|
||||
estate — a `0/0` Proxmox section means "could not look", not "nothing there".
|
||||
That reading is why the 2026-09-25 defect went unnoticed.
|
||||
|
||||
## Verification
|
||||
|
||||
```bash
|
||||
# data path, no email
|
||||
cd /root/abiba-workspace/projects/prose-contracts
|
||||
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \
|
||||
| grep -E 'pve_probe_status|node_count|nodes_online'
|
||||
|
||||
# delivery path
|
||||
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-zulip
|
||||
```
|
||||
|
||||
Regression tests: `tests/test_daily_infra_report.py` (7 tests). Four of them
|
||||
fail against the pre-fix script, which is what makes them bite.
|
||||
|
||||
## Maintains
|
||||
|
||||
- daily-infra-dashboard: { status: "ok|undelivered", transport: zulip-dm-attachment, last_check: timestamp }
|
||||
- pve-probe: { status: "ok|unreachable", last_check: timestamp }
|
||||
@@ -1,75 +0,0 @@
|
||||
# Delivery Record — HERMES-PLAYBOOK-FOR-SCOT
|
||||
|
||||
## Status: SEND-READY — awaiting Kwame's channel + recipient confirmation
|
||||
|
||||
No documented channel to Scot Murray exists in this workspace, the skills, or config
|
||||
(verified 2026-09-11 by sweep of `~/syslog/projects/murray-capital/`, `syslog-infra`
|
||||
references, `murray-harness` skill, `.hermes/memories/`, all of `~/syslog/`).
|
||||
Per the card's unblock constraints: package prepared, exact send commands written
|
||||
below, nothing transmitted. Guessing an address is out of scope.
|
||||
|
||||
## Verified artifact (single source of truth)
|
||||
|
||||
| Item | Value |
|
||||
|---|---|
|
||||
| Markdown source | `/home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.md` |
|
||||
| sha256 | `b0966a649fec96da4975ec00e627fbaba3a92a62c4a92b33bc06589c47d25f7f` |
|
||||
| Size | 28821 bytes, 377 lines |
|
||||
| Matches reviewed bytes | YES — identical to `/home/hermes/syslog/drafts/scot-hermes-playbook/` copy and to the hash recorded on card t_2c716052 |
|
||||
|
||||
## Rendered PDF (from the verified bytes, no edits)
|
||||
|
||||
| Item | Value |
|
||||
|---|---|
|
||||
| PDF | `/home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf` |
|
||||
| sha256 | `f46c0c89b4bbf6fa76c1f1c385c87753860d06bfa9e8925d6e09f27ed78a187b` |
|
||||
| Size | 96,486 bytes · 13 pages · A4 |
|
||||
| Render chain | pandoc 3.1.11.1 (gfm → html5) + WeasyPrint 62.3, stylesheet `pb.css`; reproducible via `bash render_pdf.sh` |
|
||||
| Spot-check | pdftotext shows correct title page + v0.21.1 verification note |
|
||||
|
||||
## Candidate channels — exact commands (pending Kwame's pick + address)
|
||||
|
||||
### 1. Email via syslog-email profile (recommended)
|
||||
|
||||
Mailbox ops belong to the syslog-email profile per standing rule. Send as
|
||||
jerome@sysloggh.com with both attachments.
|
||||
|
||||
```
|
||||
hermes -p syslog-email chat -q "Send an email. From jerome@sysloggh.com. \
|
||||
To: <SCOT-ADDRESS — Kwame to supply>. Subject: 'Hermes Playbook — getting real mileage out of the harness'. \
|
||||
Attach: /home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf \
|
||||
and HERMES-PLAYBOOK-FOR-SCOT.md. Body: short intro noting the PDF is the reviewed v0.21.1 playbook, \
|
||||
sha256 b0966a64… (sic, abbreviated), ask him to flag anything confusing — that feedback feeds the harness build. \
|
||||
Show me the draft before sending."
|
||||
```
|
||||
|
||||
Direct himalaya (only if Kwame wants it from the main session — normally NOT, per
|
||||
the email-routing standing rule):
|
||||
|
||||
```
|
||||
himalaya message write --to "<SCOT-ADDRESS>" --subject "Hermes Playbook — getting real mileage out of the harness" \
|
||||
--attachment .../HERMES-PLAYBOOK-FOR-SCOT.pdf --attachment .../HERMES-PLAYBOOK-FOR-SCOT.md
|
||||
himalaya message send <draft.eml>
|
||||
```
|
||||
|
||||
### 2. Telegram (only if Kwame has Scot's handle)
|
||||
|
||||
Send the PDF to Scot's handle from the gateway-connected Telegram session:
|
||||
|
||||
```
|
||||
hermes chat -q "Send the file /home/hermes/syslog/prose-contracts/deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.pdf to <SCOT-HANDLE> with a one-line intro."
|
||||
```
|
||||
|
||||
### 3. Anything else (WhatsApp, shared drive, print+hand-deliver)
|
||||
|
||||
Needs Kwame's input on mechanism; the PDF + MD at the paths above are the payload.
|
||||
|
||||
## Post-send obligations (from the card)
|
||||
|
||||
1. Record here: channel, timestamp, exact bytes + sha256 sent, any acknowledgement.
|
||||
2. Capture Scot's feedback as evidence (what he tried first, what confused him,
|
||||
which of the 17 videos he watched).
|
||||
3. Feed findings into `murray-harness` skill (+ `hermes-kanban-ops` if tooling
|
||||
lessons surface).
|
||||
4. If no reply in 7 days: ONE follow-up nudge is in scope; more is Kwame's call.
|
||||
5. Feature gaps he reports → separate card, do not widen this one.
|
||||
@@ -1,377 +0,0 @@
|
||||
# The Hermes Playbook — getting real mileage out of the harness
|
||||
|
||||
Prepared for Scot (Syslog Solution LLC). Version: Hermes Agent v0.21.1. Every CLI command below was verified live against that version on a reference install (Syslog kagentz) on 2026-09-11; anything only confirmed against the official docs is tagged DOC-ONLY.
|
||||
|
||||
You are already running Hermes next to Claude Code, on your own OpenRouter account with fast models (qwen3.8-flash, deepseek-4-flash). The question you asked: why does Hermes feel like it has less context, and what do I do about it?
|
||||
|
||||
---
|
||||
|
||||
## 1. TL;DR
|
||||
|
||||
- The context gap is not a bug. Claude Code reads the repo it sits in on every launch; a fresh Hermes install starts nearly empty by design. It gets its context from files you seed and from memory it builds over time.
|
||||
- One command closes most of the gap on day one: `hermes import-agent claude-code` carries your CLAUDE.md instructions, MCP servers, skills, and memories into Hermes (preview first with `--dry-run`).
|
||||
- Teach Hermes once, and it remembers: "save this as a skill" after any workflow you repeat. Skills auto-load when a matching task comes up — that is the learning loop.
|
||||
- Keep per-project context in an `AGENTS.md` in the repo root (git-tracked, shared with your team) and personal preferences in your persona file and persistent memory.
|
||||
- Hermes and Claude Code are not rivals: let Hermes be the always-on orchestrator (research, briefs, scheduling, messaging) and hand heavy coding to Claude Code, which Hermes can drive directly.
|
||||
|
||||
---
|
||||
|
||||
## 2. Why Hermes feels like it has less context (and why that is fixable)
|
||||
|
||||
Honest comparison, no spin:
|
||||
|
||||
| | Claude Code | Hermes (fresh install) |
|
||||
|---|---|---|
|
||||
| Where context comes from | The repo: `CLAUDE.md` auto-loaded every launch; `.claude/` folders with subagents, slash commands, hooks, skills | Config files: `AGENTS.md` in the working directory + `SOUL.md` persona + persistent memory from the Hermes home |
|
||||
| What it remembers between sessions | `~/.claude/projects/<project>/memory/` (25 KB cap) | First-class persistent memory, always injected — `MEMORY.md` / `USER.md` plus optional external providers |
|
||||
| How it learns your workflows | You write the skill/command files | It can write its own skills after learning a workflow, and a curator maintains them |
|
||||
| Out-of-the-box feel | Context-rich if you have invested in your CLAUDE.md | Quiet until you seed it |
|
||||
|
||||
That last line is the whole story. Claude Code's context is the sum of everything you built in `CLAUDE.md` and `.claude/` over months. A fresh Hermes has none of that yet — not because the harness is weaker, but because it stores context in different places and expects you to seed it (or let it build up).
|
||||
|
||||
The gap is fixable in two moves:
|
||||
|
||||
1. **Import what you already have.** `hermes import-agent claude-code` maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, and memories into Hermes equivalents. Preview with `--dry-run`; it never imports API keys; conflicts are skipped by default (`--overwrite` to change).
|
||||
2. **Let the learning loop run.** Every time you correct Hermes or finish a workflow you will repeat, tell it to remember. Within a few weeks it will have its own CLAUDE.md equivalent — built, not typed.
|
||||
|
||||
What the comparison table in our research covers, gap by gap: project instructions, instruction splits, slash commands, subagents, skills, project memory, tool permissions, MCP, session resume, cost/context visibility, headless mode, prior-setup import, hooks, and scheduled work. Each has a Hermes equivalent, and every one is documented in section 9.
|
||||
|
||||
---
|
||||
|
||||
## 3. The context stack
|
||||
|
||||
This is the order in which Hermes builds its context, and what you do at each layer.
|
||||
|
||||
**Layer 1 — Persona (`SOUL.md`).** Set up once. Your Hermes' standing identity and voice: "you are my analyst," the tone, the standing rules. Auto-injected into every session. Lives at `~/.hermes/SOUL.md` (per profile: `~/.hermes/profiles/<name>/SOUL.md`). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/configuration
|
||||
|
||||
**Layer 2 — Persistent memory.** Set up once, then feed it constantly. `MEMORY.md` / `USER.md` are always active and injected every session — this is the single biggest cure for "it forgets my project." After any correction or preference ("use this source list," "briefs go in this format"), tell Hermes to remember it. Manage with `hermes memory setup|status|off|reset` (VERIFIED-LIVE). Optional external providers exist (Honcho, Mem0, and others). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/features/memory
|
||||
|
||||
**Layer 3 — Skills.** Set up once; grows forever. Markdown procedure files that auto-load when a task matches the skill. The differentiator: after completing a workflow, ask Hermes to "save this as a skill" — it authors the skill itself, and a background curator tracks usage, archives stale ones, and keeps backups. CLI: `hermes skills list|search|install|browse|config|check|update` (VERIFIED-LIVE); in-session: `/skill <name>`, `/reload-skills` (DOC-ONLY). Docs: https://hermes-agent.nousresearch.com/docs/reference/skills-catalog and https://hermes-agent.nousresearch.com/docs/user-guide/features/curator
|
||||
|
||||
**Layer 4 — Projects.** Per workstream. `AGENTS.md` in each repo root (git-tracked, team-shared) carries project rules; Desktop Projects (`hermes project create <name>` then `add-folder`) group multi-repo work under one named workspace. Both VERIFIED-LIVE.
|
||||
|
||||
**Layer 5 — Retrieval (session store).** Automatic. All conversations land in a searchable store; Hermes can search past sessions when you ask "what did we decide last week." CLI: `hermes sessions list|browse|rename|pin|export|prune|stats` (VERIFIED-LIVE).
|
||||
|
||||
**Layer 6 — MCP (external tools).** Per integration. Plug GitHub, databases, workflow engines into the agent. `hermes mcp add|list|test|configure|picker|catalog|install` (VERIFIED-LIVE). Docs: https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp
|
||||
|
||||
Quick summary:
|
||||
|
||||
| Layer | Set up | Feed |
|
||||
|---|---|---|
|
||||
| SOUL.md persona | once | rarely |
|
||||
| Persistent memory | once | every correction/preference |
|
||||
| Skills | once | "save this as a skill" after repeated workflows |
|
||||
| AGENTS.md / Projects | once per repo/workstream | as projects evolve |
|
||||
| Session retrieval | automatic | ask |
|
||||
| MCP | once per integration | when new tools appear |
|
||||
|
||||
---
|
||||
|
||||
## 4. Top moves
|
||||
|
||||
The highest-leverage moves for your kind of work — research, evidence-graded analysis, weekly briefs — and for running alongside Claude Code. Every command verified on v0.21.1.
|
||||
|
||||
### 1. Import your Claude Code setup
|
||||
|
||||
```bash
|
||||
hermes import-agent claude-code --dry-run # preview
|
||||
hermes import-agent claude-code # migrate CLAUDE.md, MCP, skills, memories
|
||||
```
|
||||
|
||||
What it does: one-command migration of the instructions and servers that made Claude Code feel context-rich, translated into Hermes equivalents. Never imports API keys.
|
||||
Why it matters: this is the direct answer to "Hermes has no context." After this, Hermes knows your projects on day one.
|
||||
|
||||
### 2. Bring over the conversation history
|
||||
|
||||
```bash
|
||||
hermes sessions import
|
||||
```
|
||||
|
||||
What it does: imports Claude Code or Codex CLI conversations into the Hermes session store.
|
||||
Why it matters: mid-project, the new agent picks up exactly where the old one left off. `hermes --resume <id>` and `hermes sessions browse` then treat the history as native.
|
||||
|
||||
### 3. Trust your repos so project-local skills load
|
||||
|
||||
```bash
|
||||
hermes skills trust
|
||||
```
|
||||
|
||||
What it does: trusts a repo so its project-local skills (`.hermes/skills`) load — the Hermes analog of `.claude/skills/`.
|
||||
Why it matters: your evidence-grading rules can live in the repo with the project, versioned with git, and load automatically.
|
||||
|
||||
### 4. Per-directory session continuity
|
||||
|
||||
```bash
|
||||
hermes --in DIR --resume latest
|
||||
```
|
||||
|
||||
What it does: resumes the latest session for a given directory (also: `hermes -c [NAME]`, `hermes --resume <id|latest>`).
|
||||
Why it matters: every project folder gets its own continuous thread. Research on one portfolio never mixes with another.
|
||||
|
||||
### 5. Preload skills for a specific job
|
||||
|
||||
```bash
|
||||
hermes -s skill1,skill2
|
||||
```
|
||||
|
||||
What it does: preloads specific skills for the session.
|
||||
Why it matters: for a weekly brief or an evidence register, pin the exact skills that encode your grading criteria instead of hoping they auto-match.
|
||||
|
||||
### 6. Save any repeated workflow as a skill
|
||||
|
||||
In-session: "save this as a skill." (CLI: `hermes skills list|search|install|browse`.)
|
||||
|
||||
What it does: Hermes authors a skill file from the workflow you just ran.
|
||||
Why it matters: the learning loop is the whole point. Do the evidence-grading pass twice, save it, and every future run starts with the procedure loaded.
|
||||
|
||||
### 7. Fan out research with subagents
|
||||
|
||||
In-session: "delegate this to subagents." (agent-side tool `delegate_task`, no CLI.)
|
||||
|
||||
What it does: parallel subagents with isolated contexts — each gets its own conversation and terminal, only the final summary comes back.
|
||||
Why it matters: research fan-out without flooding your main context. Ten sources, ten subagents, one synthesis.
|
||||
|
||||
### 8. Make the weekly brief a cron job
|
||||
|
||||
```bash
|
||||
hermes cron create
|
||||
```
|
||||
|
||||
What it does: durable scheduler — duration or cron syntax, per-job model overrides, output chaining, delivery to messaging platforms. Manage with `hermes cron list|create|edit|pause|resume|run|remove|doctor`.
|
||||
Why it matters: a weekly brief is exactly a cron job. It runs even when you are not at the desktop, with your skills preloaded and its output delivered to you.
|
||||
|
||||
### 9. Set a standing goal for grind work
|
||||
|
||||
In-session: `/goal [text|status|pause|resume|clear]` (DOC-ONLY; CLI subcommands verified).
|
||||
|
||||
What it does: a standing objective the agent keeps working toward across turns until achieved.
|
||||
Why it matters: "keep researching until you have 5 verified sources" — the agent loops itself instead of waiting for you to say "go on."
|
||||
|
||||
### 10. Run Hermes as an MCP server for Claude Code
|
||||
|
||||
```bash
|
||||
hermes mcp serve
|
||||
```
|
||||
|
||||
What it does: exposes Hermes (persistent memory, skills, cron, sessions) to other agents as an MCP tool provider. Claude Code supports MCP clients, so it can consume Hermes.
|
||||
Why it matters: the reverse bridge. Claude Code gets the surfaces it lacks, and both tools share your knowledge base.
|
||||
|
||||
### 11. Pick the right model per task, with a safety net
|
||||
|
||||
```bash
|
||||
hermes fallback list|add|remove
|
||||
hermes -m MODEL --provider PROVIDER --reasoning high
|
||||
```
|
||||
|
||||
What it does: explicit fallback chains (a failed call rolls to a second model instead of erroring) and per-run model/provider/reasoning overrides.
|
||||
Why it matters: on OpenRouter with fast models, use `--reasoning high` for the hard analytical passes and let fallback chains keep the cheap models from stalling your brief.
|
||||
|
||||
### 12. Diagnose why responses feel thin
|
||||
|
||||
```bash
|
||||
hermes prompt-size
|
||||
```
|
||||
|
||||
What it does: byte breakdown of the system prompt + tool schemas.
|
||||
Why it matters: when output quality drops, it is usually context bloat, not model quality. This tells you what is eating the window.
|
||||
|
||||
---
|
||||
|
||||
## 5. Working alongside Claude Code
|
||||
|
||||
You run both. The proven patterns, in order of value.
|
||||
|
||||
**First: import.** `hermes import-agent claude-code` then `hermes sessions import`. After this, the "two tools that don't know each other" problem is gone — Hermes knows your projects and your history.
|
||||
|
||||
**Hermes as orchestrator, Claude Code as worker.** The installed Hermes skill for exactly this is `autonomous-ai-agents/delegate-coding-agent`. Two modes:
|
||||
|
||||
- Print mode (preferred): `claude -p '<task>' --allowedTools 'Read,Edit' --max-turns 10` — one-shot, no dialogs, structured JSON output with `session_id`, `num_turns`, `total_cost_usd`. In Hermes, just say: "delegate this coding task to Claude Code in print mode."
|
||||
- Interactive PTY via tmux: Hermes starts a tmux session, sends prompts with `send-keys`, monitors with `capture-pane`. For iterative refactor → review → fix cycles.
|
||||
|
||||
There is also a cross-agent review loop: `git diff main...feature | claude -p 'Review this diff for bugs and security issues.' --max-turns 1` — Hermes runs it, reads the findings, and fixes them itself. Claude Code becomes a reviewer Hermes coordinates.
|
||||
The skill's safety rails: explicit workdir, clean git status before launch, narrow task prompts, git diff review, targeted tests before committing.
|
||||
|
||||
**Parallel workstreams, neutral merge reconciliation.** When both agents edit the same repo and collide, do not let either resolve the conflict — both are biased toward their own side. Spawn a neutral third agent with the `merge-reconciler` skill: it classifies every conflicted hunk, resolves under an impartiality contract, verifies with build/tests, and hands back a summary naming every decision. Kanban shape: a reconciliation card assigned to a third profile, with both workers' cards as parents.
|
||||
|
||||
**Hermes as MCP server (the reverse direction).** `hermes mcp serve` (pattern in section 4, move 10). The only bridge direction Claude Code cannot offer.
|
||||
|
||||
**Desktop GUI goes to Hermes.** Claude Code has no desktop automation. `hermes computer-use install` (cua-driver; health check `hermes computer-use doctor`) drives native desktop apps background-first — never steals focus. If a task needs Excel or a native app, that part routes to Hermes while the code routes to Claude Code.
|
||||
|
||||
**Division of labor in one line:** Hermes is the always-on layer — research, briefs, scheduling, messaging, memory, and the orchestration desk. Claude Code is the deep coding worker. Hand coding-heavy tasks over; hand continuity, recall, and scheduled work to Hermes.
|
||||
|
||||
---
|
||||
|
||||
## 6. Video watch list
|
||||
|
||||
Every link verified via the YouTube oEmbed endpoint on 2026-09-11 (status PASS, title/author matched). All content is third-party ecosystem material — no official Nous Research tutorial video exists (see section 8).
|
||||
|
||||
| # | Title | Channel | Length | Link | What it demonstrates | Watch when you want to |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 1 | Learn 95% of Hermes Agent in 31 Minutes | Sharbel A. | 31:28 | https://www.youtube.com/watch?v=Ta2wg6xPaY4 | End-to-end fundamentals: install, sessions, skills, memory, the learning loop | the fastest real overview of the whole harness before touching config |
|
||||
| 2 | Hermes Agent Fundamentals In 29 Minutes | Tina Huang | 29:40 | https://www.youtube.com/watch?v=5_N84t1rUU0 | Why Hermes' memory/skills loop differs from one-shot coding agents | understand why Hermes feels different from Claude Code |
|
||||
| 3 | Every Level of Hermes Agent Explained | Jack Roberts | 25:35 | https://www.youtube.com/watch?v=6GtF_uHbGhw | Beginner to advanced ladder: memory, skills, automation, multi-agent | a map of what to learn next after the basics |
|
||||
| 4 | Hermes Agent Full Tutorial INSTALLATION + USECASES | CodeHead | 7:47 | https://www.youtube.com/watch?v=8GjyOQy19so | Install through real use cases, compact | a quick install-to-value demo to share with a colleague |
|
||||
| 5 | Hermes Agent Explained In 5 Minutes | CodeHead | 4:53 | https://www.youtube.com/watch?v=9GpWELm3_XI | Conceptual pitch of the agent and its learning loop | the elevator pitch before committing 30 minutes |
|
||||
| 6 | 100 Days With Hermes Agent in 21 Minutes | Sharbel A. | 21:19 | https://www.youtube.com/watch?v=sCa3BtpkziQ | What memory/skills accumulation looks like after months of daily use | see the payoff of the learning loop over time |
|
||||
| 7 | Hermes Agent - Crash Course for Beginners (AI Agent) | Adrian Twarog | 22:19 | https://www.youtube.com/watch?v=4sAmpcSOVEw | Beginner crash course from a well-known dev channel | a second independent explanation of the basics |
|
||||
| 8 | Hermes Agent: The Ultimate Beginner's Guide | Metics Media | 37:08 | https://www.youtube.com/watch?v=CwPUOVUdApE | Long-form beginner guide incl. setup and everyday workflows | the most thorough single walkthrough in one sitting |
|
||||
| 9 | Hermes Agent Just Killed OpenClaw (Full Tutorial) | Leon van Zyl | 19:59 | https://www.youtube.com/watch?v=jmtpYUOr7_U | Feature-by-feature tutorial (MCP config, memory, agents) | a practitioner's feature-by-feature tutorial |
|
||||
| 10 | Hermes Agent vs OpenClaw | Sharbel A. | 15:28 | https://www.youtube.com/watch?v=zwqhemjHq3E | Head-to-head comparison of the two agent harnesses | the tradeoffs between Hermes and its main alternative |
|
||||
| 11 | Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model | Tonbi's AI Garage | 15:08 | https://www.youtube.com/watch?v=8tpuky8HpXw | Hermes driven by an OpenRouter-served open model | how small open models behave inside Hermes |
|
||||
| 12 | Use This To Make The Hermes Agent Basically Free | AI LABS | 13:08 | https://www.youtube.com/watch?v=5d02TYoOzfE | Running Hermes on cheap/free model backends | cut inference costs on an OpenRouter account |
|
||||
| 13 | Hermes Agent The 24/7 Self-Evolving AI Agent! | WorldofAI | 9:15 | https://www.youtube.com/watch?v=cu2fgknmemA | Always-on operation: gateway, cron, background automation | turn Hermes from a chat window into a 24/7 assistant |
|
||||
| 14 | Hermes Co-Founder on Building an AI Agent That Improves Itself \| Karan Malhotra | Peter Yang | 46:45 | https://www.youtube.com/watch?v=UWjh5Z4s8jY | Interview on design philosophy (self-improving agents, skills as memory) | where the product is going |
|
||||
| 15 | Hermes Agent: Agents that grow with you \| Episode #357 | Practical AI | 47:34 | https://www.youtube.com/watch?v=UTZhvPXnmwA | Podcast-depth technical discussion of the agent architecture | the engineering story behind the learning loop |
|
||||
| 16 | Did Hermes Agent just kill OpenClaw? (full guide) | Alex Finn | 13:55 | https://www.youtube.com/watch?v=tP6yf22OJdI | Guide-style comparison/switch content | a switcher's guide perspective |
|
||||
| 17 | Hermes Agent: Why Everyone's Ditching OpenClaw in 2026 | Luke Alexander AI | 18:03 | https://www.youtube.com/watch?v=1UgXUjT-QtI | Comparison content | more comparison context |
|
||||
|
||||
Suggested order: 1 or 5 first (whichever mood you are in), then 2, then 6 once you have a few weeks of use under your belt.
|
||||
|
||||
---
|
||||
|
||||
## 7. Your first 7 days
|
||||
|
||||
One action per day, each finishable in 15 minutes.
|
||||
|
||||
**Day 1 — Import.** `hermes import-agent claude-code --dry-run`, review the preview, then run it without the flag. Your CLAUDE.md context now lives in Hermes.
|
||||
|
||||
**Day 2 — Write your SOUL.md.** Open `~/.hermes/SOUL.md` and write who this agent is for you: its role, your tone, three standing rules (e.g., how to grade evidence, where briefs go, how to flag uncertainty). Ten lines is plenty.
|
||||
|
||||
**Day 3 — Per-directory sessions.** Pick your most active project folder. Work one task there via `hermes --in DIR --resume latest`. Notice the thread is separate from everything else.
|
||||
|
||||
**Day 4 — First skill.** Finish a small repeated workflow (a source-check pass, a brief section). At the end, say "save this as a skill." Next day, watch it load by itself.
|
||||
|
||||
**Day 5 — One cron job.** `hermes cron create` for a small daily check (inbox digest, a price or news watch, whatever you already do by hand). Deliver it somewhere you actually look.
|
||||
|
||||
**Day 6 — Hand a task to Claude Code.** In Hermes: "delegate this coding task to Claude Code in print mode." Read the JSON result. This is the bridge working.
|
||||
|
||||
**Day 7 — Recall test.** Ask Hermes "what did we decide last week about [your project]?" If it can answer from the session store, the stack is working. If not, `/compress` the bloat and try `hermes prompt-size` to see what is eating the window.
|
||||
|
||||
---
|
||||
|
||||
## 8. What NOT to expect
|
||||
|
||||
- **A bigger context window than you have.** Model choice does not change the window size. Fast models on OpenRouter (qwen3.8-flash, deepseek-4-flash) are cheap and quick, but they carry fewer bytes per turn than a frontier model. The harness compresses automatically near the limit — you will not watch a meter like Claude Code's `/context` — but compression is lossy. For the heaviest analytical passes, use `--reasoning high` and a larger model for that run.
|
||||
- **Model choice as a silver bullet.** What a different model buys: better reasoning, better tool-calling, more reliable long-horizon work. What it does not buy: memory of your projects, your workflows, or last week's decisions. That lives in your context stack, not the model.
|
||||
- **Desktop = everything.** The desktop app is a thin client over a local agent: config, memory, skills, sessions, cron, and kanban all live in the Hermes home, not in the window. Close the window and the work keeps living; that is a feature, not a bug.
|
||||
- **Cron limits.** Cron jobs are durable, but they run on their own budgets: wall-clock caps, per-job model overrides, and delivery depends on configured platforms. A job is not an infinite second brain — design it as a bounded task with a bounded output.
|
||||
- **It will still need to be told things twice.** If you did not save it as memory or a skill, the next session does not know. The learning loop only works if you trigger it. "Remember this" and "save this as a skill" are deliberate moves, not magic.
|
||||
- **Official tutorial videos.** None exist from Nous Research; the watch list is verified third-party content. The docs (hermes-agent.nousresearch.com/docs) are the authoritative source, and `/help` inside a session lists the exact commands your version supports.
|
||||
- **Slash commands behave like the CLI does.** The slash registry is version-dependent; anything tagged DOC-ONLY here was confirmed against the docs but not exercised live from a headless session. `/help` in your own session is the final word.
|
||||
|
||||
---
|
||||
|
||||
## 9. Appendix: command reference
|
||||
|
||||
Tags: **VERIFIED-LIVE** = confirmed against `hermes --help` / `hermes <cmd> --help` on v0.21.1 (2026.9.7), reference install (Syslog kagentz), 2026-09-11. **DOC-ONLY** = confirmed against the official docs (slash commands run inside a chat session and were not exercised from a headless research run; their CLI subcommands were verified live).
|
||||
|
||||
### Setup & health
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes setup` | Interactive setup wizard | VERIFIED-LIVE |
|
||||
| `hermes doctor [--fix] [--live]` | Diagnose config/deps; `--fix` auto-repairs | VERIFIED-LIVE |
|
||||
| `hermes status [--all] [--deep]` | Component status | VERIFIED-LIVE |
|
||||
| `hermes config show/edit/get/set/unset/path/env-path/check/migrate` | View/edit config | VERIFIED-LIVE |
|
||||
| `hermes update` | Update Hermes to latest | VERIFIED-LIVE |
|
||||
|
||||
### The Claude Code bridge (highest value for you)
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes import-agent claude-code [--dry-run] [--overwrite] [--yes]` | One-command import of a Claude Code setup: maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, memories into Hermes equivalents. Never imports API keys. | VERIFIED-LIVE |
|
||||
| `hermes import-agent codex` | Same for Codex CLI setups | VERIFIED-LIVE |
|
||||
| `hermes sessions import` | Import a Claude Code or Codex CLI session/conversation into Hermes | VERIFIED-LIVE |
|
||||
| `hermes skills trust` | Trust a repo so its project-local skills (`.hermes/skills`) load — the Hermes analog of `.claude/skills/` | VERIFIED-LIVE |
|
||||
|
||||
### Daily driving
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes` / `hermes chat` | Interactive session | VERIFIED-LIVE |
|
||||
| `hermes -c [NAME]` / `hermes --resume <id\|latest>` | Resume by name or ID | VERIFIED-LIVE |
|
||||
| `hermes --in DIR --resume latest` | Resume the latest session for a directory | VERIFIED-LIVE |
|
||||
| `hermes -z "PROMPT"` | One-shot: prints ONLY the final answer (scripting/CI); tools, memory, and AGENTS.md still load | VERIFIED-LIVE |
|
||||
| `hermes chat -q "PROMPT"` | Single-query mode | VERIFIED-LIVE |
|
||||
| `hermes -m MODEL --provider PROVIDER --reasoning LEVEL` | Per-run model/provider/reasoning overrides (`none…ultra`) | VERIFIED-LIVE |
|
||||
| `hermes -s SKILL1,SKILL2` | Preload specific skills for the session | VERIFIED-LIVE |
|
||||
| `hermes -t TOOLSETS` | Restrict toolsets for this run | VERIFIED-LIVE |
|
||||
| `hermes -w` | Isolated git worktree session (parallel agents on one repo) | VERIFIED-LIVE |
|
||||
| `hermes chat --checkpoints` | Enable filesystem checkpoints (`/rollback` to restore) | VERIFIED-LIVE |
|
||||
| `hermes chat --max-turns N` / `--run-budget SECONDS` | Cap loop iterations / wall-clock budget | VERIFIED-LIVE |
|
||||
| `hermes -yolo` | Bypass command approval prompts (use with care) | VERIFIED-LIVE |
|
||||
| `hermes pause` / `hermes resume` | Emergency stop / lift (pauses cron, kanban dispatch, gateway turns) | VERIFIED-LIVE |
|
||||
|
||||
### Context & memory management
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes memory setup/status/off/reset` | External memory provider management (built-in MEMORY.md/USER.md always active) | VERIFIED-LIVE |
|
||||
| `hermes sessions list/browse/rename/pin/export/prune/stats` | Session store management | VERIFIED-LIVE |
|
||||
| `hermes skills list/search/install/inspect/browse/config/check/update` | Skill management | VERIFIED-LIVE |
|
||||
| `hermes skills trust/untrust` | Repo-local skill trust | VERIFIED-LIVE |
|
||||
| `hermes curator status/run/pause/pin/...` | Background skill maintenance (auto-archive, backups) | VERIFIED-LIVE |
|
||||
| `hermes prompt-size` | Byte breakdown of system prompt + tool schemas (context-bloat diagnosis) | VERIFIED-LIVE |
|
||||
| `hermes insights [--days N]` | Usage analytics | VERIFIED-LIVE |
|
||||
|
||||
### Tools, MCP, integrations
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes tools` (interactive) / `list/enable/disable` | Per-platform toolset toggles; MCP tools as `server:tool` | VERIFIED-LIVE |
|
||||
| `hermes mcp add/remove/list/test/configure/picker/catalog/install` | MCP server management (incl. one-click catalog installs) | VERIFIED-LIVE |
|
||||
| `hermes mcp serve` | Run Hermes AS an MCP server for other agents | VERIFIED-LIVE |
|
||||
| `hermes computer-use install/status/doctor` | Desktop-control backend (cua-driver) | VERIFIED-LIVE |
|
||||
| `hermes gateway run/install/start/status/setup` | Messaging gateway (Telegram, Discord, Slack, WhatsApp, …) | VERIFIED-LIVE |
|
||||
| `hermes send` | Send a message to a configured platform (scripts/cron/CI) | VERIFIED-LIVE |
|
||||
|
||||
### Automation & multi-agent
|
||||
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes cron list/create/edit/pause/resume/run/remove/doctor` | Scheduled jobs (durable, multi-platform delivery) | VERIFIED-LIVE |
|
||||
| `hermes cron notepad` | Durable per-job key-value notepad across runs | VERIFIED-LIVE |
|
||||
| `hermes kanban create/list/show/link/complete/swarm/...` | Durable multi-profile task board (40+ verbs) | VERIFIED-LIVE |
|
||||
| `hermes kanban swarm` | Generate a parallel-workers → verifier → synthesizer task graph | VERIFIED-LIVE |
|
||||
| `hermes project create/list/add-folder/bind-board` | Named multi-folder workspaces (desktop Projects) | VERIFIED-LIVE |
|
||||
| `hermes profile list/create/use/alias/export/import` | Isolated Hermes instances | VERIFIED-LIVE |
|
||||
| `hermes auth add/list/priority/reset` | Pooled credentials per provider (rotation) | VERIFIED-LIVE |
|
||||
| `hermes fallback list/add/remove` | Fallback model chain (auto-rollover on failure) | VERIFIED-LIVE |
|
||||
| `hermes model` | Interactive model/provider picker | VERIFIED-LIVE |
|
||||
|
||||
### In-session slash commands (DOC-ONLY)
|
||||
|
||||
Source: https://hermes-agent.nousresearch.com/docs/reference/slash-commands
|
||||
|
||||
| Command | What it does |
|
||||
|---|---|
|
||||
| `/help` | List all commands (authoritative in your version) |
|
||||
| `/new` (`/reset`) | Fresh session |
|
||||
| `/resume [name]` | Resume a named/recent session |
|
||||
| `/branch` (`/fork`) | Branch the current session |
|
||||
| `/compress` | Manually compress context (auto-compression also exists) |
|
||||
| `/undo` | Remove last exchange |
|
||||
| `/retry` | Resend last message |
|
||||
| `/title [name]` | Name the session |
|
||||
| `/save` | Save conversation to file |
|
||||
| `/history` | Show conversation history |
|
||||
| `/skill <name>` | Load a skill into the session |
|
||||
| `/skills` | Search/install skills |
|
||||
| `/reload-skills` | Re-scan skill directory |
|
||||
| `/tools` / `/toolsets` | Manage tools |
|
||||
| `/goal [text]` | Set a standing goal the agent works toward across turns (`/goal status/pause/clear` to manage) |
|
||||
| `/background <prompt>` | Run a prompt in the background |
|
||||
| `/queue <prompt>` | Queue a prompt for the next turn |
|
||||
| `/steer <prompt>` | Inject a course-correction after the next tool call without interrupting |
|
||||
| `/agents` | Show active agents and running tasks |
|
||||
| `/cron` | Manage cron jobs in-session |
|
||||
| `/kanban` | Multi-profile collaboration board in-session |
|
||||
| `/model [name]` | Show/change model mid-session |
|
||||
| `/reasoning [level]` | Set reasoning effort |
|
||||
| `/voice [on\|off\|tts]` | Voice mode |
|
||||
| `/rollback [N]` | Restore filesystem checkpoint (needs `--checkpoints`) |
|
||||
| `/usage` | Token usage |
|
||||
| `/insights [days]` | Usage analytics |
|
||||
| `/platforms` | Gateway platform status |
|
||||
|
||||
Note: Hermes compresses automatically near the context limit; no manual threshold watch is needed the way Claude Code's `/context` grid is.
|
||||
Binary file not shown.
@@ -1,100 +0,0 @@
|
||||
# Review Results: Scot Murray Hermes Playbook (t_fefdf30b)
|
||||
|
||||
**VERDICT: APPROVED-WITH-FIXES**
|
||||
|
||||
## Summary
|
||||
The playbook is well-structured, factually accurate, and provides genuine value for a new Hermes user. All 17 video links verified live (17/17 PASS), all 10 source URLs resolved successfully, all 26+ CLI commands verified against v0.21.1, no client data leaks detected, and all 9 required sections present with substantive content. One minor documentation accuracy issue requires correction.
|
||||
|
||||
## Per-Check Results
|
||||
|
||||
### 1. COMMANDS: ✅ PASS
|
||||
- 26 top-level commands and subcommands verified live on Hermes Agent v0.21.1 (2026.9.7)
|
||||
- All VERIFIED-LIVE tags confirmed: `hermes import-agent claude-code --dry-run`, `hermes skills trust`, `hermes mcp serve`, `hermes prompt-size`, `hermes fallback`, `hermes curator`, etc.
|
||||
- All DOC-ONLY commands (in-session slash commands) confirmed against official docs
|
||||
- No fabricated or non-existent commands found
|
||||
|
||||
### 2. VIDEO LINKS: ✅ PASS
|
||||
- All 17 YouTube URLs verified via oEmbed endpoint
|
||||
- **17/17 PASS** - All titles and channels match the documentation claims
|
||||
- Videos: https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=<ID>&format=json
|
||||
- Example verified: Ta2wg6xPaY4 → "Learn 95% of Hermes Agent in 31 Minutes" | Sharbel A. ✅
|
||||
|
||||
### 3. SOURCES: ✅ PASS
|
||||
- 10/10 URLs in sources.md resolved successfully (HTTP 200)
|
||||
- No dead links or inaccessible URLs found
|
||||
- All official docs and GitHub repo accessible
|
||||
|
||||
### 4. COMPLETENESS: ✅ PASS
|
||||
- All 9 required sections present and substantive:
|
||||
- 1. TL;DR ✅
|
||||
- 2. Context gap explanation ✅
|
||||
- 3. Context stack ✅
|
||||
- 4. Top moves (12 items) ✅
|
||||
- 5. Working alongside Claude Code ✅
|
||||
- 6. Video watch list (17 videos) ✅
|
||||
- 7. 7-day ramp ✅
|
||||
- 8. What NOT to expect ✅
|
||||
- 9. Appendix: command reference ✅
|
||||
- Top moves count: 12/12 (within 12 limit) ✅
|
||||
- 7-day ramp is actionable with specific commands ✅
|
||||
|
||||
### 5. CLIENT-DATA LEAK: ✅ PASS
|
||||
- **No private financial data or personal data found**
|
||||
- Grep patterns searched: murray, jds, portfolio, allocation, holding, ticker, position, dollar, 192.168.68.17, syslog solution llc
|
||||
- Only mentions of "portfolio" are generic workflow descriptions, not specific financial data
|
||||
- No Murray Capital/JDS portfolio details, positions, or dollar figures found
|
||||
|
||||
### 6. HONESTY/OVER-CLAIM: ✅ PASS (1 minor issue)
|
||||
- **Minor issue found:** The "save this as a skill" workflow description is slightly misleading
|
||||
- Book says: "after you complete a workflow twice, ask Hermes to 'save this as a skill'"
|
||||
- Reality: The workflow works on a single workflow completion (not after two)
|
||||
- The language "after you complete a workflow twice" suggests a minimum repetition requirement that doesn't exist
|
||||
- **Recommendation:** Change to "after completing a workflow, ask Hermes to save this as a skill"
|
||||
- No major over-claims about features that don't exist
|
||||
- All model claims are accurate for OpenRouter-only setup
|
||||
- Honest about "no official Nous Research tutorial videos exist" ✅
|
||||
|
||||
### 7. USEFULNESS: ✅ PASS
|
||||
- **Strongest section:** Section 4 "Top moves" - provides 12 highly actionable, verified commands
|
||||
- **Strongest section:** Section 6 "Video watch list" - all links verified, titles/channels accurate
|
||||
- **Strongest section:** Section 7 "Your first 7 days" - practical, incremental onboarding plan
|
||||
- **Weakest section:** Section 2 "Why Hermes feels like it has less context" - could benefit from more concrete examples
|
||||
- Overall: Would genuinely help a new user close the context gap with actionable, verified steps
|
||||
|
||||
## Prioritized Fixes
|
||||
|
||||
### SHOULD-FIX
|
||||
1. **Fix "save as skill" workflow description** - Change "after you complete a workflow twice" to "after completing a workflow" (Section 3, paragraph 4)
|
||||
- This is the only minor issue found
|
||||
- Doesn't affect functionality but could create false expectations about repetition requirements
|
||||
|
||||
### NIT
|
||||
- None identified - all content is accurate and well-organized
|
||||
|
||||
## Edits Applied
|
||||
|
||||
Applied the SHOULD-FIX correction directly to the playbook:
|
||||
- **Section 3, Layer 3 (Skills):** Changed "after you complete a workflow twice" → "after completing a workflow"
|
||||
|
||||
---
|
||||
|
||||
## FINAL SUMMARY
|
||||
|
||||
**Verdict: APPROVED-WITH-FIXES**
|
||||
|
||||
**Per-check results:**
|
||||
- Check 1 (COMMANDS): ✅ PASS - 26+ commands verified live
|
||||
- Check 2 (VIDEO LINKS): ✅ PASS - 17/17 valid with matching titles/channels
|
||||
- Check 3 (SOURCES): ✅ PASS - 10/10 URLs resolved
|
||||
- Check 4 (COMPLETENESS): ✅ PASS - 9/9 sections, 12 top moves, actionable 7-day ramp
|
||||
- Check 5 (CLIENT-DATA LEAK): ✅ PASS - No private data found
|
||||
- Check 6 (HONESTY/OVER-CLAIM): ✅ PASS - 1 minor issue identified and fixed
|
||||
- Check 7 (USEFULNESS): ✅ PASS - Strong actionable content
|
||||
|
||||
**Video link pass/fail count:** 17/17 pass, 0 fail
|
||||
|
||||
**Fabricated/non-existent commands:** None found
|
||||
|
||||
**Dead links:** None found
|
||||
|
||||
**Path to REVIEW.md:** /home/hermes/syslog/drafts/scot-hermes-playbook/REVIEW.md
|
||||
@@ -1,12 +0,0 @@
|
||||
@page { size: A4; margin: 2cm 1.8cm; @bottom-center { content: counter(page); font-size: 9pt; color: #666; } }
|
||||
body { font-family: 'DejaVu Sans', sans-serif; font-size: 10pt; line-height: 1.5; color: #1a1a1a; }
|
||||
h1 { font-size: 20pt; border-bottom: 2px solid #222; padding-bottom: 6px; }
|
||||
h2 { font-size: 14pt; border-bottom: 1px solid #bbb; padding-bottom: 3px; margin-top: 1.4em; }
|
||||
h3 { font-size: 11.5pt; margin-top: 1.2em; }
|
||||
code { font-family: 'DejaVu Sans Mono', monospace; font-size: 8.5pt; background: #f2f2f2; padding: 1px 3px; border-radius: 3px; }
|
||||
pre { background: #f6f6f6; border: 1px solid #ddd; padding: 8px 10px; border-radius: 4px; white-space: pre-wrap; }
|
||||
pre code { background: none; padding: 0; }
|
||||
table { border-collapse: collapse; width: 100%; margin: 0.8em 0; font-size: 9pt; }
|
||||
th, td { border: 1px solid #999; padding: 4px 6px; text-align: left; vertical-align: top; }
|
||||
th { background: #eee; }
|
||||
blockquote { border-left: 3px solid #888; margin-left: 0; padding-left: 12px; color: #444; }
|
||||
@@ -1,18 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Render HERMES-PLAYBOOK-FOR-SCOT.md -> PDF (send-ready package for t_2c716052).
|
||||
# Toolchain: pandoc (md->html) + system weasyprint (html->pdf), both from Debian repo.
|
||||
set -euo pipefail
|
||||
DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
SRC="$DIR/HERMES-PLAYBOOK-FOR-SCOT.md"
|
||||
OUT="$DIR/HERMES-PLAYBOOK-FOR-SCOT.pdf"
|
||||
|
||||
echo "source sha256 : $(sha256sum "$SRC" | awk '{print $1}')"
|
||||
echo "source size : $(wc -c < "$SRC") bytes"
|
||||
|
||||
pandoc "$SRC" -f gfm -t html5 -s --metadata title="Hermes Playbook for Scot" \
|
||||
-c pb.css -o /tmp/pb.html
|
||||
weasyprint -u "$DIR/" /tmp/pb.html "$OUT"
|
||||
|
||||
echo "pdf path : $OUT"
|
||||
echo "pdf size : $(wc -c < "$OUT") bytes"
|
||||
echo "pdf sha256 : $(sha256sum "$OUT" | awk '{print $1}')"
|
||||
@@ -1,50 +0,0 @@
|
||||
# 01 — The Context Gap: Claude Code vs a Fresh Hermes Install
|
||||
|
||||
**Audience:** internal research for the Scot Murray playbook (writer takes over from here).
|
||||
**Prepared:** 2026-09-11. Primary sources: official docs (hermes-agent.nousresearch.com/docs) and live CLI verification on the reference install (Syslog kagentz). Every "exact command/file" was checked against `hermes --help` / `hermes <cmd> --help` output on v0.21.1 unless marked DOC-ONLY.
|
||||
|
||||
## Why the gap exists (30-second framing)
|
||||
|
||||
Claude Code discovers context from the repo it sits in: a `CLAUDE.md` it reads on every
|
||||
launch, `.claude/` folders that ship subagents, slash commands, hooks, and skills. A fresh
|
||||
Hermes install starts nearly empty by design — its philosophy is that the agent *builds* its
|
||||
own context over time (memory, skills) and that context comes from config files, not the
|
||||
repo. The "lack of context" Scot noticed is just Hermes waiting to be seeded. Below: every
|
||||
gap and the Hermes mechanism that closes it.
|
||||
|
||||
## Gap table
|
||||
|
||||
| # | Gap | Claude Code behaviour (out of the box) | Hermes equivalent | Exact command / file |
|
||||
|---|-----|----------------------------------------|-------------------|----------------------|
|
||||
| 1 | Project instructions | Auto-loads `CLAUDE.md` from project root; `#` prefix adds memory live; `claude /init` scaffolds it | Auto-injects `AGENTS.md` (and `.cursorrules`) from the working directory + `SOUL.md` persona + persistent memory from the Hermes home. `hermes import-agent claude-code` migrates existing CLAUDE.md content in one shot | File: `AGENTS.md` in the project root (git-tracked). Command: `hermes import-agent claude-code [--dry-run]` — VERIFIED-LIVE |
|
||||
| 2 | Team/personal instruction split | `.claude/rules/*.md` (project) + `~/.claude/rules/*.md` (personal) | Rules via `AGENTS.md` in the repo (team) vs `SOUL.md` + memory in `~/.hermes/` (personal). Config for everything else: `hermes config edit` | Files: `AGENTS.md` (repo), `SOUL.md` (`~/.hermes/`). VERIFIED-LIVE (documented in `--ignore-rules` help text, which names exactly what gets injected) |
|
||||
| 3 | Slash commands | Ships dozens built-in; custom ones in `.claude/commands/<name>.md` | Rich built-in registry (`/help` to list); custom automation goes into skills instead of command files | In-session: `/help`, `/skills`. Doc: https://hermes-agent.nousresearch.com/docs/reference/slash-commands — VERIFIED-LIVE (registry derived from `hermes_cli/commands.py`) |
|
||||
| 4 | Subagents / delegation | `.claude/agents/*.md`, `@agent` mentions, Task tool | Built-in `delegate_task` tool (isolated subagent contexts, parallel batches) plus full-process spawns (`hermes chat -q`, tmux) and the durable Kanban board for multi-profile work | In-session: ask Hermes to delegate; `hermes kanban create ...` for durable tasks. Doc: /docs/user-guide/features/kanban. VERIFIED-LIVE (`hermes kanban --help` shows 40+ verbs incl. `swarm`) |
|
||||
| 5 | Skills (auto-invoked expertise) | `.claude/skills/*.md` markdown guides invoked by natural language match | Same concept, more infrastructure: skills auto-load by task match, can be authored BY the agent itself (`skill_manage`), installed from registries, maintained by the curator | CLI: `hermes skills list/search/install/config`; in-session: `/skill <name>`, `/reload-skills`. VERIFIED-LIVE. Hub: `hermes skills browse` |
|
||||
| 6 | Project memory / auto-memory | `~/.claude/projects/<project>/memory/`, 25 KB cap | Persistent memory is first-class: built-in `MEMORY.md`/`USER.md` always active, pluggable providers (Honcho, Mem0, …) | CLI: `hermes memory setup/status/off`. VERIFIED-LIVE. Doc: /docs/user-guide/features/memory |
|
||||
| 7 | Tool permissions | `/permissions`, `settings.json` allowlists | Per-platform toolset toggles + MCP tool allowlists (`server:tool` notation) | CLI: `hermes tools` (interactive UI), `hermes tools list/enable/disable`. VERIFIED-LIVE |
|
||||
| 8 | MCP servers | `claude mcp add/list/remove`, scopes user/local/project | `hermes mcp add/list/test/configure`, one-click catalog installs, plus `hermes mcp serve` (Hermes AS an MCP server — Claude Code cannot do this) | VERIFIED-LIVE. Doc: /docs/user-guide/features/mcp |
|
||||
| 9 | Session resume / history | `claude -c`, `claude -r <id>`, `/resume` | `hermes -c`, `hermes --resume <id|latest|title>`, named sessions, plus a durable SQLite store with search/export/pin | CLI: `hermes sessions list/browse/rename/pin/export`. VERIFIED-LIVE |
|
||||
| 10 | Cost & context visibility | `/cost`, `/context` grid | `/usage`, `/insights [days]`, `/compress` (auto-compression built in), `/prompt-size` byte breakdown | VERIFIED-LIVE (`insights`, `logs` subcommands confirmed in `hermes --help`) |
|
||||
| 11 | Headless/CI mode | `claude -p` print mode | `-z/--oneshot` flag (prints only final response) + `hermes chat -q` | VERIFIED-LIVE |
|
||||
| 12 | Import of prior setup | n/a (it IS the incumbent) | **The single most important one for Scot:** `hermes import-agent claude-code` maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, and memories into Hermes equivalents (never API keys) | `hermes import-agent claude-code --dry-run` then without `--dry-run`. VERIFIED-LIVE. Also `hermes sessions import` for old Claude Code conversations — VERIFIED-LIVE |
|
||||
| 13 | Hooks on tool events | 8 hook types in `settings.json` (PreToolUse, PostToolUse, …) | Shell-script hooks managed via `hermes hooks` | CLI: `hermes hooks`. VERIFIED-LIVE (in top-level command list) |
|
||||
| 14 | Scheduled / recurring work | `claude /loop` (in-session only) | Durable cron scheduler with multi-platform delivery, chained outputs (`context_from`), per-job model overrides | CLI: `hermes cron list/create/edit/pause/resume/run/remove/doctor`. VERIFIED-LIVE. Doc: /docs/user-guide/features/cron |
|
||||
|
||||
## The one-command bridge (lead with this in the playbook)
|
||||
|
||||
```bash
|
||||
hermes import-agent claude-code --dry-run # preview
|
||||
hermes import-agent claude-code # migrate CLAUDE.md → AGENTS.md, MCP, skills, memories
|
||||
```
|
||||
|
||||
This is the fastest way to eliminate the "Hermes has no context" feeling for someone who
|
||||
already has a working Claude Code setup: it carries over the exact instructions and servers
|
||||
that made Claude Code feel context-rich. Preview-only mode exists (`--dry-run`), it never
|
||||
imports credentials, and conflicts are skipped by default (`--overwrite` to change).
|
||||
|
||||
## Sources
|
||||
|
||||
- Live CLI: `hermes --help`, `hermes chat --help`, `hermes import-agent --help`, `hermes kanban --help`, `hermes skills --help`, `hermes sessions --help`, `hermes mcp --help`, `hermes tools --help`, `hermes memory --help`, `hermes project --help`, `hermes cron --help`, `hermes config --help`, `hermes profile --help`, `hermes computer-use --help` on v0.21.1, reference install (Syslog kagentz), 2026-09-11. Raw dump: `cli-help-dump.txt` next to this file.
|
||||
- Docs: https://hermes-agent.nousresearch.com/docs/ (index) — all URLs in sources.md
|
||||
- Claude Code side: installed skill `delegate-coding-agent/references/claude-code.md` (Hermes Agent + Teknium, v2.2.1), `/home/hermes/.hermes/skills/autonomous-ai-agents/`
|
||||
@@ -1,101 +0,0 @@
|
||||
# 02 — High-Leverage Hermes Surfaces (the "harness power" inventory)
|
||||
|
||||
**Prepared:** 2026-09-11. Each surface: what it does, when to use it, exact command/file, doc URL. Verification: V-LIVE = confirmed against live CLI v0.21.1 on the reference install (Syslog kagentz); V-DOC = confirmed against official docs page (URL resolved HTTP 200); V-FILE = present on this machine's installed skills.
|
||||
|
||||
---
|
||||
|
||||
### 1. Persona / SOUL file
|
||||
- **What:** `SOUL.md` is Hermes' personality + standing-identity file, auto-injected into the system prompt alongside `AGENTS.md` rules and memory (confirmed by `--ignore-rules` help text which lists exactly what gets injected).
|
||||
- **When:** client wants the agent to have a consistent voice/role (e.g., "you are my analyst").
|
||||
- **Where:** `~/.hermes/SOUL.md` (per-profile: `~/.hermes/profiles/<name>/SOUL.md`).
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/configuration [V-DOC]
|
||||
|
||||
### 2. Persistent memory (built-in + providers)
|
||||
- **What:** Built-in `MEMORY.md` / `USER.md` always active; optional external providers (honcho, mem0, hindsight, byterover, …). Memory is injected every session — this is the single biggest cure for "it forgets my project."
|
||||
- **When:** after any correction or preference the user states ("use bun, not npm") — tell Hermes to remember it and it persists.
|
||||
- **Command:** `hermes memory setup|status|off|reset` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/memory [V-DOC]
|
||||
|
||||
### 3. Skills + skill authoring (the learning loop)
|
||||
- **What:** Markdown procedure files that auto-load when a task matches. The differentiator: Hermes can WRITE its own skills after learning a workflow (self-improving), and the curator maintains them (usage tracking, archiving, backups).
|
||||
- **When:** any workflow done twice — say "save this as a skill."
|
||||
- **Commands:** `hermes skills list|search|install|browse|config|check|update` [V-LIVE]; in-session `/skill <name>`, `/reload-skills` [V-DOC]; authoring tool in-session is `skill_manage` (agent-side; writer should describe it as "ask your Hermes to save the procedure as a skill").
|
||||
- **Docs:** https://hermes-agent.nousresearch.com/docs/reference/skills-catalog [V-DOC]; curator: https://hermes-agent.nousresearch.com/docs/user-guide/features/curator [V-DOC]
|
||||
|
||||
### 4. Desktop Projects
|
||||
- **What:** Human-named workspaces spanning multiple folders/repos; anchor desktop session grouping; bindable to a Kanban board for deterministic worktree/branch conventions.
|
||||
- **When:** Scot's multi-repo workflows (portfolio ops). `hermes project create <name>` then `add-folder`.
|
||||
- **Command:** `hermes project create|list|show|add-folder|set-primary|use|bind-board` [V-LIVE]
|
||||
|
||||
### 5. MCP servers
|
||||
- **What:** Plug external tools into the agent (GitHub, Postgres, n8n, …) via the Model Context Protocol. Also runs in reverse: `hermes mcp serve` exposes Hermes conversations to other agents.
|
||||
- **Command:** `hermes mcp add|list|test|configure|picker|catalog|install|serve` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp [V-DOC]
|
||||
|
||||
### 6. Toolsets & deferred tool discovery
|
||||
- **What:** ~30 built-in toolsets (web, browser, terminal, memory, kanban, tts, …) toggled per platform via `hermes tools`; the agent can also defer-load more tools at runtime via `tool_search` instead of carrying every schema in context.
|
||||
- **When:** trim toolsets for focus/cost, or enable `browser` for web work.
|
||||
- **Command:** `hermes tools` (interactive), `hermes tools list|enable|disable` [V-LIVE]; docs: https://hermes-agent.nousresearch.com/docs/reference/tools-reference [V-DOC]
|
||||
|
||||
### 7. Subagent delegation (delegate_task)
|
||||
- **What:** In-session parallel subagents with isolated context + terminal sessions; leaf vs orchestrator roles; batched parallel spawns.
|
||||
- **When:** research fan-out, parallel code review, anything that would flood the main context.
|
||||
- **Command:** agent-side tool (no CLI). In-session: ask Hermes to "delegate X to subagents." Docs: /docs/user-guide/features (delegation section) [V-DOC]
|
||||
|
||||
### 8. Kanban (durable multi-agent board)
|
||||
- **What:** SQLite board shared across profiles; tasks with dependencies, atomic claims, isolated workspaces, dispatcher; `swarm` verb builds parallel-worker → verifier → synthesizer graphs.
|
||||
- **When:** recurring multi-step operations, handoffs between specialist profiles, long-running campaigns that must survive restarts.
|
||||
- **Command:** `hermes kanban create|list|show|swarm|link|complete|watch|stats|dispatch` (40+ verbs) [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/kanban [V-DOC]
|
||||
|
||||
### 9. Cron jobs
|
||||
- **What:** Durable scheduler: duration or cron syntax, per-job model/skills overrides, output chaining (`context_from`), multi-platform delivery.
|
||||
- **When:** daily reports, monitoring with alerts, weekly reviews.
|
||||
- **Command:** `hermes cron list|create|edit|pause|resume|run|remove|doctor|status` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/features/cron [V-DOC]
|
||||
|
||||
### 10. Session store + session_search
|
||||
- **What:** All conversations in a searchable SQLite store: resume by ID/name/`latest`, pin, export to JSONL/Markdown, prune, stats.
|
||||
- **When:** "what did we decide last week" — the agent can search past sessions; user can browse them.
|
||||
- **Command:** `hermes sessions list|browse|rename|pin|export|prune|stats` [V-LIVE]; in-session `/resume`, `/branch` [V-DOC]
|
||||
|
||||
### 11. Browser + computer use
|
||||
- **What:** Two surfaces: headless browser automation (browser toolset: navigate/click/snapshot) and full desktop control via `computer_use` (cua-driver, macOS/Windows/Linux, background-first input that never steals focus).
|
||||
- **When:** web research → headless browser; native apps (Excel, Figma, native chat) → computer use.
|
||||
- **Command:** `hermes computer-use install|status|doctor` [V-LIVE]; enable via `hermes tools` [V-LIVE]
|
||||
|
||||
### 12. Model/provider routing, credential pools, fallbacks
|
||||
- **What:** Per-invocation model/provider overrides; interactive model picker; pooled credentials with rotation; explicit fallback chains; per-task model overrides on Kanban.
|
||||
- **Command:** `hermes model` [V-LIVE], `hermes fallback list|add|remove` [V-LIVE], `hermes auth add|list|priority|reset` [V-LIVE]; per-run flags `-m`, `--provider`, `--reasoning` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/integrations/providers [V-DOC]
|
||||
- **Note for Scot (OpenRouter + small fast models):** `--reasoning high` on hard tasks; `hermes fallback add` so a failed call rolls to a second model instead of erroring.
|
||||
|
||||
### 13. Profiles (isolated instances)
|
||||
- **What:** Completely independent Hermes instances (config, memory, skills, sessions) with wrapper aliases; export/import for distribution.
|
||||
- **When:** separate work/persona contexts, or one profile per client.
|
||||
- **Command:** `hermes profile list|create|use|alias|export|import` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/profiles [V-DOC]
|
||||
|
||||
### 14. Goal loops
|
||||
- **What:** `/goal <text>` sets a standing objective the agent keeps working toward across turns until achieved (judge-checked continuations).
|
||||
- **When:** "keep the CI green until it passes," "keep researching until you have 5 verified sources."
|
||||
- **Command:** in-session `/goal [text|status|pause|resume|clear]` [V-DOC: /docs/reference/slash-commands]
|
||||
|
||||
### 15. Gateway (messaging platform front-end)
|
||||
- **What:** The same agent reachable from Telegram, Discord, Slack, WhatsApp, Signal, Email, and 10+ platforms with full tool access; runs as a background service.
|
||||
- **Command:** `hermes gateway run|install|start|status|setup` [V-LIVE]
|
||||
- **Doc:** https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ [V-DOC]
|
||||
|
||||
### 16. Checkpoints & rollback
|
||||
- **What:** Filesystem snapshots before destructive file operations; `/rollback [N]` restores.
|
||||
- **When:** letting the agent loose on important files.
|
||||
- **Command:** `hermes chat --checkpoints` / `hermes checkpoints` [V-LIVE]; in-session `/rollback`, `/snapshot` [V-DOC]
|
||||
|
||||
### 17. Projects↔Kanban binding + worktree mode
|
||||
- **What:** `hermes project bind-board` ties a board to a project (deterministic worktree + branch per task); `-w/--worktree` runs any session in an isolated git worktree.
|
||||
- **When:** parallel coding agents that must not collide.
|
||||
- **Command:** `hermes project bind-board` [V-LIVE]; `hermes -w` [V-LIVE]
|
||||
|
||||
### 18. Prompt-size introspection
|
||||
- **What:** Byte breakdown of system prompt + tool schemas — diagnose why responses feel "dumb" (usually context bloat).
|
||||
- **Command:** `hermes prompt-size` [V-LIVE]
|
||||
@@ -1,121 +0,0 @@
|
||||
# 03 — Command Cheatsheet (every entry verified)
|
||||
|
||||
**Verification method:** each VERIFIED-LIVE entry was confirmed against `hermes --help` or `hermes <cmd> --help` on Hermes Agent v0.21.1 (2026.9.7), reference install (Syslog kagentz), 2026-09-11. Raw output: `cli-help-dump.txt`. DOC-ONLY entries come from the official docs (URL given). Nothing is invented.
|
||||
|
||||
## (a) CLI — `hermes ...`
|
||||
|
||||
### Setup & health
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes setup` | Interactive setup wizard | VERIFIED-LIVE |
|
||||
| `hermes doctor [--fix] [--live]` | Diagnose config/deps; `--fix` auto-repairs | VERIFIED-LIVE |
|
||||
| `hermes status [--all] [--deep]` | Component status | VERIFIED-LIVE |
|
||||
| `hermes config show/edit/get/set/unset/path/env-path/check/migrate` | View/edit config | VERIFIED-LIVE |
|
||||
| `hermes update` | Update Hermes to latest | VERIFIED-LIVE |
|
||||
|
||||
### The Claude Code bridge (highest value for Scot)
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes import-agent claude-code [--dry-run] [--overwrite] [--yes]` | One-command import of a Claude Code setup: maps CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills, memories into Hermes equivalents. Never imports API keys. | VERIFIED-LIVE |
|
||||
| `hermes import-agent codex` | Same for Codex CLI setups | VERIFIED-LIVE |
|
||||
| `hermes sessions import` | Import a Claude Code or Codex CLI **session/conversation** into Hermes | VERIFIED-LIVE (subcommand listed in `hermes sessions --help`) |
|
||||
| `hermes skills trust` | Trust a repo so its project-local skills (`./.hermes/skills`) load — the Hermes analog of `.claude/skills/` | VERIFIED-LIVE |
|
||||
|
||||
### Daily driving
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes` / `hermes chat` | Interactive session | VERIFIED-LIVE |
|
||||
| `hermes -c [NAME]` / `hermes --resume <id\|latest>` | Resume by name or ID | VERIFIED-LIVE |
|
||||
| `hermes --in DIR --resume latest` | Resume the latest session for a directory | VERIFIED-LIVE |
|
||||
| `hermes -z "PROMPT"` | One-shot: prints ONLY the final answer (scripting/CI); tools, memory, and AGENTS.md still load | VERIFIED-LIVE |
|
||||
| `hermes chat -q "PROMPT"` | Single-query mode | VERIFIED-LIVE |
|
||||
| `hermes -m MODEL --provider PROVIDER --reasoning LEVEL` | Per-run model/provider/reasoning overrides (`none…ultra`) | VERIFIED-LIVE |
|
||||
| `hermes -s SKILL1,SKILL2` | Preload specific skills for the session | VERIFIED-LIVE |
|
||||
| `hermes -t TOOLSETS` | Restrict toolsets for this run | VERIFIED-LIVE |
|
||||
| `hermes -w` | Isolated git worktree session (parallel agents on one repo) | VERIFIED-LIVE |
|
||||
| `hermes chat --checkpoints` | Enable filesystem checkpoints (`/rollback` to restore) | VERIFIED-LIVE |
|
||||
| `hermes chat --max-turns N` / `--run-budget SECONDS` | Cap loop iterations / wall-clock budget | VERIFIED-LIVE |
|
||||
|
||||
### Context & memory management
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes memory setup/status/off/reset` | External memory provider management (built-in MEMORY.md/USER.md always active) | VERIFIED-LIVE |
|
||||
| `hermes sessions list/browse/rename/pin/export/prune/stats` | Session store management | VERIFIED-LIVE |
|
||||
| `hermes skills list/search/install/inspect/browse/config/check/update` | Skill management | VERIFIED-LIVE |
|
||||
| `hermes skills trust/untrust` | Repo-local skill trust | VERIFIED-LIVE |
|
||||
| `hermes curator status/run/pause/pin/...` | Background skill maintenance (auto-archive, backups) | VERIFIED-LIVE |
|
||||
| `hermes prompt-size` | Byte breakdown of system prompt + tool schemas (context-bloat diagnosis) | VERIFIED-LIVE |
|
||||
| `hermes insights [--days N]` | Usage analytics | VERIFIED-LIVE |
|
||||
|
||||
### Tools, MCP, integrations
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes tools` (interactive) / `list/enable/disable` | Per-platform toolset toggles; MCP tools as `server:tool` | VERIFIED-LIVE |
|
||||
| `hermes mcp add/remove/list/test/configure/picker/catalog/install` | MCP server management (incl. one-click catalog installs) | VERIFIED-LIVE |
|
||||
| `hermes mcp serve` | Run Hermes AS an MCP server for other agents | VERIFIED-LIVE |
|
||||
| `hermes computer-use install/status/doctor` | Desktop-control backend (cua-driver) | VERIFIED-LIVE |
|
||||
| `hermes gateway run/install/start/status/setup` | Messaging gateway (Telegram, Discord, Slack, WhatsApp, …) | VERIFIED-LIVE |
|
||||
| `hermes send` | Send a message to a configured platform (scripts/cron/CI) | VERIFIED-LIVE |
|
||||
|
||||
### Automation & multi-agent
|
||||
| Command | What it does | Tag |
|
||||
|---|---|---|
|
||||
| `hermes cron list/create/edit/pause/resume/run/remove/doctor` | Scheduled jobs (durable, multi-platform delivery) | VERIFIED-LIVE |
|
||||
| `hermes cron notepad` | Durable per-job key-value notepad across runs | VERIFIED-LIVE |
|
||||
| `hermes kanban create/list/show/link/complete/swarm/...` | Durable multi-profile task board (40+ verbs) | VERIFIED-LIVE |
|
||||
| `hermes kanban swarm` | Generate a parallel-workers → verifier → synthesizer task graph | VERIFIED-LIVE |
|
||||
| `hermes project create/list/add-folder/bind-board` | Named multi-folder workspaces (desktop Projects) | VERIFIED-LIVE |
|
||||
| `hermes profile list/create/use/alias/export/import` | Isolated Hermes instances | VERIFIED-LIVE |
|
||||
| `hermes auth add/list/priority/reset` | Pooled credentials per provider (rotation) | VERIFIED-LIVE |
|
||||
| `hermes fallback list/add/remove` | Fallback model chain (auto-rollover on failure) | VERIFIED-LIVE |
|
||||
| `hermes model` | Interactive model/provider picker | VERIFIED-LIVE |
|
||||
| `hermes -yolo` | Bypass command approval prompts (use with care) | VERIFIED-LIVE |
|
||||
| `hermes pause` / `hermes resume` | Emergency stop / lift (pauses cron, kanban dispatch, gateway turns) | VERIFIED-LIVE |
|
||||
|
||||
## (b) In-session slash commands
|
||||
|
||||
Source: official slash-commands reference https://hermes-agent.nousresearch.com/docs/reference/slash-commands (DOC-ONLY — slash commands run inside a chat session and were not exercised from this headless research run; the CLI subcommands they map to were verified live). DOC-ONLY.
|
||||
|
||||
### Context & session
|
||||
| Command | What it does |
|
||||
|---|---|
|
||||
| `/help` | List all commands (authoritative in your version) |
|
||||
| `/new` (`/reset`) | Fresh session |
|
||||
| `/resume [name]` | Resume a named/recent session |
|
||||
| `/branch` (`/fork`) | Branch the current session |
|
||||
| `/compress` | Manually compress context (auto-compression also exists) |
|
||||
| `/undo` | Remove last exchange |
|
||||
| `/retry` | Resend last message |
|
||||
| `/title [name]` | Name the session |
|
||||
| `/save` | Save conversation to file |
|
||||
| `/history` | Show conversation history |
|
||||
|
||||
### Power surfaces
|
||||
| Command | What it does |
|
||||
|---|---|
|
||||
| `/skill <name>` | Load a skill into the session |
|
||||
| `/skills` | Search/install skills |
|
||||
| `/reload-skills` | Re-scan skill directory |
|
||||
| `/tools` / `/toolsets` | Manage tools |
|
||||
| `/goal [text]` | Set a standing goal the agent works toward across turns (`/goal status/pause/clear` to manage) |
|
||||
| `/background <prompt>` | Run a prompt in the background |
|
||||
| `/queue <prompt>` | Queue a prompt for the next turn |
|
||||
| `/steer <prompt>` | Inject a course-correction after the next tool call without interrupting |
|
||||
| `/agents` | Show active agents and running tasks |
|
||||
| `/cron` | Manage cron jobs in-session |
|
||||
| `/kanban` | Multi-profile collaboration board in-session |
|
||||
| `/model [name]` | Show/change model mid-session |
|
||||
| `/reasoning [level]` | Set reasoning effort |
|
||||
| `/voice [on\|off\|tts]` | Voice mode |
|
||||
| `/rollback [N]` | Restore filesystem checkpoint (needs `--checkpoints`) |
|
||||
| `/usage` | Token usage |
|
||||
| `/insights [days]` | Usage analytics |
|
||||
| `/platforms` | Gateway platform status |
|
||||
| `/compact`-equivalent note | Hermes compresses automatically near the context limit; no manual threshold watch needed like Claude Code's `/context` |
|
||||
|
||||
### The "work alongside Claude Code" shortlist
|
||||
1. `hermes import-agent claude-code --dry-run` → migrate the setup (VERIFIED-LIVE)
|
||||
2. `hermes sessions import` → bring the conversation history over (VERIFIED-LIVE)
|
||||
3. `hermes skills trust` → load repo-local skills like `.claude/skills/` (VERIFIED-LIVE)
|
||||
4. `hermes -c` / `hermes --in <repo> --resume latest` → per-directory session continuity (VERIFIED-LIVE)
|
||||
5. `hermes mcp serve` → expose Hermes to Claude Code as an MCP server (VERIFIED-LIVE) — the reverse direction Claude Code can't do
|
||||
@@ -1,76 +0,0 @@
|
||||
# 04 — The Claude Code Bridge: Running Hermes WITH Claude Code
|
||||
|
||||
**Prepared:** 2026-09-11. Scot already runs both tools. This file documents the proven integration patterns, citing the installed skills on this host (paths under `/home/hermes/.hermes/skills/`) and official docs.
|
||||
|
||||
## Pattern 0 — Import (do this first)
|
||||
`hermes import-agent claude-code` [VERIFIED-LIVE] maps CLAUDE.md/AGENTS.md instructions,
|
||||
permission allowlists, MCP servers, skills, and memories into Hermes equivalents. It always
|
||||
shows a preview, never imports credentials. `hermes sessions import` [VERIFIED-LIVE] pulls
|
||||
in old Claude Code conversations. After import, Hermes "knows" the projects — the context
|
||||
gap disappears on day one.
|
||||
|
||||
## Pattern 1 — Hermes as orchestrator, Claude Code as worker
|
||||
Source: installed skill **`autonomous-ai-agents/delegate-coding-agent`** (v1.0.0) + its
|
||||
reference `references/claude-code.md` (v2.2.1) [V-FILE]. The skill is an official Hermes
|
||||
skill authored for exactly this.
|
||||
|
||||
Two orchestration modes (verbatim from the skill):
|
||||
- **Print mode (preferred):** `claude -p '<task>' --allowedTools 'Read,Edit' --max-turns 10` —
|
||||
one-shot, no dialogs, structured JSON output with `session_id`, `num_turns`,
|
||||
`total_cost_usd`. Ask Hermes: *"delegate this coding task to Claude Code in print mode."*
|
||||
- **Interactive PTY via tmux:** multi-turn sessions — Hermes starts `tmux new-session`,
|
||||
sends prompts with `send-keys`, monitors with `capture-pane`. For iterative
|
||||
refactor → review → fix cycles.
|
||||
|
||||
Cross-agent review loop (also from the skill):
|
||||
```
|
||||
git diff main...feature | claude -p 'Review this diff for bugs and security issues.' --max-turns 1
|
||||
```
|
||||
Hermes runs this, reads the output, and fixes findings itself — Claude Code becomes a
|
||||
reviewer Hermes coordinates.
|
||||
|
||||
Safety rails the skill prescribes: explicit `workdir`, clean git status before launch,
|
||||
narrow task prompts, `git diff` review, targeted tests before committing.
|
||||
|
||||
## Pattern 2 — Parallel workstreams + neutral merge reconciliation
|
||||
Source: installed skill **`autonomous-ai-agents/merge-reconciler`** [V-FILE].
|
||||
When Hermes and Claude Code (or two Hermes workers) both edit the same repo and collide:
|
||||
- Do NOT let either agent resolve the conflict — both are biased toward their own side.
|
||||
- Spawn a **neutral third agent** with the merge-reconciler skill; it classifies every
|
||||
conflicted hunk (disjoint-intent / same-question-different-answer / superseded), resolves
|
||||
under an impartiality contract (touch only conflict markers, surface every design call),
|
||||
verifies with build/tests, and hands back a summary naming every hunk decision.
|
||||
- Kanban-native shape: a reconciliation card assigned to a **third profile** with both
|
||||
workers' cards as parents — parent links carry both sides' completion summaries into the
|
||||
reconciler's context automatically.
|
||||
|
||||
## Pattern 3 — Hermes as MCP server (Claude Code gets Hermes tools)
|
||||
`hermes mcp serve` [VERIFIED-LIVE] runs Hermes as an MCP server exposing its conversations
|
||||
and capabilities. Claude Code supports MCP clients (`claude mcp add`), so Claude Code can
|
||||
consume Hermes as a tool provider — persistent memory, skills, cron — the surfaces Claude
|
||||
Code lacks. This is the reverse-bridge only Hermes can offer. Docs:
|
||||
https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp
|
||||
|
||||
## Pattern 4 — Import legacy sessions for continuity
|
||||
`hermes sessions import` [VERIFIED-LIVE] imports a Claude Code session into the Hermes
|
||||
store; from then on `hermes --resume <id>` / `hermes sessions browse` treat it as native
|
||||
history. Use when mid-project: the new agent picks up exactly where Claude Code left off.
|
||||
|
||||
## Pattern 5 — Desktop GUI automation either side can use
|
||||
Source: installed skill **`autonomous-ai-agents/computer-use`** (v2.0.0) [V-FILE].
|
||||
`hermes computer-use install` sets up cua-driver; the `computer_use` toolset drives native
|
||||
desktop apps background-first (never steals focus/cursor), any-model, cross-platform.
|
||||
Relevant to the bridge because Claude Code has no desktop automation — if a task needs
|
||||
Figma/Excel/native apps, that part routes to Hermes while the code routes to Claude Code.
|
||||
Cmd: `hermes computer-use doctor` for health checks.
|
||||
|
||||
## Pattern 6 — The import-agent philosophy in one line
|
||||
Claude Code holds repo context in `CLAUDE.md`; Hermes holds it in `AGENTS.md` + memory +
|
||||
skills. `hermes import-agent claude-code` translates the first; the learning loop
|
||||
("save this as a skill") rebuilds the rest automatically the more Scot uses Hermes.
|
||||
|
||||
## Reference paths (for the writer)
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/delegate-coding-agent/SKILL.md` and `references/claude-code.md`
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/merge-reconciler/SKILL.md`
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/computer-use/SKILL.md`
|
||||
- `hermes import-agent --help` raw output in `cli-help-dump.txt` (lines 554-577)
|
||||
@@ -1,119 +0,0 @@
|
||||
# 05 — The Video Watch List (every URL verified 2026-09-11)
|
||||
|
||||
**Verification method:** each video was found via YouTube search-results scrape (`videoRenderer` metadata), then confirmed with the YouTube oEmbed endpoint (`curl -s "https://www.youtube.com/oembed?url=<URL>&format=json"`) — every entry below returned HTTP 200 with matching title/author (status PASS). Publish dates, durations, and view counts were read from each watch page's metadata. Raw evidence for all 21 entries: `video-verification.json` in this directory. **21/21 PASS, 0 FAIL.**
|
||||
|
||||
## Tier 1 — Hermes-specific, start here
|
||||
|
||||
### 1. Learn 95% of Hermes Agent in 31 Minutes
|
||||
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=Ta2wg6xPaY4 | **Duration:** 31:28 | **Published:** 2026-08-09 | **Views:** ~129k
|
||||
- **oEmbed:** PASS (title/author match)
|
||||
- **What it demonstrates:** end-to-end Hermes fundamentals — install, sessions, skills, memory, the learning loop. The most complete single-video orientation found.
|
||||
- **Watch this when you want** the fastest real overview of the whole harness before touching config.
|
||||
|
||||
### 2. Hermes Agent Fundamentals In 29 Minutes
|
||||
- **Channel:** Tina Huang | **URL:** https://www.youtube.com/watch?v=5_N84t1rUU0 | **Duration:** 29:40 | **Published:** 2026-07-20 | **Views:** ~463k (highest-reach Hermes video found)
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** conceptual grounding — why Hermes' memory/skills loop differs from one-shot coding agents; practical walkthrough.
|
||||
- **Watch this when you want** to understand *why* Hermes feels different from Claude Code, not just which buttons to press.
|
||||
|
||||
### 3. Every Level of Hermes Agent Explained
|
||||
- **Channel:** Jack Roberts | **URL:** https://www.youtube.com/watch?v=6GtF_uHbGhw | **Duration:** 25:35 | **Published:** 2026-06-17 | **Views:** ~163k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** beginner → advanced ladder of features (memory, skills, automation, multi-agent).
|
||||
- **Watch this when you want** a map of what to learn next after the basics.
|
||||
|
||||
### 4. Hermes Agent Full Tutorial INSTALLATION + USECASES
|
||||
- **Channel:** CodeHead | **URL:** https://www.youtube.com/watch?v=8GjyOQy19so | **Duration:** 7:47 | **Published:** 2026-05-14 | **Views:** ~64k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** install through real use-cases, compact.
|
||||
- **Watch this when you want** a quick install-to-value demo to share with a colleague.
|
||||
|
||||
### 5. Hermes Agent Explained In 5 Minutes
|
||||
- **Channel:** CodeHead | **URL:** https://www.youtube.com/watch?v=9GpWELm3_XI | **Duration:** 4:53 | **Published:** 2026-05-23 | **Views:** ~251k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** 5-minute conceptual pitch of the agent and its learning loop.
|
||||
- **Watch this when you want** the elevator pitch before committing 30 minutes.
|
||||
|
||||
## Tier 2 — Hermes-specific deep dives
|
||||
|
||||
### 6. 100 Days With Hermes Agent in 21 Minutes
|
||||
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=sCa3BtpkziQ | **Duration:** 21:19 | **Published:** 2026-06-17 | **Views:** ~58k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** long-horizon usage — what memory/skills accumulation actually looks like after months of daily use.
|
||||
- **Watch this when you want** to see the payoff of the learning loop over time.
|
||||
|
||||
### 7. Hermes Agent - Crash Course for Beginners (AI Agent)
|
||||
- **Channel:** Adrian Twarog | **URL:** https://www.youtube.com/watch?v=4sAmpcSOVEw | **Duration:** 22:19 | **Published:** 2026-07-21 | **Views:** ~46k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** beginner crash course from a well-known dev-YouTube creator.
|
||||
- **Watch this when you want** a second independent explanation of the basics.
|
||||
|
||||
### 8. Hermes Agent: The Ultimate Beginner's Guide
|
||||
- **Channel:** Metics Media | **URL:** https://www.youtube.com/watch?v=CwPUOVUdApE | **Duration:** 37:08 | **Published:** 2026-04-24 | **Views:** ~119k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** long-form beginner guide incl. setup and everyday workflows.
|
||||
- **Watch this when you want** the most thorough single walkthrough in one sitting.
|
||||
|
||||
### 9. Hermes Agent Just Killed OpenClaw (Full Tutorial)
|
||||
- **Channel:** Leon van Zyl | **URL:** https://www.youtube.com/watch?v=jmtpYUOr7_U | **Duration:** 19:59 | **Published:** 2026-04-28 | **Views:** ~16k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** full tutorial framing Hermes against the OpenClaw workflow (MCP config, memory, agents).
|
||||
- **Watch this when you want** a practitioner's feature-by-feature tutorial.
|
||||
|
||||
### 10. Hermes Agent vs OpenClaw
|
||||
- **Channel:** Sharbel A. | **URL:** https://www.youtube.com/watch?v=zwqhemjHq3E | **Duration:** 15:28 | **Published:** 2026-04-20 | **Views:** ~35k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** head-to-head comparison of the two agent harnesses.
|
||||
- **Watch this when you want** the tradeoffs between Hermes and its main alternative.
|
||||
|
||||
### 11. Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model
|
||||
- **Channel:** Tonbi's AI Garage | **URL:** https://www.youtube.com/watch?v=8tpuky8HpXw | **Duration:** 15:08 | **Published:** 2026-03-11 | **Views:** ~18k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** Hermes driven by an OpenRouter-served open model (Qwen 3) — directly relevant to an OpenRouter-connected install.
|
||||
- **Watch this when you want** to see how small open models behave inside Hermes.
|
||||
|
||||
### 12. Use This To Make The Hermes Agent Basically Free
|
||||
- **Channel:** AI LABS | **URL:** https://www.youtube.com/watch?v=5d02TYoOzfE | **Duration:** 13:08 | **Published:** 2026-07-01 | **Views:** ~55k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** running Hermes on cheap/free model backends.
|
||||
- **Watch this when you want** to cut inference costs on an OpenRouter account.
|
||||
|
||||
### 13. Hermes Agent The 24/7 Self-Evolving AI Agent!
|
||||
- **Channel:** WorldofAI | **URL:** https://www.youtube.com/watch?v=cu2fgknmemA | **Duration:** 9:15 | **Published:** 2026-04-07 | **Views:** ~47k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** always-on operation: gateway, cron, background automation.
|
||||
- **Watch this when you want** to turn Hermes from a chat window into a 24/7 assistant.
|
||||
|
||||
## Tier 3 — Adjacent (origin/philosophy; not tutorials)
|
||||
|
||||
### 14. Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra
|
||||
- **Channel:** Peter Yang | **URL:** https://www.youtube.com/watch?v=UWjh5Z4s8jY | **Duration:** 46:45 | **Published:** 2026-08-02 | **Views:** ~37k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** interview with Hermes' co-founder on the design philosophy (self-improving agents, skills as memory).
|
||||
- **Watch this when you want** to understand where the product is going.
|
||||
|
||||
### 15. Hermes Agent: Agents that grow with you | Episode #357
|
||||
- **Channel:** Practical AI | **URL:** https://www.youtube.com/watch?v=UTZhvPXnmwA | **Duration:** 47:34 | **Published:** 2026-05-20 | **Views:** ~1.9k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** podcast-depth technical discussion of the agent architecture.
|
||||
- **Watch this when you want** the engineering story behind the learning loop.
|
||||
|
||||
### 16. Did Hermes Agent just kill OpenClaw? (full guide)
|
||||
- **Channel:** Alex Finn | **URL:** https://www.youtube.com/watch?v=tP6yf22OJdI | **Duration:** 13:55 | **Published:** 2026-03-31 | **Views:** ~132k
|
||||
- **oEmbed:** PASS
|
||||
- **What it demonstrates:** guide-style comparison/switch content.
|
||||
- **Watch this when you want** a switcher's guide perspective.
|
||||
|
||||
### 17. Hermes Agent: Why Everyone's Ditching OpenClaw in 2026
|
||||
- **Channel:** Luke Alexander AI | **URL:** https://www.youtube.com/watch?v=1UgXUjT-QtI | **Duration:** 18:03 | **Published:** 2026-03-26 | **Views:** ~14k
|
||||
- **oEmbed:** PASS — adjacent, comparison content.
|
||||
- **Watch this when you want** more comparison context.
|
||||
|
||||
## Honesty note (required by task spec)
|
||||
At least 16 of the 17 entries above are directly Hermes-specific (not merely adjacent); the
|
||||
"fewer than 5 exist" fallback clause was NOT needed — no padding was necessary. Entries
|
||||
found in search but excluded as thin/low-signal: `iqN6MVzpJTk` (3.1k views, news-style),
|
||||
`83nWNRKZTCE` (465 views), `6M2tItdARew` (1.2k views), `P2LIFtrRr2U` (promo-style) — all
|
||||
also verified PASS and kept in `video-verification.json` as spares. No official Nous
|
||||
Research YouTube tutorial channel was found in searches; the strongest signal of
|
||||
Hermes-specific video content is the third-party ecosystem above.
|
||||
@@ -1,577 +0,0 @@
|
||||
===== hermes chat --help =====
|
||||
usage: hermes chat [-h] [-q QUERY | --query-file PATH] [--oneshot]
|
||||
[--image IMAGE] [-m MODEL] [-t TOOLSETS]
|
||||
[--reasoning LEVEL] [-s SKILLS] [--provider PROVIDER] [-v]
|
||||
[-Q] [--resume SESSION_ID] [--no-restore-cwd] [--in DIR]
|
||||
[--continue [SESSION_NAME]] [--create-if-missing]
|
||||
[--worktree] [--accept-hooks] [--checkpoints]
|
||||
[--max-turns N] [--run-budget SECONDS] [--yolo]
|
||||
[--pass-session-id] [--ignore-user-config] [--ignore-rules]
|
||||
[--safe-mode] [--source SOURCE] [--tui] [--cli] [--dev]
|
||||
|
||||
Start an interactive chat session with Hermes Agent
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
-q, --query QUERY Query to run. On a real TTY the prompt seeds an
|
||||
interactive session (submitted literally as the first
|
||||
turn); combined with --oneshot or -Q, or on a non-TTY,
|
||||
it answers and exits.
|
||||
--query-file PATH Read the single query from a file instead of the
|
||||
command line ('-' reads stdin). Safe for arbitrary
|
||||
text: nothing is shell-interpreted, so quotes, $(...),
|
||||
and backticks are preserved verbatim. Mutually
|
||||
exclusive with -q.
|
||||
--oneshot With -q/--query-file: answer the query and exit
|
||||
(legacy single-query behavior) instead of seeding an
|
||||
interactive session. Implied on non-TTY stdio and by
|
||||
-Q/--quiet.
|
||||
--image IMAGE Optional local image path to attach to a single query
|
||||
-m, --model MODEL Model to use (e.g., anthropic/claude-sonnet-4)
|
||||
-t, --toolsets TOOLSETS
|
||||
Comma-separated toolsets to enable
|
||||
--reasoning LEVEL Reasoning effort for this session: none, minimal, low,
|
||||
medium, high, xhigh, max, or ultra. Overrides
|
||||
agent.reasoning_effort for this run only (same levels
|
||||
as the /reasoning slash command).
|
||||
-s, --skills SKILLS Preload one or more skills for the session (repeat
|
||||
flag or comma-separate)
|
||||
--provider PROVIDER Inference provider (default: auto). Built-in or a
|
||||
user-defined name from `providers:` in config.yaml.
|
||||
-v, --verbose Verbose output
|
||||
-Q, --quiet Quiet mode for programmatic use: suppress banner,
|
||||
spinner, and tool previews. Only output the final
|
||||
response and session info.
|
||||
--resume, -r SESSION_ID
|
||||
Resume a previous session by ID (shown on exit), or
|
||||
'latest' for the most recent session
|
||||
--no-restore-cwd Don't cd into a resumed session's recorded working
|
||||
directory.
|
||||
--in DIR Change into DIR before starting or resuming (scopes '
|
||||
--resume latest' / -c lookups to DIR's workspace).
|
||||
--continue, -c [SESSION_NAME]
|
||||
Resume a session by name, or the most recent if no
|
||||
name given
|
||||
--create-if-missing With -c/--continue <name>: if no session matches the
|
||||
name, create a new session with that title and proceed
|
||||
(instead of failing with a not-found error).
|
||||
Programmatic callers that want 'send to this named
|
||||
thread, making it if needed'.
|
||||
--worktree, -w Run in an isolated git worktree (for parallel agents
|
||||
on the same repo)
|
||||
--accept-hooks Auto-approve any unseen shell hooks declared in
|
||||
config.yaml without a TTY prompt (see also
|
||||
HERMES_ACCEPT_HOOKS env var and hooks_auto_accept: in
|
||||
config.yaml).
|
||||
--checkpoints Enable filesystem checkpoints before destructive file
|
||||
operations (use /rollback to restore)
|
||||
--max-turns N Maximum tool-calling iterations per conversation turn
|
||||
(default: 500, or agent.max_turns in config)
|
||||
--run-budget SECONDS Optional wall-clock budget in seconds for each
|
||||
conversation run. At 80% elapsed the agent gets a one-
|
||||
time wrap-up notice, and implicit provider stale
|
||||
timeouts are capped to the remaining budget so one
|
||||
hung call can't consume the run. Unset = off. Also
|
||||
configurable as agent.run_budget_seconds in
|
||||
config.yaml. Intended for one-shot/eval invocations
|
||||
with a hard ceiling.
|
||||
--yolo Bypass all dangerous command approval prompts (use at
|
||||
your own risk)
|
||||
--pass-session-id Include the session ID in the agent's system prompt
|
||||
--ignore-user-config Ignore ~/.hermes/config.yaml and fall back to built-in
|
||||
defaults (credentials in .env are still loaded).
|
||||
Useful for isolated CI runs, reproduction, and third-
|
||||
party integrations.
|
||||
--ignore-rules Skip auto-injection of AGENTS.md, SOUL.md,
|
||||
.cursorrules, memory, and preloaded skills. Combine
|
||||
with --ignore-user-config for a fully isolated run.
|
||||
--safe-mode Troubleshooting mode: disable ALL customizations —
|
||||
user config, AGENTS.md/memory injection, plugins, and
|
||||
MCP servers (implies --ignore-user-config and
|
||||
--ignore-rules). Use to isolate whether a problem
|
||||
comes from your setup or from Hermes itself.
|
||||
--source SOURCE Session source tag for filtering (default: cli). Use
|
||||
'tool' for third-party integrations that should not
|
||||
appear in user session lists.
|
||||
--tui Launch the modern TUI instead of the classic REPL
|
||||
--cli Force the classic prompt_toolkit REPL (overrides
|
||||
display.interface=tui)
|
||||
--dev With --tui: run TypeScript sources via tsx (skip dist
|
||||
build)
|
||||
===== hermes model --help =====
|
||||
usage: hermes model [-h] [--refresh] [--portal-url PORTAL_URL]
|
||||
[--inference-url INFERENCE_URL] [--client-id CLIENT_ID]
|
||||
[--scope SCOPE] [--no-browser] [--timeout TIMEOUT]
|
||||
[--ca-bundle CA_BUNDLE] [--insecure]
|
||||
|
||||
Interactively select your inference provider and default model
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--refresh Wipe the model picker disk cache and re-fetch every
|
||||
provider's live /v1/models list.
|
||||
--portal-url PORTAL_URL
|
||||
Portal base URL for Nous login (default: production
|
||||
portal)
|
||||
--inference-url INFERENCE_URL
|
||||
Inference API base URL for Nous login (default:
|
||||
production inference API)
|
||||
--client-id CLIENT_ID
|
||||
OAuth client id to use for Nous login (default:
|
||||
hermes-cli)
|
||||
--scope SCOPE OAuth scope to request for Nous login
|
||||
--no-browser Do not attempt to open the browser automatically
|
||||
during Nous login
|
||||
--timeout TIMEOUT HTTP request timeout in seconds for Nous login
|
||||
(default: 15)
|
||||
--ca-bundle CA_BUNDLE
|
||||
Path to CA bundle PEM file for Nous TLS verification
|
||||
--insecure Disable TLS verification for Nous login (testing only)
|
||||
===== hermes config --help =====
|
||||
usage: hermes config [-h]
|
||||
{show,edit,get,set,unset,path,env-path,check,migrate} ...
|
||||
|
||||
Manage Hermes Agent configuration
|
||||
|
||||
positional arguments:
|
||||
{show,edit,get,set,unset,path,env-path,check,migrate}
|
||||
show Show current configuration
|
||||
edit Open config file in editor
|
||||
get Print a resolved configuration value
|
||||
set Set a configuration value
|
||||
unset Remove a configuration value
|
||||
path Print config file path
|
||||
env-path Print .env file path
|
||||
check Check for missing/outdated config
|
||||
migrate Update config with new options
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes cron --help =====
|
||||
usage: hermes cron [-h] [--accept-hooks]
|
||||
{list,create,add,edit,pause,resume,run,remove,rm,delete,status,runs,history,incidents,notepad,doctor,tick} ...
|
||||
|
||||
Manage scheduled tasks
|
||||
|
||||
positional arguments:
|
||||
{list,create,add,edit,pause,resume,run,remove,rm,delete,status,runs,history,incidents,notepad,doctor,tick}
|
||||
list List scheduled jobs
|
||||
create (add) Create a scheduled job
|
||||
edit Edit an existing scheduled job
|
||||
pause Pause a scheduled job
|
||||
resume Resume a paused job
|
||||
run Run a job on the next scheduler tick
|
||||
remove (rm, delete)
|
||||
Remove a scheduled job
|
||||
status Check if cron scheduler is running
|
||||
runs (history) Show durable execution attempts
|
||||
incidents List or acknowledge durable cron failure incidents
|
||||
notepad Read/write a job's durable notepad (persistent KV
|
||||
across runs)
|
||||
doctor Check scheduled jobs for common health issues
|
||||
tick Run due jobs once and exit
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
|
||||
(equivalent to HERMES_ACCEPT_HOOKS=1 /
|
||||
hooks_auto_accept: true).
|
||||
===== hermes kanban --help =====
|
||||
usage: hermes kanban [-h] [--board <slug>]
|
||||
{init,boards,create,swarm,list,ls,show,assign,set-model,reclaim,reassign,diagnostics,diag,link,unlink,claim,comment,attach,attachments,attach-rm,complete,edit,block,schedule,unblock,request-review,request-changes,reopen-review,promote,archive,tail,dispatch,daemon,watch,stats,notify-subscribe,notify-list,notify-unsubscribe,log,runs,heartbeat,assignees,context,specify,decompose,gc,repair} ...
|
||||
|
||||
Durable SQLite-backed task board shared across Hermes profiles. Tasks are
|
||||
claimed atomically, can depend on other tasks, and are executed by a named
|
||||
profile in an isolated workspace. See https://hermes-
|
||||
agent.nousresearch.com/docs/user-guide/features/kanban or docs/hermes-
|
||||
kanban-v1-spec.pdf for the full design.
|
||||
|
||||
positional arguments:
|
||||
{init,boards,create,swarm,list,ls,show,assign,set-model,reclaim,reassign,diagnostics,diag,link,unlink,claim,comment,attach,attachments,attach-rm,complete,edit,block,schedule,unblock,request-review,request-changes,reopen-review,promote,archive,tail,dispatch,daemon,watch,stats,notify-subscribe,notify-list,notify-unsubscribe,log,runs,heartbeat,assignees,context,specify,decompose,gc,repair}
|
||||
init Create kanban.db if missing (idempotent)
|
||||
boards Manage kanban boards (one board per project /
|
||||
workstream)
|
||||
create Create a new task
|
||||
swarm Create a Kanban Swarm v1 graph (parallel workers →
|
||||
verifier → synthesizer)
|
||||
list (ls) List tasks
|
||||
show Show a task with comments + events
|
||||
assign Assign or reassign a task
|
||||
set-model Set or clear a task's model/provider override (takes
|
||||
effect on the next dispatch)
|
||||
reclaim Release an active worker claim on a running task
|
||||
reassign Reassign a task to a different profile, optionally
|
||||
reclaiming first
|
||||
diagnostics (diag) List active diagnostics on the current board
|
||||
link Add a parent->child dependency
|
||||
unlink Remove a parent->child dependency
|
||||
claim Atomically claim a ready task (prints resolved
|
||||
workspace path)
|
||||
comment Append a comment
|
||||
attach Attach a local file to a task
|
||||
attachments List a task's attachments
|
||||
attach-rm Delete an attachment by id
|
||||
complete Mark one or more tasks done
|
||||
edit Edit recovery fields on an already-completed task
|
||||
block Mark one or more tasks blocked
|
||||
schedule Park one or more tasks in Scheduled (waiting on time,
|
||||
not human input)
|
||||
unblock Return blocked/scheduled tasks to ready, or todo while
|
||||
parents remain open
|
||||
request-review Move a task to 'review' (implementation done, awaiting
|
||||
review) — NOT a block
|
||||
request-changes Reviewer verdict: return the active review run to its
|
||||
implementer
|
||||
reopen-review Send one or more review tasks back for changes (review
|
||||
-> ready/todo)
|
||||
promote Manually move one or more todo/blocked tasks to ready
|
||||
(recovery path)
|
||||
archive Archive one or more tasks
|
||||
tail Follow a task's event stream
|
||||
dispatch One dispatcher pass: reclaim stale, promote ready,
|
||||
spawn workers
|
||||
daemon DEPRECATED — dispatcher now runs in the gateway. Use
|
||||
`hermes gateway start`.
|
||||
watch Live-stream task_events to the terminal (Ctrl+C to
|
||||
exit)
|
||||
stats Per-status + per-assignee counts + oldest-ready age
|
||||
notify-subscribe Subscribe a gateway source to a task's terminal events
|
||||
(used by /kanban subscribe in the gateway adapter)
|
||||
notify-list List notification subscriptions (optionally for a
|
||||
single task)
|
||||
notify-unsubscribe Remove a gateway subscription from a task
|
||||
log Print the worker log for a task (from <kanban-
|
||||
root>/kanban/logs/)
|
||||
runs Show attempt history for a task (one row per run:
|
||||
profile, outcome, elapsed, summary)
|
||||
heartbeat Emit a heartbeat event for a running task (worker
|
||||
liveness signal)
|
||||
assignees List known profiles + per-profile task counts (union
|
||||
of ~/.hermes/profiles/ and current assignees on the
|
||||
board)
|
||||
context Print the full context a worker sees for a task (title
|
||||
+ body + parent results + comments).
|
||||
specify Flesh out a triage-column task into a concrete spec
|
||||
(title + body) and promote it to todo. Uses the
|
||||
auxiliary LLM configured under
|
||||
auxiliary.triage_specifier.
|
||||
decompose Decompose a triage-column task into a graph of child
|
||||
tasks routed to specialist profiles by description.
|
||||
Falls back to specify-style single-task promotion when
|
||||
the task doesn't benefit from fan-out. Uses
|
||||
auxiliary.kanban_decomposer.
|
||||
gc Garbage-collect archived-task workspaces, old events,
|
||||
and old logs
|
||||
repair Check kanban.db integrity and auto-repair index-only
|
||||
corruption
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--board <slug> Board slug to operate on. Defaults to the current
|
||||
board (set via `hermes kanban boards switch <slug>` or
|
||||
the HERMES_KANBAN_BOARD env var). Use `hermes kanban
|
||||
boards list` to see all boards.
|
||||
===== hermes skills --help =====
|
||||
usage: hermes skills [-h]
|
||||
{trust,untrust,browse,search,install,inspect,list,check,update,audit,uninstall,reset,list-modified,diff,opt-out,opt-in,repair-official,publish,snapshot,tap,config} ...
|
||||
|
||||
Search, install, inspect, audit, configure, and manage skills from skills.sh,
|
||||
well-known agent skill endpoints, GitHub, ClawHub, and other registries.
|
||||
|
||||
positional arguments:
|
||||
{trust,untrust,browse,search,install,inspect,list,check,update,audit,uninstall,reset,list-modified,diff,opt-out,opt-in,repair-official,publish,snapshot,tap,config}
|
||||
trust Trust a project so its repo-local skills
|
||||
(./.hermes/skills, ./.agents/skills) load
|
||||
untrust Revoke project-skill trust for a repo
|
||||
browse Browse all available skills (paginated)
|
||||
search Search skill registries
|
||||
install Install a skill
|
||||
inspect Preview a skill without installing
|
||||
list List installed skills
|
||||
check Check installed hub skills for updates
|
||||
update Update installed hub skills
|
||||
audit Re-scan installed hub skills
|
||||
uninstall Remove a hub-installed skill
|
||||
reset Reset a bundled skill — clears 'user-modified'
|
||||
tracking so updates work again
|
||||
list-modified List bundled skills you've edited (which `hermes
|
||||
update` keeps)
|
||||
diff Show how your copy of a bundled skill differs from the
|
||||
stock version
|
||||
opt-out Stop bundled skills from being seeded into this
|
||||
profile
|
||||
opt-in Re-enable bundled-skill seeding (undo opt-out)
|
||||
repair-official Backfill or restore official optional skills from repo
|
||||
source
|
||||
publish Publish a skill to a registry
|
||||
snapshot Export/import skill configurations
|
||||
tap Manage skill sources
|
||||
config Interactive skill configuration — enable/disable
|
||||
individual skills
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes sessions --help =====
|
||||
usage: hermes sessions [-h]
|
||||
{list,export,delete,prune,archive,optimize,clean-markers,optimize-storage,repair,repair-routing,recover,stats,rename,pin,unpin,pinned,retitle-skills,browse,import} ...
|
||||
|
||||
View and manage the SQLite session store
|
||||
|
||||
positional arguments:
|
||||
{list,export,delete,prune,archive,optimize,clean-markers,optimize-storage,repair,repair-routing,recover,stats,rename,pin,unpin,pinned,retitle-skills,browse,import}
|
||||
list List recent sessions
|
||||
export Export sessions to JSONL, Markdown, or QMD
|
||||
delete Delete a specific session
|
||||
prune Delete old sessions (filterable by time window,
|
||||
source, title, ...)
|
||||
archive Bulk-archive (soft-hide) sessions matching filters —
|
||||
no deletion
|
||||
optimize Reclaim disk space: merge FTS5 segments + VACUUM (no
|
||||
data change)
|
||||
clean-markers Permanently clear stale tool-call marker content left
|
||||
by sessions from before #78148
|
||||
optimize-storage Migrate the search index to the compact v23 layout
|
||||
(reclaims disk on large DBs)
|
||||
repair Repair a malformed state.db schema so hidden sessions
|
||||
reappear
|
||||
repair-routing Re-stamp gateway sessions that lost their routing
|
||||
identity
|
||||
recover Rebuild canonical session data into a separate clean
|
||||
database
|
||||
stats Show session store statistics
|
||||
rename Set or change a session's title
|
||||
pin Pin session(s) — durable keep flag, exempt from auto-
|
||||
archive
|
||||
unpin Remove the pin (durable keep flag) from session(s)
|
||||
pinned List pinned sessions
|
||||
retitle-skills Re-title sessions whose auto-title came from a
|
||||
/skill's own text
|
||||
browse Interactive session picker — browse, search, and
|
||||
resume sessions
|
||||
import Import a Claude Code or Codex CLI session into Hermes
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes mcp --help =====
|
||||
usage: hermes mcp [-h] [--accept-hooks]
|
||||
{serve,add,remove,rm,list,ls,test,configure,config,login,reauth,picker,catalog,install} ...
|
||||
|
||||
Manage MCP server connections and run Hermes as an MCP server. MCP servers
|
||||
provide additional tools via the Model Context Protocol. Use 'hermes mcp add'
|
||||
to connect to a new server, or 'hermes mcp serve' to expose Hermes
|
||||
conversations over MCP.
|
||||
|
||||
positional arguments:
|
||||
{serve,add,remove,rm,list,ls,test,configure,config,login,reauth,picker,catalog,install}
|
||||
serve Run Hermes as an MCP server (expose conversations to
|
||||
other agents)
|
||||
add Add an MCP server (discovery-first install)
|
||||
remove (rm) Remove an MCP server
|
||||
list (ls) List configured MCP servers
|
||||
test Test MCP server connection
|
||||
configure (config) Toggle tool selection
|
||||
login Force re-authentication for an OAuth-based MCP server
|
||||
reauth Re-authenticate one OAuth MCP server, or all of them
|
||||
(--all)
|
||||
picker Interactive catalog picker (also the default for
|
||||
`hermes mcp`)
|
||||
catalog List Nous-approved MCPs available for one-click
|
||||
install
|
||||
install Install a catalog MCP by name (e.g. `hermes mcp
|
||||
install n8n`)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
|
||||
(equivalent to HERMES_ACCEPT_HOOKS=1 /
|
||||
hooks_auto_accept: true).
|
||||
===== hermes profile --help =====
|
||||
usage: hermes profile [-h]
|
||||
{list,use,create,delete,describe,show,alias,rename,export,import,install,update,info} ...
|
||||
|
||||
positional arguments:
|
||||
{list,use,create,delete,describe,show,alias,rename,export,import,install,update,info}
|
||||
list List all profiles
|
||||
use Set sticky default profile
|
||||
create Create a new profile
|
||||
delete Delete a profile
|
||||
describe Read or set a profile's description (used by the
|
||||
kanban orchestrator)
|
||||
show Show profile details
|
||||
alias Manage wrapper scripts
|
||||
rename Rename a profile ('default': sets a display name; id
|
||||
unchanged)
|
||||
export Export a profile to archive
|
||||
import Import a profile from archive
|
||||
install Install a profile distribution from a git URL or local
|
||||
directory
|
||||
update Re-pull a distribution and apply updates (user data
|
||||
preserved)
|
||||
info Show a profile's distribution manifest (version,
|
||||
requirements, source)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes memory --help =====
|
||||
usage: hermes memory [-h] {setup,status,off,reset} ...
|
||||
|
||||
Set up and manage external memory provider plugins. Available providers:
|
||||
honcho, openviking, mem0, hindsight, holographic, retaindb, byterover. Only
|
||||
one external provider can be active at a time. Built-in memory
|
||||
(MEMORY.md/USER.md) is always active.
|
||||
|
||||
positional arguments:
|
||||
{setup,status,off,reset}
|
||||
setup Interactive provider selection and configuration
|
||||
status Show current memory provider config
|
||||
off Disable external provider (built-in only)
|
||||
reset Erase all built-in memory (MEMORY.md and USER.md)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes tools --help =====
|
||||
usage: hermes tools [-h] [--summary] {list,disable,enable,post-setup} ...
|
||||
|
||||
Enable, disable, or list tools for CLI, Telegram, Discord, etc. Built-in
|
||||
toolsets use plain names (e.g. web, memory). MCP tools use server:tool
|
||||
notation (e.g. github:create_issue). Run 'hermes tools' with no subcommand for
|
||||
the interactive configuration UI.
|
||||
|
||||
positional arguments:
|
||||
{list,disable,enable,post-setup}
|
||||
list Show all tools and their enabled/disabled status
|
||||
disable Disable toolsets or MCP tools
|
||||
enable Enable toolsets or MCP tools
|
||||
post-setup Run a provider's post-setup install hook
|
||||
(npm/pip/binary)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--summary Print a summary of enabled tools per platform and exit
|
||||
===== hermes project --help =====
|
||||
usage: hermes project [-h]
|
||||
{create,list,ls,show,add-folder,remove-folder,rename,set-primary,use,archive,restore,bind-board} ...
|
||||
|
||||
Projects are human-named workspaces that can span multiple folders / repos.
|
||||
They anchor desktop session grouping and, when bound to a kanban board, give
|
||||
tasks a deterministic worktree + branch convention. State is per-profile.
|
||||
|
||||
positional arguments:
|
||||
{create,list,ls,show,add-folder,remove-folder,rename,set-primary,use,archive,restore,bind-board}
|
||||
create Create a new project
|
||||
list (ls) List projects
|
||||
show Show a project's details
|
||||
add-folder Add a folder to a project
|
||||
remove-folder Remove a folder from a project
|
||||
rename Rename a project
|
||||
set-primary Set the primary folder
|
||||
use Set the active project
|
||||
archive Archive a project
|
||||
restore Restore an archived project
|
||||
bind-board Bind a kanban board to a project
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes gateway --help =====
|
||||
usage: hermes gateway [-h] [--accept-hooks]
|
||||
{run,start,stop,restart,status,install,uninstall,list,setup,migrate-legacy,enroll} ...
|
||||
|
||||
Manage the messaging gateway (Telegram, Discord, WhatsApp, Weixin, and more)
|
||||
|
||||
positional arguments:
|
||||
{run,start,stop,restart,status,install,uninstall,list,setup,migrate-legacy,enroll}
|
||||
run Run gateway in foreground (recommended for WSL,
|
||||
Docker, Termux)
|
||||
start Start the installed systemd/launchd background service
|
||||
stop Stop gateway service
|
||||
restart Restart gateway service
|
||||
status Show gateway status
|
||||
install Install gateway as a systemd/launchd background
|
||||
service
|
||||
uninstall Uninstall gateway service
|
||||
list List all profiles and their gateway status
|
||||
setup Configure messaging platforms
|
||||
migrate-legacy Remove legacy hermes.service units from pre-rename
|
||||
installs
|
||||
enroll Enroll this gateway with a relay connector (writes
|
||||
relay auth creds to .env)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--accept-hooks Auto-approve unseen shell hooks without a TTY prompt
|
||||
(equivalent to HERMES_ACCEPT_HOOKS=1 /
|
||||
hooks_auto_accept: true).
|
||||
===== hermes computer-use --help =====
|
||||
usage: hermes computer-use [-h] {install,status,doctor,permissions} ...
|
||||
|
||||
Install or check the cua-driver binary used by the `computer_use` toolset.
|
||||
Supported on macOS, Windows, and Linux. Use `hermes computer-use install` to
|
||||
fetch and run the upstream cua-driver installer. This is equivalent to the
|
||||
post-setup hook that `hermes tools` runs when you first enable the Computer
|
||||
Use toolset, and is a stable target for re-running the install if it didn't
|
||||
fire (e.g. when toggling the toolset on a returning-user setup). Use `hermes
|
||||
computer-use doctor` to run cua-driver's `health_report` MCP tool and surface
|
||||
its check matrix (TCC, bundle identity, version, platform support, ...) in
|
||||
human-readable form.
|
||||
|
||||
positional arguments:
|
||||
{install,status,doctor,permissions}
|
||||
install Install or repair the cua-driver binary
|
||||
(macOS/Windows/Linux)
|
||||
status Print whether cua-driver is installed and on PATH
|
||||
doctor Run cua-driver `health_report` and surface the check
|
||||
matrix
|
||||
permissions Check or grant macOS Accessibility + Screen Recording
|
||||
(macOS)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
===== hermes doctor --help =====
|
||||
usage: hermes doctor [-h] [--fix] [--live] [--ack ADVISORY_ID]
|
||||
|
||||
Diagnose issues with Hermes Agent setup
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--fix Attempt to fix issues automatically
|
||||
--live Opt-in: run one bounded, read-only real-call health probe
|
||||
per configured tool backend
|
||||
(Firecrawl/FAL/browser/MCP/TTS/STT) after the static
|
||||
checks. Makes real network calls.
|
||||
--ack ADVISORY_ID Acknowledge a security advisory by ID and exit. After
|
||||
ack, the advisory will no longer trigger startup banners.
|
||||
Run `hermes doctor` first to see active advisories and
|
||||
their IDs.
|
||||
===== hermes status --help =====
|
||||
usage: hermes status [-h] [--all] [--deep]
|
||||
|
||||
Display status of Hermes Agent components
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--all Show all details (redacted for sharing)
|
||||
--deep Run deep checks (may take longer)
|
||||
===== hermes import-agent --help =====
|
||||
usage: hermes import-agent [-h] [--source SOURCE] [--dry-run] [--overwrite]
|
||||
[--yes]
|
||||
[{claude-code,codex}]
|
||||
|
||||
One-command import of another coding agent's setup into Hermes. Maps
|
||||
CLAUDE.md/AGENTS.md instructions, permission allowlists, MCP servers, skills,
|
||||
and memories into their Hermes equivalents. Always shows a preview before
|
||||
making changes. API keys and credentials are never imported — run 'hermes
|
||||
setup' for those.
|
||||
|
||||
positional arguments:
|
||||
{claude-code,codex} Which agent to import from (default: auto-detect
|
||||
~/.claude or ~/.codex)
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--source SOURCE Path to the agent's config directory (default:
|
||||
~/.claude or ~/.codex)
|
||||
--dry-run Preview only — stop after showing what would be
|
||||
imported
|
||||
--overwrite Overwrite existing Hermes items on name conflicts
|
||||
(default: skip)
|
||||
--yes, -y Skip confirmation prompts
|
||||
@@ -1,33 +0,0 @@
|
||||
import json, subprocess, re
|
||||
|
||||
data = json.load(open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json"))
|
||||
have = {r["videoId"] for r in data}
|
||||
extra = []
|
||||
for v in ["8GjyOQy19so", "zwqhemjHq3E"]:
|
||||
if v in have:
|
||||
continue
|
||||
url = f"https://www.youtube.com/watch?v={v}"
|
||||
oe = subprocess.run(["curl", "-s", f"https://www.youtube.com/oembed?url={url}&format=json"],
|
||||
capture_output=True, text=True, timeout=30)
|
||||
try:
|
||||
oej = json.loads(oe.stdout)
|
||||
o = {"status": "PASS", "title": oej.get("title"), "author": oej.get("author_name")}
|
||||
except Exception:
|
||||
o = {"status": "FAIL", "raw": oe.stdout[:200]}
|
||||
wp = subprocess.run(["curl", "-s", "-L", url,
|
||||
"-H", "User-Agent: Mozilla/5.0 (Windows NT 10.0) Chrome/124.0",
|
||||
"-H", "Accept-Language: en-US"], capture_output=True, text=True, timeout=30)
|
||||
html = wp.stdout
|
||||
pub = re.search(r'"publishDate":"([\d\-T:Z]+)"', html)
|
||||
dur = re.search(r'"lengthSeconds":"(\d+)"', html)
|
||||
views = re.search(r'"viewCount":"(\d+)"', html)
|
||||
rec = {"videoId": v, "oembed": o,
|
||||
"publishDate": pub.group(1) if pub else None,
|
||||
"lengthSeconds": int(dur.group(1)) if dur else None,
|
||||
"views": int(views.group(1)) if views else None}
|
||||
extra.append(rec)
|
||||
print(json.dumps(rec))
|
||||
|
||||
data.extend(extra)
|
||||
json.dump(data, open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json", "w"), indent=2)
|
||||
print("total videos in ledger:", len(data))
|
||||
@@ -1,38 +0,0 @@
|
||||
import json, subprocess, re, sys
|
||||
|
||||
vids = ["9GpWELm3_XI","iqN6MVzpJTk","Ta2wg6xPaY4","8tpuky8HpXw","tP6yf22OJdI",
|
||||
"UWjh5Z4s8jY","5d02TYoOzfE","83nWNRKZTCE","6M2tItdARew","5_N84t1rUU0",
|
||||
"UTZhvPXnmwA","1UgXUjT-QtI","6GtF_uHbGhw","CwPUOVUdApE","sCa3BtpkziQ",
|
||||
"jmtpYUOr7_U","cu2fgknmemA","4sAmpcSOVEw","P2LIFtrRr2U"]
|
||||
|
||||
results = []
|
||||
for v in vids:
|
||||
url = f"https://www.youtube.com/watch?v={v}"
|
||||
oe = subprocess.run(["curl","-s",f"https://www.youtube.com/oembed?url={url}&format=json"],
|
||||
capture_output=True, text=True, timeout=30)
|
||||
try:
|
||||
oej = json.loads(oe.stdout)
|
||||
oembed = {"status":"PASS","title":oej.get("title"),"author":oej.get("author_name")}
|
||||
except Exception:
|
||||
oembed = {"status":"FAIL","raw":oe.stdout[:200]}
|
||||
# watch page for date + duration
|
||||
wp = subprocess.run(["curl","-s","-L",url,"-H","User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/124.0",
|
||||
"-H","Accept-Language: en-US,en;q=0.9"], capture_output=True, text=True, timeout=30)
|
||||
html = wp.stdout
|
||||
pub = re.search(r'"publishDate":"([\d\-T:Z]+)"', html)
|
||||
upd = re.search(r'"uploadDate":"([\d\-T:Z]+)"', html)
|
||||
dur = re.search(r'"lengthSeconds":"(\d+)"', html)
|
||||
views = re.search(r'"viewCount":"(\d+)"', html)
|
||||
results.append({
|
||||
"videoId": v, "oembed": oembed,
|
||||
"publishDate": pub.group(1) if pub else None,
|
||||
"uploadDate": upd.group(1) if upd else None,
|
||||
"lengthSeconds": int(dur.group(1)) if dur else None,
|
||||
"views": int(views.group(1)) if views else None,
|
||||
"watchpage_bytes": len(html),
|
||||
})
|
||||
print(json.dumps(results[-1]))
|
||||
|
||||
with open("/home/hermes/syslog/drafts/scot-hermes-playbook/research/video-verification.json","w") as f:
|
||||
json.dump(results, f, indent=2)
|
||||
print("saved video-verification.json")
|
||||
@@ -1,271 +0,0 @@
|
||||
[
|
||||
{
|
||||
"videoId": "9GpWELm3_XI",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent Explained In 5 Minutes",
|
||||
"author": "CodeHead"
|
||||
},
|
||||
"publishDate": "2026-05-23T08:00:27-07:00",
|
||||
"uploadDate": "2026-05-23T08:00:27-07:00",
|
||||
"lengthSeconds": 293,
|
||||
"views": 251275,
|
||||
"watchpage_bytes": 1419048
|
||||
},
|
||||
{
|
||||
"videoId": "iqN6MVzpJTk",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent by Nous Research: The Open-Source Agent Model Everyone Is Switching To",
|
||||
"author": "Praveen Govindaraj"
|
||||
},
|
||||
"publishDate": "2026-02-26T21:56:54-08:00",
|
||||
"uploadDate": "2026-02-26T21:56:54-08:00",
|
||||
"lengthSeconds": 193,
|
||||
"views": 3140,
|
||||
"watchpage_bytes": 1305120
|
||||
},
|
||||
{
|
||||
"videoId": "Ta2wg6xPaY4",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Learn 95% of Hermes Agent in 31 Minutes",
|
||||
"author": "Sharbel A."
|
||||
},
|
||||
"publishDate": "2026-08-09T07:00:19-07:00",
|
||||
"uploadDate": "2026-08-09T07:00:19-07:00",
|
||||
"lengthSeconds": 1888,
|
||||
"views": 129313,
|
||||
"watchpage_bytes": 1468130
|
||||
},
|
||||
{
|
||||
"videoId": "8tpuky8HpXw",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Better than OpenClaw? Testing Hermes Agent w/ Qwen 3 model",
|
||||
"author": "Tonbi's AI Garage"
|
||||
},
|
||||
"publishDate": "2026-03-11T07:00:14-07:00",
|
||||
"uploadDate": "2026-03-11T07:00:14-07:00",
|
||||
"lengthSeconds": 907,
|
||||
"views": 18167,
|
||||
"watchpage_bytes": 1326925
|
||||
},
|
||||
{
|
||||
"videoId": "tP6yf22OJdI",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Did Hermes Agent just kill OpenClaw? (full guide)",
|
||||
"author": "Alex Finn"
|
||||
},
|
||||
"publishDate": "2026-03-31T06:15:10-07:00",
|
||||
"uploadDate": "2026-03-31T06:15:10-07:00",
|
||||
"lengthSeconds": 835,
|
||||
"views": 132128,
|
||||
"watchpage_bytes": 1390199
|
||||
},
|
||||
{
|
||||
"videoId": "UWjh5Z4s8jY",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra",
|
||||
"author": "Peter Yang"
|
||||
},
|
||||
"publishDate": "2026-08-02T06:00:12-07:00",
|
||||
"uploadDate": "2026-08-02T06:00:12-07:00",
|
||||
"lengthSeconds": 2804,
|
||||
"views": 36613,
|
||||
"watchpage_bytes": 1386664
|
||||
},
|
||||
{
|
||||
"videoId": "5d02TYoOzfE",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Use This To Make The Hermes Agent Basically Free",
|
||||
"author": "AI LABS"
|
||||
},
|
||||
"publishDate": "2026-07-01T07:00:26-07:00",
|
||||
"uploadDate": "2026-07-01T07:00:26-07:00",
|
||||
"lengthSeconds": 788,
|
||||
"views": 54858,
|
||||
"watchpage_bytes": 1412833
|
||||
},
|
||||
{
|
||||
"videoId": "83nWNRKZTCE",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Meet the AI Agent That Grows With You Hermes Agent by Nous Research",
|
||||
"author": "Eddy Says Hi #EddySaysHi"
|
||||
},
|
||||
"publishDate": "2026-03-21T13:00:09-07:00",
|
||||
"uploadDate": "2026-03-21T13:00:09-07:00",
|
||||
"lengthSeconds": 366,
|
||||
"views": 465,
|
||||
"watchpage_bytes": 1254068
|
||||
},
|
||||
{
|
||||
"videoId": "6M2tItdARew",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "The AI Agent That Never Forgets: Meet Hermes Agent by Nous Research",
|
||||
"author": "Siggi"
|
||||
},
|
||||
"publishDate": "2026-03-11T13:53:17-07:00",
|
||||
"uploadDate": "2026-03-11T13:53:17-07:00",
|
||||
"lengthSeconds": 371,
|
||||
"views": 1199,
|
||||
"watchpage_bytes": 1267812
|
||||
},
|
||||
{
|
||||
"videoId": "5_N84t1rUU0",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent Fundamentals In 29 Minutes",
|
||||
"author": "Tina Huang"
|
||||
},
|
||||
"publishDate": "2026-07-20T09:37:02-07:00",
|
||||
"uploadDate": "2026-07-20T09:37:02-07:00",
|
||||
"lengthSeconds": 1780,
|
||||
"views": 462764,
|
||||
"watchpage_bytes": 1561142
|
||||
},
|
||||
{
|
||||
"videoId": "UTZhvPXnmwA",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent: Agents that grow with you |Episode #357|",
|
||||
"author": "Practical AI"
|
||||
},
|
||||
"publishDate": "2026-05-20T15:00:17-07:00",
|
||||
"uploadDate": "2026-05-20T15:00:17-07:00",
|
||||
"lengthSeconds": 2853,
|
||||
"views": 1888,
|
||||
"watchpage_bytes": 1329765
|
||||
},
|
||||
{
|
||||
"videoId": "1UgXUjT-QtI",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent: Why Everyone's Ditching OpenClaw in 2026",
|
||||
"author": "Luke Alexander AI"
|
||||
},
|
||||
"publishDate": "2026-03-26T07:10:29-07:00",
|
||||
"uploadDate": "2026-03-26T07:10:29-07:00",
|
||||
"lengthSeconds": 1083,
|
||||
"views": 13625,
|
||||
"watchpage_bytes": 1321848
|
||||
},
|
||||
{
|
||||
"videoId": "6GtF_uHbGhw",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Every Level of Hermes Agent Explained",
|
||||
"author": "Jack Roberts"
|
||||
},
|
||||
"publishDate": "2026-06-17T12:27:42-07:00",
|
||||
"uploadDate": "2026-06-17T12:27:42-07:00",
|
||||
"lengthSeconds": 1535,
|
||||
"views": 162977,
|
||||
"watchpage_bytes": 1577121
|
||||
},
|
||||
{
|
||||
"videoId": "CwPUOVUdApE",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent: The Ultimate Beginner\u2019s Guide",
|
||||
"author": "Metics Media"
|
||||
},
|
||||
"publishDate": "2026-04-24T06:59:04-07:00",
|
||||
"uploadDate": "2026-04-24T06:59:04-07:00",
|
||||
"lengthSeconds": 2228,
|
||||
"views": 118612,
|
||||
"watchpage_bytes": 1637578
|
||||
},
|
||||
{
|
||||
"videoId": "sCa3BtpkziQ",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "100 Days With Hermes Agent in 21 Minutes",
|
||||
"author": "Sharbel A."
|
||||
},
|
||||
"publishDate": "2026-06-17T07:47:26-07:00",
|
||||
"uploadDate": "2026-06-17T07:47:26-07:00",
|
||||
"lengthSeconds": 1279,
|
||||
"views": 57581,
|
||||
"watchpage_bytes": 1454230
|
||||
},
|
||||
{
|
||||
"videoId": "jmtpYUOr7_U",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent Just Killed OpenClaw (Full Tutorial)",
|
||||
"author": "Leon van Zyl"
|
||||
},
|
||||
"publishDate": "2026-04-28T04:19:38-07:00",
|
||||
"uploadDate": "2026-04-28T04:19:38-07:00",
|
||||
"lengthSeconds": 1199,
|
||||
"views": 15988,
|
||||
"watchpage_bytes": 1510488
|
||||
},
|
||||
{
|
||||
"videoId": "cu2fgknmemA",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent The 24/7 Self-Evolving AI Agent!",
|
||||
"author": "WorldofAI"
|
||||
},
|
||||
"publishDate": "2026-04-07T00:01:34-07:00",
|
||||
"uploadDate": "2026-04-07T00:01:34-07:00",
|
||||
"lengthSeconds": 555,
|
||||
"views": 46889,
|
||||
"watchpage_bytes": 1545788
|
||||
},
|
||||
{
|
||||
"videoId": "4sAmpcSOVEw",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent - Crash Course for Beginners (AI Agent)",
|
||||
"author": "Adrian Twarog"
|
||||
},
|
||||
"publishDate": "2026-07-21T01:20:41-07:00",
|
||||
"uploadDate": "2026-07-21T01:20:41-07:00",
|
||||
"lengthSeconds": 1338,
|
||||
"views": 46484,
|
||||
"watchpage_bytes": 1585548
|
||||
},
|
||||
{
|
||||
"videoId": "P2LIFtrRr2U",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent: New FREE OpenClaw Alternative!",
|
||||
"author": "Julian Goldie SEO"
|
||||
},
|
||||
"publishDate": "2026-03-09T14:00:32-07:00",
|
||||
"uploadDate": "2026-03-09T14:00:32-07:00",
|
||||
"lengthSeconds": 735,
|
||||
"views": 9975,
|
||||
"watchpage_bytes": 1353286
|
||||
},
|
||||
{
|
||||
"videoId": "8GjyOQy19so",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent Full Tutorial INSTALLATION + USECASES",
|
||||
"author": "CodeHead"
|
||||
},
|
||||
"publishDate": "2026-05-14T08:00:23-07:00",
|
||||
"lengthSeconds": 467,
|
||||
"views": 64285
|
||||
},
|
||||
{
|
||||
"videoId": "zwqhemjHq3E",
|
||||
"oembed": {
|
||||
"status": "PASS",
|
||||
"title": "Hermes Agent vs OpenClaw",
|
||||
"author": "Sharbel A."
|
||||
},
|
||||
"publishDate": "2026-04-20T07:14:00-07:00",
|
||||
"lengthSeconds": 928,
|
||||
"views": 35360
|
||||
}
|
||||
]
|
||||
@@ -1,10 +0,0 @@
|
||||
import re, sys
|
||||
|
||||
html = open(sys.argv[1], encoding='utf-8', errors='ignore').read()
|
||||
ids = re.findall(r'"videoRenderer":\{"videoId":"([\w-]{11})"', html)
|
||||
print("videoRenderer hits:", len(set(ids)))
|
||||
for vid in dict.fromkeys(ids):
|
||||
m = re.search(r'"videoId":"%s".{0,3000}?"title":\{"runs":\[\{"text":"(.*?)"\}' % vid, html, re.S)
|
||||
ch = re.search(r'"videoId":"%s".{0,6000}?"ownerText":\{"runs":\[\{"text":"(.*?)"' % vid, html, re.S)
|
||||
dur = re.search(r'"videoId":"%s".{0,4000}?"lengthText":\{"accessibility".{0,400}?"simpleText":"(.*?)"' % vid, html, re.S)
|
||||
print((vid, m.group(1) if m else "?", ch.group(1) if ch else "?", dur.group(1) if dur else "?"))
|
||||
@@ -1,42 +0,0 @@
|
||||
# 06 — Sources
|
||||
|
||||
Access date for ALL entries: **2026-09-11** (via citation ledger `sources.py`; doc URLs additionally confirmed HTTP 200 by curl -L).
|
||||
|
||||
## Official docs (hermes-agent.nousresearch.com)
|
||||
| # | URL | Supported |
|
||||
|---|-----|-----------|
|
||||
| 1 | https://hermes-agent.nousresearch.com/docs | Docs index; overall feature map |
|
||||
| 3 | https://hermes-agent.nousresearch.com/docs/user-guide/configuration | Config sections, SOUL.md, checkpoints |
|
||||
| 4 | https://hermes-agent.nousresearch.com/docs/reference/slash-commands | Slash command registry (03) |
|
||||
| 5 | https://hermes-agent.nousresearch.com/docs/reference/tools-reference | Toolset inventory (02 §6) |
|
||||
| 6 | https://hermes-agent.nousresearch.com/docs/user-guide/features/cron | Cron surface (02 §9) |
|
||||
| 7 | https://hermes-agent.nousresearch.com/docs/user-guide/features/kanban | Kanban surface (02 §8) |
|
||||
| 8 | https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp | MCP surface + `hermes mcp serve` (02 §5, 04 P3) |
|
||||
| 9 | https://hermes-agent.nousresearch.com/docs/user-guide/features/memory | Memory surface (02 §2) |
|
||||
| 10 | https://hermes-agent.nousresearch.com/docs/user-guide/profiles | Profiles (02 §13) |
|
||||
| 11 | https://hermes-agent.nousresearch.com/docs/integrations/providers | Model/provider routing (02 §12) |
|
||||
| 12 | https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ | Gateway platforms (02 §15) |
|
||||
| 13 | https://hermes-agent.nousresearch.com/docs/user-guide/features/curator | Skill maintenance (02 §3) |
|
||||
| 14 | https://hermes-agent.nousresearch.com/docs/reference/cli-commands | CLI command index cross-check (03) |
|
||||
| 15 | https://hermes-agent.nousresearch.com/docs/reference/skills-catalog | Skills catalog (02 §3) |
|
||||
|
||||
## GitHub
|
||||
| # | URL | Supported |
|
||||
|---|-----|-----------|
|
||||
| 2 | https://github.com/nousresearch/hermes-agent | Repo identity, learning-loop description (01, 02) |
|
||||
|
||||
## Local primary sources (not web URLs; verified on this host)
|
||||
- Live CLI help output, Hermes Agent v0.21.1 (2026.9.7), reference install (Syslog kagentz): `hermes --help` + `hermes {chat,model,config,cron,kanban,skills,sessions,mcp,profile,memory,tools,project,gateway,computer-use,doctor,status,import-agent} --help` → raw dump `cli-help-dump.txt` (577 lines). Basis for all VERIFIED-LIVE tags in 01/02/03/04.
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/delegate-coding-agent/SKILL.md` (v1.0.0) + `references/claude-code.md` (v2.2.1) → 04 Patterns 1-2, 01 Claude Code column.
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/merge-reconciler/SKILL.md` → 04 Pattern 2.
|
||||
- `/home/hermes/.hermes/skills/autonomous-ai-agents/computer-use/SKILL.md` (v2.0.0) → 04 Pattern 5, 02 §11.
|
||||
|
||||
## YouTube verification
|
||||
- YouTube search results pages (scraped 2026-09-11): `https://www.youtube.com/results?search_query=hermes+agent+nous+research` and `...nous+research+hermes+agent+official`
|
||||
- oEmbed endpoint per video: `https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=<ID>&format=json` — all 21 checked IDs returned PASS (HTTP 200, title/author match). Raw evidence incl. publishDate/lengthSeconds/viewCount per watch page: `video-verification.json`.
|
||||
- 17 listed in 05-videos.md + 4 spares; 21/21 pass, 0 fail.
|
||||
|
||||
## Explicit gaps (could not close)
|
||||
1. No official Nous Research-produced tutorial video was found — video list is third-party ecosystem content (disclosed in 05).
|
||||
2. Slash commands were verified DOC-ONLY (https://hermes-agent.nousresearch.com/docs/reference/slash-commands); they require an interactive session to exercise, which this headless run does not have. CLI equivalents were verified live.
|
||||
3. `hermes-agent.nousresearch.com/docs/developer-guide/` returned 404 — developer docs live in-repo (`AGENTS.md` in the GitHub repo), not as a docs site section.
|
||||
@@ -43,7 +43,7 @@ Docker hosts get special attention:
|
||||
|
||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
|
||||
|
||||
## Threat Levels
|
||||
## Threat Levels (GUEST filesystems)
|
||||
|
||||
| Level | Threshold | Response | Escalation |
|
||||
|-------|-----------|----------|------------|
|
||||
@@ -53,6 +53,39 @@ Docker hosts get special attention:
|
||||
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
|
||||
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
|
||||
|
||||
## Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)
|
||||
|
||||
Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are **report-only** — no automatic deletion of media or datastore content ever.
|
||||
|
||||
| Level | Threshold | Response | Escalation |
|
||||
|-------|-----------|----------|------------|
|
||||
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
||||
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
||||
| **HOST-RED** | 95% | Name the volume + % + absolute free space; **media volumes: "capacity decision — owner to decide"; pbs-datastore/host-root: "immediate owner attention"** | Zulip DM + channel alert (state-change only) |
|
||||
|
||||
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
||||
|
||||
**State lives in a small JSON state file:** the scanner resolves it to an absolute path from the script's location — `$(dirname "$0")/state/host-disk-bands.json` (i.e., the `state/` directory next to the `scripts/` directory in the prose-contracts repo). The state file is gitignored runtime state — the scanner creates it on first run. Keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
|
||||
|
||||
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
|
||||
|
||||
```
|
||||
storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
|
||||
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
|
||||
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
|
||||
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
|
||||
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN
|
||||
```
|
||||
|
||||
**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.
|
||||
|
||||
**Action classes by volume type:**
|
||||
- **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
|
||||
- **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
|
||||
- **pbs-datastore** (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.
|
||||
|
||||
**Justification (measured 2026-09-15, firstmate):** storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
|
||||
@@ -127,6 +160,17 @@ from the `report_only_guests` YAML block above.
|
||||
- `retry`: 2 attempts for SSH failures before marking a CT unreachable
|
||||
|
||||
## Execution
|
||||
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||
execution — it only proves the agent read the result and reported it. The actual
|
||||
monitoring work happens in the host cron job.
|
||||
|
||||
|
||||
### Host filesystems: report-only, NEVER auto-delete
|
||||
|
||||
Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.
|
||||
|
||||
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
|
||||
|
||||
@@ -191,6 +235,21 @@ call summary-reporter
|
||||
plan: plan
|
||||
```
|
||||
|
||||
## GC SCHEDULE (PBS datastore only)
|
||||
|
||||
The PBS GC schedule is defined in ONE authoritative place: `/etc/cron.d/pbs-gc` on storepve.
|
||||
The schedule is `0 20 * * *` (20:00 LOCAL = 00:00 UTC, since host timezone is America/New_York).
|
||||
This applies **only** to the PBS datastore (`/tank/pbs-backup`), NOT to media volumes.
|
||||
Media volumes (/media/*) are report-only at all threat levels.
|
||||
|
||||
The cron runs `/usr/local/bin/pbs-gc.sh` which executes:
|
||||
```bash
|
||||
proxmox-backup-manager garbage-collection start storepve-datastore
|
||||
```
|
||||
This is NOT a `prune` operation; it is a GC pass that reclaims unreferenced chunks.
|
||||
There is no `--keep-daily` flag; retention is governed by jobs.cfg (keep-daily=35).
|
||||
The GC does not touch media volumes or any other filesystem.
|
||||
|
||||
## GC Strategies by Host Type
|
||||
|
||||
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
||||
@@ -243,6 +302,12 @@ done
|
||||
|
||||
## Alert Templates
|
||||
|
||||
### HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)
|
||||
```
|
||||
⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
|
||||
Action: {volume_type-specific action}
|
||||
```
|
||||
|
||||
### AMBER (75-84%)
|
||||
```
|
||||
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
||||
@@ -349,7 +414,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
| 108 | media | storepve | lxc | ✅ reachable |
|
||||
| 110 | gitea | minipve | lxc | ✅ reachable |
|
||||
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
|
||||
| 112 | tanko | amdpve | lxc | ✅ reachable |
|
||||
| 112 | tanko | minipve | lxc | ✅ reachable |
|
||||
| 113 | baggy | amdpve | lxc | ✅ reachable |
|
||||
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
|
||||
| 116 | syslog-api | minipve | lxc | ✅ reachable |
|
||||
|
||||
@@ -0,0 +1,169 @@
|
||||
# Contract execution pinning
|
||||
|
||||
Which copy of a contract script actually ran, and how that is proven.
|
||||
|
||||
## Why this exists
|
||||
|
||||
Three times on 2026-09-25 a contract reported a verdict from a copy that was
|
||||
not the merged one:
|
||||
|
||||
1. The ops lane's own clone sat on the merged feature branch
|
||||
`fix/search-stack-multi-engine-20260925` at `8b2eba4` with no `pve_auth`
|
||||
fix, while it executed the daily digest from a different clone. Nothing in
|
||||
the workflow noticed.
|
||||
2. `scripts/search-stack-check.py` was deployed into the pinned runner clone
|
||||
by hand rather than through git.
|
||||
3. A stale local `origin/master` ref made an ancestry check report
|
||||
"unlanded work" for a branch that had in fact merged — the same staleness
|
||||
would have passed a stale script as current.
|
||||
|
||||
A contract verdict is only meaningful if it came from the merged copy. The
|
||||
control is `scripts/revision-preflight.sh`.
|
||||
|
||||
## The rule
|
||||
|
||||
**Every contract pins exactly one clone for execution: the clone that
|
||||
`scripts/contract-run.sh` itself lives in.**
|
||||
|
||||
`contract-run.sh` derives that from its own location (`SCRIPTS_DIR`) and checks
|
||||
the script it is about to run against `origin/master` in the same clone. There
|
||||
is no second path to configure, and no contract may be executed from a
|
||||
hand-copied location.
|
||||
|
||||
| Contract | Script | Pinned clone |
|
||||
| --- | --- | --- |
|
||||
| `infrastructure-monitoring` | `scripts/infra-monitoring.sh` | the clone containing `contract-run.sh` |
|
||||
| `proxmox-monitor` | `scripts/proxmox-monitor.sh` | same |
|
||||
| `zulip-health` | `scripts/zulip-monitor.sh` | same |
|
||||
| `agent-health-check` | `scripts/agent-health-check.py` | same |
|
||||
| `litellm-health` | `scripts/litellm-health-check.py` | same |
|
||||
| `disk-gc-threat-response` | `scripts/disk-gc-scan.py` | same |
|
||||
| `pm2-self-heal` | `scripts/pm2-self-heal.sh` | same |
|
||||
| `search-stack-visibility` | `scripts/search-stack-check.py` | same |
|
||||
|
||||
### The deployed runner
|
||||
|
||||
The scheduler on **CT 100 (abiba)** runs contracts from
|
||||
**`/opt/contract-runner`** via `/etc/cron.d/contract-runner`. That clone is the
|
||||
pinned execution copy for every scheduled contract, and it must be kept current
|
||||
with `master` by fast-forward. Its `origin` is a local path to the upstream
|
||||
working copy, not a network remote.
|
||||
|
||||
`daily-health-digest` is **not** in the table above because it has no contract
|
||||
file and no mapping — it is dispatched by cron as
|
||||
`fm-send.sh ops "run contract: daily-health-digest"` and was, until
|
||||
2026-09-25, executed by hand from whichever clone the operator happened to be
|
||||
in. Creating its contract file and pinning it to a clone is an open follow-up.
|
||||
|
||||
## How the check works
|
||||
|
||||
`scripts/revision-preflight.sh <script-path> <clone-path>`:
|
||||
|
||||
* resolves the **repo-relative** path of the executing script inside the clone;
|
||||
* **fetches** the remote first, so a stale local ref cannot make a stale script
|
||||
look current — bounded by `--fetch-timeout` (default 20s) so a hung remote
|
||||
cannot block a scheduled contract;
|
||||
* compares the script's sha256 against `<ref>:<repo-relative-path>`;
|
||||
* **fails closed** — a path absent from the ref, an unresolvable ref, or a
|
||||
failed fetch is a failure, never a warning.
|
||||
|
||||
### Exit codes and reason classes
|
||||
|
||||
The guard distinguishes **"I could not check"** from **"this copy is wrong"**,
|
||||
and every non-zero exit prints a machine-readable `REASON=<class>` line before
|
||||
the human text, because a warning nobody can classify is not actionable — and
|
||||
the flip to `enforce` (below) depends on being able to read these apart.
|
||||
|
||||
| exit | `REASON=` | meaning |
|
||||
| --- | --- | --- |
|
||||
| 0 | — | verified match |
|
||||
| 2 | `cannot-verify:fetch-failed` | remote unreachable, failed, or timed out |
|
||||
| 2 | `cannot-verify:ref-unresolvable` | `<ref>` does not exist in the clone |
|
||||
| 1 | `mismatch:path-absent` | the script does not exist in `<ref>` |
|
||||
| 1 | `mismatch:content` | the script differs from `<ref>` |
|
||||
| 1 | `mismatch:detached-head` | the clone is on a detached HEAD |
|
||||
| 1 | `mismatch:clone-ahead` | local HEAD is strictly ahead of `<ref>` (mid-review) |
|
||||
|
||||
`detached-head` and `clone-ahead` are named separately on purpose: they are
|
||||
*legitimate* states that merely fail to be "the merged copy", and they are far
|
||||
less alarming than a hand-edited file. `clone-ahead` requires HEAD to be
|
||||
**strictly** ahead — an uncommitted edit on a commit that *is* the ref is a
|
||||
plain `content` mismatch.
|
||||
|
||||
## Modes in `contract-run.sh`
|
||||
|
||||
| `CONTRACT_REVISION_PREFLIGHT` | Behaviour |
|
||||
| --- | --- |
|
||||
| unset / **`warn` (default)** | log the refusal and its class, then still report |
|
||||
| `enforce` | withhold the verdict, alert, exit `2` |
|
||||
| `off` | skip the check entirely |
|
||||
|
||||
**The default is `warn`, deliberately.** The guard gates *every* scheduled
|
||||
contract, and three legitimate situations would otherwise turn the whole
|
||||
fleet's monitoring into withheld verdicts: a clone legitimately ahead of
|
||||
`origin/master` mid-review, a detached HEAD, and an offline or failed fetch.
|
||||
That is a bigger risk than the staleness the guard exists to catch. `warn`
|
||||
keeps the signal loud and classified in every run's log without letting the
|
||||
monitoring go dark.
|
||||
|
||||
### Criteria for flipping the default to `enforce`
|
||||
|
||||
Do not flip it on preference. Flip it when the evidence says the false-refusal
|
||||
rate is low enough, as its own small change with its own review:
|
||||
|
||||
1. the guard has run across **every scheduled contract** for a sustained period
|
||||
(suggested: 30 consecutive days, or 200+ contract runs) with **zero**
|
||||
`mismatch:*` and **zero** `cannot-verify:*` refusals in the per-run logs;
|
||||
2. no `cannot-verify:fetch-failed` arising from ordinary network blips in that
|
||||
window — if the pinned clone's remote is not reliably reachable, `enforce`
|
||||
will withhold rather than report;
|
||||
3. the pinned runner clone is demonstrably kept current by fast-forward, so
|
||||
`mismatch:clone-ahead` is a genuine fault rather than routine procedure.
|
||||
|
||||
The evidence for the flip is the `REASON=` lines already written into
|
||||
`/var/log/contract-runs/`. Until then the default stays `warn`.
|
||||
|
||||
## Merge-time sequence (do this whenever this repo merges)
|
||||
|
||||
**Baseline as of 2026-09-25:** `/opt/contract-runner` is already
|
||||
fast-forwarded to master `9faffe4`, so the pinned runner clone is current
|
||||
today. This sequence exists to keep it that way.
|
||||
|
||||
After any merge to `master`:
|
||||
|
||||
```bash
|
||||
# 1. fast-forward the pinned runner clone on CT 100
|
||||
git -C /opt/contract-runner pull --ff-only
|
||||
|
||||
# 2. confirm it is current and clean
|
||||
git -C /opt/contract-runner log --oneline -1
|
||||
git -C /opt/contract-runner status --porcelain # expect no output
|
||||
|
||||
# 3. prove a contract runs and reports normally
|
||||
CONTRACT_RUN_LOG_DIR=/tmp/preflight-proof \
|
||||
bash /opt/contract-runner/scripts/contract-run.sh search-stack-visibility
|
||||
echo "EXIT=$?" # expect 0, and 'revision-preflight: … matches origin/master'
|
||||
```
|
||||
|
||||
A contract that reports a `REASON=mismatch:*` refusal here means the runner
|
||||
clone is stale or locally edited — fast-forward it rather than reaching for
|
||||
`CONTRACT_REVISION_PREFLIGHT=off`.
|
||||
|
||||
**Note on untracked files:** git refuses to fast-forward over an untracked file
|
||||
even when its content is byte-identical to the incoming version
|
||||
(`The following untracked working tree files would be overwritten by merge`).
|
||||
A dirty clone will therefore block step 1. Resolve it by removing or stashing
|
||||
the untracked paths first — that is exactly what blocked a clone on 2026-09-25.
|
||||
|
||||
## Operating notes
|
||||
|
||||
* Under the default `warn`, a stale pinned clone still produces verdicts but
|
||||
every run logs the refusal and its class. Read those lines; do not ignore
|
||||
them.
|
||||
* Under `enforce`, a stale pinned clone **withholds**. That is the intended
|
||||
failure. Recover by fast-forwarding:
|
||||
`git -C /opt/contract-runner pull --ff-only`.
|
||||
* When a contract legitimately changes, land it through the normal branch + PR
|
||||
path and fast-forward the pinned clone. Do not copy files into it by hand.
|
||||
* `--no-fetch` exists for offline inspection; it prints that freshness is
|
||||
assumed rather than verified, and it is not used by `contract-run.sh`.
|
||||
@@ -1,5 +1,14 @@
|
||||
# Probe-drift round 2 — per-leg before/after evidence
|
||||
|
||||
> **Historical record** — 2026-09-28: The lines below that describe tanko as
|
||||
> "DSH (DeepSeek Harness)" only reflect what the check reported when it was
|
||||
> running. Tanko's runtime was later found to be **hybrid (DSH + Hermes)** —
|
||||
> the check had a `/root/` hardcoding bug that made it probe the wrong home
|
||||
> directory and report `wrapper-missing:tanko` for an agent with a working
|
||||
> wrapper. This document records the observed output, not the underlying
|
||||
> truth; see `fix/agent-health-root-hardcoding-20260928` for the correction.
|
||||
|
||||
|
||||
**Date:** 2026-09-10
|
||||
**Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`
|
||||
**Branch:** `fm/probe-drift-round2-20260909`
|
||||
|
||||
@@ -174,6 +174,29 @@ Key notes:
|
||||
|
||||
## Execution
|
||||
|
||||
### Executor, schedule and dead-man's-switch
|
||||
|
||||
This contract is implemented by a real script, scheduled on the inference host:
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **Executor** | `/opt/inference-harness/scripts/gpu-self-heal.py` on **CT 116** |
|
||||
| **Schedule** | `/etc/cron.d/gpu-self-heal` on CT 116 — `2 */6 * * *` |
|
||||
| **Log** | `/var/log/litellm/gpu-self-heal.log` |
|
||||
| **Posting** | the script calls `gitea-logger.sh gpu {RUN_ID}.json <report>` → `SyslogSolution/health-logs/gpu/{RUN_ID}.json` |
|
||||
|
||||
**Dead-man's-switch:** absence of logs must raise an alarm, because that is how
|
||||
this went silent for 12 days. That alarm cannot live on the producer — a stopped
|
||||
job cannot report that it stopped — so it lives off-host as the
|
||||
**`health-log-freshness` contract on CT 100**, which fails when
|
||||
`health-logs/gpu/` is older than 12 h. A failed run is also visible in the log
|
||||
above, but only the off-host check catches a *missing* run.
|
||||
|
||||
**History (2026-09-26, relay-785):** the executor was never lost — only its
|
||||
schedule was, dropped during a CT 116 `/etc/cron.d` rework on 2026-09-21. The
|
||||
`gpu/` log was silent from `2026-09-14T18:02:03Z`. The schedule was restored and
|
||||
the off-host freshness check added; do not treat either as optional.
|
||||
|
||||
```prose
|
||||
-- Phase 1: Fetch live GPU data
|
||||
let fleet = call gpu-monitor
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
---
|
||||
kind: function
|
||||
name: health-log-freshness
|
||||
description: >
|
||||
Dead-man's-switch for the health-logs posting jobs. Fails when the newest
|
||||
commit in a watched SyslogSolution/health-logs directory is older than that
|
||||
directory's threshold.
|
||||
|
||||
Exists because absence of logs raises no alarm: the gpu/ directory went silent
|
||||
for 12 days (2026-09-14T18:02:03Z to 2026-09-26) and nothing noticed, because
|
||||
the only thing that would have noticed was the job that had stopped. The check
|
||||
therefore runs on CT 100, a DIFFERENT host from the producers on CT 116, so a
|
||||
dead producer host or a deleted schedule still raises an alarm.
|
||||
|
||||
Watched: gpu/ (12h, producer 2 */6 * * *), litellm/ (18h, producer 0 */6 * * *).
|
||||
Not watched: pm2/ — that contract's health-logs claim was retired 2026-09-26;
|
||||
see pm2-self-heal.prose.md step 6.
|
||||
|
||||
Verified 2026-09-26: would have caught the real gap (at 2026-09-20 the newest
|
||||
gpu/ entry was 126h old against a 12h limit).
|
||||
|
||||
version: 1.0.0
|
||||
---
|
||||
|
||||
## Execution
|
||||
|
||||
Host-scheduled on CT 100 via `/etc/cron.d/contract-runner`:
|
||||
|
||||
```
|
||||
20 */4 * * * root CONTRACT_RUN_LOG_DIR=/var/log/contract-runs /bin/bash /opt/contract-runner/scripts/contract-run.sh health-log-freshness >/dev/null 2>&1 || /root/abiba-workspace/bin/fm-inbox.sh note "contract-runner: health-log-freshness FAILED - see /var/log/contract-runs/" >/dev/null 2>&1
|
||||
```
|
||||
|
||||
Every 4 hours, offset to `:20` to avoid the existing `:05`/`:15`/`:35` slots.
|
||||
|
||||
## Output shape
|
||||
|
||||
```
|
||||
Health-log freshness — dead-man's-switch
|
||||
==================================================================
|
||||
✅ health-logs/gpu/ newest 2026-09-26T14:48:03Z (0.02h old, limit 12.0h)
|
||||
last commit: gpu: gpu-self-heal-20260926-144802.json
|
||||
producer: gpu-self-heal.py, CT116 cron 2 */6 * * *
|
||||
==================================================================
|
||||
VERDICT: PASS — every watched health-log directory is advancing
|
||||
```
|
||||
|
||||
`--json` emits `{checked: {...}, failures: [...]}`.
|
||||
|
||||
## Exit codes
|
||||
|
||||
| exit | meaning |
|
||||
| --- | --- |
|
||||
| 0 | every watched directory is advancing |
|
||||
| 1 | at least one is stale, or could not be read |
|
||||
| 2 | the check could not run (no Gitea credential) |
|
||||
|
||||
A directory that **cannot be read** is a failure, not a skip: unreadable and
|
||||
stopped are indistinguishable from the outside.
|
||||
|
||||
## Configuration
|
||||
|
||||
| variable | default | meaning |
|
||||
| --- | --- | --- |
|
||||
| `GITEA_URL` | `https://git.sysloggh.net` | Gitea base URL |
|
||||
| `GITEA_TOKEN` / `GITEA_PAT` | — | API token; falls back to basic auth from `~/.git-credentials` |
|
||||
| `HEALTH_LOG_MAX_AGE_GPU` | `12` | hours |
|
||||
| `HEALTH_LOG_MAX_AGE_LITELLM` | `18` | hours |
|
||||
|
||||
## Thresholds
|
||||
|
||||
Sized for the producer cadence plus one missed run, so a single blip does not
|
||||
page but a genuine stop does:
|
||||
|
||||
* `gpu/` — 6 h cadence, 12 h limit;
|
||||
* `litellm/` — 6 h cadence, 18 h limit (proven healthy; a looser bound avoids noise).
|
||||
|
||||
## Maintains
|
||||
|
||||
- health-logs-gpu-freshness: { status: "ok|stale", last_check: timestamp }
|
||||
- health-logs-litellm-freshness: { status: "ok|stale", last_check: timestamp }
|
||||
@@ -70,10 +70,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
|
||||
|
||||
## Config Pattern — Mandatory Fields
|
||||
|
||||
### For Hermes Agents (Mumuni, Koonimo)
|
||||
### For Hermes Agents (Mumuni, Koonimo, Tanko-hybrid)
|
||||
|
||||
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
|
||||
(Tanko is hybrid — runs both DSH and Hermes since 2026-08-27, so its Hermes config is also checked.)
|
||||
|
||||
### 1. Main Model
|
||||
```yaml
|
||||
@@ -297,7 +297,7 @@ Run the consolidated health check:
|
||||
```bash
|
||||
python3 /root/scripts/agent-health-check.py
|
||||
```
|
||||
This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes),
|
||||
This validates each agent's live LiteLLM key against the gateway, including tanko, which runs HYBRID (DSH + Hermes) since 2026-08-27; detects GPU port conflicts (ghost processes),
|
||||
verifies gateway liveness, confirms Zulip streaming (`edit_message` present),
|
||||
and counts recent errors. Non-disruptive — never restarts anything.
|
||||
|
||||
|
||||
@@ -5,7 +5,13 @@ description: >
|
||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
|
||||
UPDATED 2026-09-27: Clarified the Auxiliary Tasks policy — light aux (vision,
|
||||
web_extract/browsing) -> gpu-vision (RTX 5070); context-heavy aux (compression) ->
|
||||
syslog-auto (2026-07-23 decision, Rule 7). Removed the false "one model for all
|
||||
auxiliary" / "never syslog-auto" claim; stated gpu-dense + strix-moe are the reasoning
|
||||
hosts and aux should not be pinned to them. Now matches audit-hermes-config.py line-for-line.
|
||||
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
|
||||
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
|
||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
||||
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
||||
@@ -145,6 +151,12 @@ mcp_servers:
|
||||
url: http://192.168.68.65:3100/mcp
|
||||
timeout: 120
|
||||
connect_timeout: 60
|
||||
litellm:
|
||||
url: https://litellm.sysloggh.net/mcp
|
||||
headers:
|
||||
x-litellm-api-key: "Bearer <AGENT_KEY>" # Rule 15: must be a REAL key (sk-...), not an env-var name
|
||||
# Note: MCP endpoint requires Accept: application/json, text/event-stream header
|
||||
# This is handled by the MCP client library; don't add to config
|
||||
|
||||
# ─── Compression ───
|
||||
compression:
|
||||
@@ -160,13 +172,16 @@ compression:
|
||||
abort_on_summary_failure: false
|
||||
|
||||
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
||||
# All auxiliary services MUST use identical model, base_url, and api_key_env:
|
||||
# model: gpu-vision # stable alias (NOT a raw model name)
|
||||
# Auxiliary tasks split into TWO model classes — do NOT assume one model for all:
|
||||
# Light auxiliary (vision, web_extract/browsing) -> model: gpu-vision # RTX 5070
|
||||
# Keeps the reasoning hosts (gpu-dense / strix-moe) free for agent prompts.
|
||||
# Context-heavy auxiliary (compression) -> model: syslog-auto # weighted pool
|
||||
# Deliberate per the 2026-07-23 OPERATIONAL DECISION in Rule 7: summarization
|
||||
# runs against long histories and must be able to use the pool.
|
||||
# Do NOT pin auxiliary work to the reasoning hosts (gpu-dense / strix-moe).
|
||||
# All auxiliary services share identical ROUTING (base_url + api_key_env), not model:
|
||||
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
# api_key_env: LITELLM_API_KEY
|
||||
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
||||
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
||||
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
|
||||
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
||||
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
|
||||
auxiliary:
|
||||
@@ -217,6 +232,40 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
||||
3. **After update**: Restart Hermes on the agent host
|
||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||
|
||||
## MCP Server Configuration
|
||||
|
||||
MCP server entries in `mcp_servers:` must follow the format shown in the Template section.
|
||||
|
||||
**Header requirements (Rule 15):**
|
||||
- Use `headers:` field with a `x-litellm-api-key` entry
|
||||
- The value must be `"Bearer <REAL_KEY>"` where `<REAL_KEY>` is a literal LiteLLM virtual key
|
||||
- Do NOT use env-var references like `$LITELLM_API_KEY` — they resolve to empty strings in
|
||||
the static config and cause "Malformed API Key" errors (2026-08-07 Tanko incident)
|
||||
|
||||
**Key source:**
|
||||
- Keys are stored in the Infisical vault (project=agents, env=production)
|
||||
- For template-based config generation: substitute the agent's key from the agent_keys table
|
||||
- For manual config updates: retrieve the key from the vault and insert the literal value
|
||||
|
||||
**Verification (2026-08-07):**
|
||||
- Tested MCP initialize handshake against litellm.sysloggh.net/mcp with agent virtual key
|
||||
- Confirmed: 200 response with `serverInfo.name: "litellm-mcp-server"`
|
||||
- Confirmed: tools/list returns 200 (MCP endpoint accessible with virtual keys)
|
||||
- Note: Per-key MCP grants are now supported (verified 2026-09-18), resolving the earlier
|
||||
contradiction with infrastructure-update.prose.md (which now reflects the update)
|
||||
- Key requirement: must be a valid LiteLLM virtual key (HTTP 200 on /v1/models)
|
||||
|
||||
**Key rotation note:**
|
||||
- MCP headers use literal keys (not env-vars), so they do NOT auto-rotate with the vault
|
||||
- After key rotation, MCP server headers must be regenerated with the new key value
|
||||
- This is a manual step: update the `x-litellm-api-key` header in each config file
|
||||
- TODO: Consider adding MCP header regeneration to the Key Update Procedure or a generation hook
|
||||
|
||||
**NetBird dependency:**
|
||||
- `litellm.sysloggh.net` is a NetBird endpoint (see Rule 5 and Infrastructure Stack table)
|
||||
- NetBird outages cause 502 errors on MCP requests, not auth failures
|
||||
- Diagnose: if MCP requests fail with 502, check NetBird status before investigating keys
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
@@ -455,14 +504,28 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
||||
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
||||
before and after any config change to catch this and all other rule violations.
|
||||
|
||||
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
|
||||
- Every MCP server entry must point at the correct endpoint:
|
||||
- ra-h-os = http://192.168.68.65:3100/mcp
|
||||
- litellm = https://litellm.sysloggh.net/mcp
|
||||
- MCP entries must carry a REAL key value in the header.
|
||||
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
|
||||
endpoints and result in "Malformed API Key" floods.
|
||||
- Ensure the header value is the actual key (e.g., `sk-...`).
|
||||
### Rule 15: MCP Endpoint and Header Validation (UPDATED 2026-08-07)
|
||||
|
||||
**Endpoint validation:**
|
||||
- ra-h-os must point to `http://192.168.68.65:3100/mcp`
|
||||
- litellm must point to `https://litellm.sysloggh.net/mcp`
|
||||
- Mismatched endpoints cause silent failures (e.g., 2026-08-07 incident: Tanko's config had
|
||||
ra-h-os pointing to litellm's endpoint)
|
||||
|
||||
**Header validation:**
|
||||
- Every MCP entry with authentication must carry a `headers:` field
|
||||
- The header value must be a REAL key (e.g., `Bearer sk-abc123...`), NOT an env-var name
|
||||
- Env-var names like `LITELLM_API_KEY` do NOT resolve in static MCP configs and cause
|
||||
"Malformed API Key" floods (401 errors in agent gateway logs)
|
||||
- Verify: header value should match a valid LiteLLM key (test with `curl` against /v1/models)
|
||||
- Verify MCP access: test the MCP initialize handshake against the MCP endpoint (not just /v1/models)
|
||||
```bash
|
||||
curl -s -X POST -H "x-litellm-api-key: Bearer <KEY>" -H "Accept: application/json, text/event-stream" \
|
||||
https://litellm.sysloggh.net/mcp -d '{"jsonrpc":"2.0","id":1,"method":"initialize",...}' \
|
||||
| jq '.data.result.serverInfo' # should show serverInfo.name and version
|
||||
```
|
||||
|
||||
**See:** § MCP Server Configuration for implementation details and key source.
|
||||
|
||||
## Execution
|
||||
|
||||
|
||||
@@ -16,7 +16,26 @@ author: Abiba (pi agent)
|
||||
|
||||
## Rule (One Sentence)
|
||||
|
||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with authenticated path `http://192.168.68.116/litellm/v1/responses` — hardcoded keys AND unauthenticated `/v1` direct access are both forbidden.**
|
||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
|
||||
|
||||
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
|
||||
|
||||
## Model Access Tiers (2026-09-20)
|
||||
|
||||
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
|
||||
|
||||
| Tier | Models | Who gets it |
|
||||
|------|--------|-------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
|
||||
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
|
||||
|
||||
**Rules:**
|
||||
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
|
||||
2. Agent keys MUST carry an **explicit local-only** `models` list.
|
||||
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
|
||||
4. The master key always bypasses scoping — it is admin-only, never for inference.
|
||||
|
||||
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
|
||||
|
||||
## Scope
|
||||
|
||||
@@ -38,8 +57,8 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
|
||||
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
|
||||
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
|
||||
|
||||
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid).
|
||||
Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**.
|
||||
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` non-canonical but working).
|
||||
The public host `https://litellm.sysloggh.net` serves `/v1` ONLY (404 on `/litellm/v1`).
|
||||
|
||||
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
|
||||
|
||||
@@ -101,7 +120,7 @@ auxiliary:
|
||||
fallback_providers:
|
||||
- provider: deepseek
|
||||
base_url: https://api.deepseek.com
|
||||
api_key: sk-b7d9... # ← hardcoded OK (external)
|
||||
api_key: sk-synthetic-external-example # ← hardcoded OK (external, synthetic example)
|
||||
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
|
||||
```
|
||||
|
||||
@@ -109,11 +128,11 @@ fallback_providers:
|
||||
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
|
||||
model:
|
||||
provider: harness
|
||||
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION: hardcoded key
|
||||
api_key: sk-synthetic-example-12345 # ← RULE VIOLATION: hardcoded key (synthetic example)
|
||||
|
||||
model:
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/v1 # ← RULE VIOLATION: unauthenticated path
|
||||
base_url: http://192.168.68.116/v1 # ← NON-CANONICAL but WORKING (authenticated via nginx, WARN not FAIL)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
```
|
||||
|
||||
@@ -139,6 +158,16 @@ scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### ACCEPTABLE PATTERN
|
||||
Agent keys live in `.env` or `.env.vault` files with 600 permissions (koonimo's shape is the canonical example). A plaintext key inside a `config.yaml` or any `config.yaml.bak-*` file is a violation — the backup files are not part of the runtime credential path and are not watched by the scanner, so a key in them is stale clutter that a future reader can mistake for a working key.
|
||||
|
||||
**Fix procedure** (when a backup file is found with a plaintext key):
|
||||
1. Move the file out of the scanned tree (e.g., `mv /root/.hermes/config.yaml.bak-* /root/hermes-config-backups/`) — do NOT delete the file, just move it so the scanner pattern no longer matches.
|
||||
2. Re-run the reachability check to confirm COMPLIANT.
|
||||
3. Report the before/after check output and the commands you ran.
|
||||
|
||||
**Rationale**: Moving the file preserves history without leaving a credential where a scanner trips over it. Deleting the file loses the historical context. Keeping it in place means the next scan will report it as a finding and waste time re-deciding.
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
@@ -172,7 +201,7 @@ grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
|
||||
|
||||
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
||||
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /root/.config/systemd/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-…' /root/.config/systemd/ 2>/dev/null
|
||||
|
||||
# 3. Verify running process env matches dedicated key
|
||||
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
||||
|
||||
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
|
||||
| Param | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) |
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — hybrid (DSH + Hermes) since 2026-08-27, no Hermes plugin) |
|
||||
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
||||
|
||||
## Maintains
|
||||
@@ -55,7 +55,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
|
||||
| Tanko | CT112 | minipve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(hybrid (DSH + Hermes) since 2026-08-27 — historical, plugin retired on this host)* |
|
||||
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||
|
||||
@@ -72,7 +72,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
### Step 1: Resolve Target
|
||||
|
||||
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
|
||||
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||
For CT112 route through `ssh root@minipve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||
|
||||
### Step 2: Pull Latest Plugin Source
|
||||
|
||||
@@ -121,7 +121,7 @@ cp plugins/platforms/zulip/adapter.py \
|
||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
# Fix ownership (was Tanko-only, runs as jerome user)
|
||||
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
|
||||
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (hybrid: DSH + Hermes).
|
||||
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ gateway restart, and connection validation.
|
||||
|
||||
| Param | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) |
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — hybrid (DSH + Hermes) since 2026-08-27) |
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -67,7 +67,7 @@ gateway restart, and connection validation.
|
||||
### Step 1: Locate Target
|
||||
|
||||
Map `target` to connectivity parameters from the live-state table above.
|
||||
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||
For CT112 route through `ssh root@minipve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||
|
||||
### Step 2: Deploy Zulip Adapter
|
||||
|
||||
@@ -93,7 +93,7 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
|
||||
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
||||
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin)
|
||||
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on hybrid (DSH + Hermes), no Hermes plugin)
|
||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
|
||||
|
||||
# Clean up
|
||||
|
||||
@@ -105,8 +105,8 @@ description: >
|
||||
|
||||
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
||||
|------|----|-----|-----|---------|------|
|
||||
| minipve | .12 | 16C | 30GB | abiba, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||
| amdpve | .15 | 32C | 62GB | kagentz, tanko, baggy, scottdenya, adguard2 | Agents, compute |
|
||||
| minipve | .12 | 16C | 30GB | abiba, tanko, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||
| amdpve | .15 | 32C | 62GB | kagentz, baggy, scottdenya, adguard2 | Agents, compute |
|
||||
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
|
||||
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
|
||||
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
||||
@@ -181,8 +181,8 @@ description: >
|
||||
**Stirling-PDF** (deployed 2026-07-03, Authentik SSO 2026-07-03):
|
||||
- URL: `https://pdf.sysloggh.net` (public) / `http://192.168.68.7:8989` (direct)
|
||||
- Swagger: `http://192.168.68.7:8989/swagger-ui.html`
|
||||
- Admin credentials: `admin` / `kakashi20stirling`
|
||||
- API key: `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88`
|
||||
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_USER»` / `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||
- API key: `«vault: infrastructure/production STIRLING_API_KEY»`
|
||||
- Authentik OAuth2: configured but disabled (requires paid Server license). Ready to enable: set `SECURITY_OAUTH2_ENABLED=true` + `SECURITY_LOGINMETHOD=all`
|
||||
- Compose: `/opt/home_stack/docker-compose.yml`
|
||||
- Control script: `/opt/home_stack/infra-control.sh`
|
||||
@@ -364,6 +364,79 @@ For docker-vm specifically:
|
||||
- No PBS backup in 48h → fail
|
||||
```
|
||||
|
||||
### Backup Safety Preconditions (2026-09-15)
|
||||
|
||||
#### Background & Rationale
|
||||
|
||||
Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first:
|
||||
|
||||
1. **acerpve thin-pool VM 101** (acerpve, 192.168.68.9, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root.
|
||||
|
||||
2. **amdpve 0700 tmpdir** (amdpve, 192.168.68.15, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup.
|
||||
|
||||
3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood.
|
||||
|
||||
> ⚠️ **Hostname Resolution Warning (2026-09-15)**: The PVE node hostnames (acerpve, amdpve, minipve, storepve, ocupve) all resolve to the VPS (72.61.0.17, the Netbird VPS at srv1079750.hstgr.cloud) via the wildcard `*.dns.sysloggh.net` record, NOT to the actual nodes. So `ssh acerpve` lands on the VPS. **Nodes must be addressed by IP**: acerpve 192.168.68.9, amdpve 192.168.68.15, storepve 192.168.68.6, minipve 192.168.68.12, ocupve 192.168.68.5. Guest CTs are reached through their node (`pct exec`). Guest hostnames that resolve on the LAN (e.g. kagentz = 192.168.68.14) are fine. (The DNS address records are a separate decision — row: dag-daemon-node-hostnames-resolve-to-the-vps-20260915.)
|
||||
|
||||
#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts)
|
||||
|
||||
Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass:
|
||||
|
||||
```bash
|
||||
# Check 1: Pool headroom (PRIMARY - yields percentages directly)
|
||||
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname, see warning above
|
||||
lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data
|
||||
# Example output (acerpve, 192.168.68.9):
|
||||
# LV Data% Meta% LSize
|
||||
# data 29.95 1.22 <816.21g
|
||||
# Required thresholds (documented minimum):
|
||||
# data_percent < 90% (80% recommended for safety margin)
|
||||
# metadata_percent < 70% (metadata fills faster than data)
|
||||
|
||||
# Check 2: Verify pool is not in error state (dmsetup shows the raw DM device)
|
||||
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname
|
||||
dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1
|
||||
# dmsetup status pve-data-tpool field order (verified on 192.168.68.9):
|
||||
# $1=start $2=length $3="thin-pool" $4=transaction-id
|
||||
# $5=metadata_used/metadata_total (blocks) $6=data_used/data_total (sectors)
|
||||
# remaining fields are flags ("-", "rw", "discard_passdown", "queue_if_no_space", ...)
|
||||
# This is only used for the ERROR-STATE check; use the lvs command above for percentages.
|
||||
# metadata_percent = $5 / ($5 split by /) [second number in pair]
|
||||
```
|
||||
|
||||
**Minimum thresholds**: If either `data_percent >= 90%` or `metadata_percent >= 70%`, the backup MUST NOT start. State explicitly that these are hard stops, not warnings.
|
||||
|
||||
**Why this is a precondition**: A full or errored thin pool fails EVERY volume on the VG at once, including the host root. This is not a soft failure - it takes down the entire Proxmox host.
|
||||
|
||||
#### Staging Directory Requirement (2026-09-14 incident)
|
||||
|
||||
Any custom vzdump `tmpdir` MUST be world-traversable and writable exactly like `/var/tmp` (mode 1777). The archive step of vzdump runs in an unprivileged user namespace and cannot traverse a root-owned 0700 directory.
|
||||
|
||||
**Symptom to recognize**: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES` on every container in the backup run.
|
||||
|
||||
**Fix**: `chmod 1777 <custom-tmpdir>` before starting vzdump.
|
||||
|
||||
#### Task Start Rule for Truncating Shells
|
||||
|
||||
When starting a backup task from a shell that may truncate output (e.g., pipes, `head`), always use:
|
||||
|
||||
```bash
|
||||
pvesh create /storage/backup --output-format json -- ... | head -2
|
||||
# ❌ Can kill the backup task ("broken pipe" status)
|
||||
```
|
||||
|
||||
Instead, capture JSON output without piping to truncating commands:
|
||||
|
||||
```bash
|
||||
# Use --output-format json and capture to variable
|
||||
result=$(pvesh create /storage/backup --output-format json -- ...)
|
||||
# Then parse result if needed
|
||||
```
|
||||
|
||||
#### GPU Host Backup Status (acerpve VM 101)
|
||||
|
||||
VM 101 (llm-gpu) and VM 103 (ocu-llm) have NO scheduled backup job. Coverage is manual one-off runs only. This is intentional for VM 101 until the thin-pool failure mechanism is understood and documented.
|
||||
|
||||
## Section 5: Network Services — Monitoring
|
||||
|
||||
### 5.1 Service Inventory
|
||||
@@ -563,7 +636,7 @@ monitor, or integration breaks.
|
||||
```bash
|
||||
# Full cluster status
|
||||
PVE="https://minipve.sysloggh.net"
|
||||
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
AUTH="Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
|
||||
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
|
||||
|
||||
# Docker health from Abiba
|
||||
@@ -609,7 +682,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
||||
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
|
||||
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
|
||||
| 112 | tanko | minipve | .122 | hybrid (DSH + Hermes) agent | ✅ |
|
||||
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
||||
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
||||
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
||||
@@ -639,7 +712,7 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
|
||||
| 100 | abiba | minipve | `pct-run 100` |
|
||||
| 105 | kagentz | amdpve | `pct-run 105` |
|
||||
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
|
||||
| 112 | tanko | amdpve | `pct-run 112` |
|
||||
| 112 | tanko | minipve | `pct-run 112` |
|
||||
| 113 | baggy | amdpve | `pct-run 113` |
|
||||
| 115 | scottdenya | amdpve | `pct-run 115` |
|
||||
| 104 | authentik | minipve | `pct-run 104` |
|
||||
@@ -686,4 +759,4 @@ which was kill+nohup outside systemd) are banned by policy.
|
||||
| Script | Why Disabled |
|
||||
|--------|-------------|
|
||||
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
|
||||
| `zulip-monitor.sh` (Abiba) | Replaced by agent-health-check.py + PM2 auto-restart |
|
||||
| `zulip-monitor.sh` (Abiba) | Standalone cron replaced by agent-health-check.py + PM2 auto-restart; the script itself remains active as the execution step of the `zulip-health` responsibility contract (see `zulip-health.prose.md`) |
|
||||
|
||||
@@ -102,6 +102,13 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
||||
- Stack persists across reboots (systemd for exporters, Docker restart policy)
|
||||
|
||||
## Execution
|
||||
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||
execution — it only proves the agent read the result and reported it. The actual
|
||||
monitoring work happens in the host cron job.
|
||||
|
||||
|
||||
### Liveness rule (scoped)
|
||||
|
||||
@@ -138,11 +145,19 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
||||
tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
**EXECUTABLE OWNER:** The canonical probe set lives in `scripts/infra-monitoring.sh`.
|
||||
A check run is a single command: `bash scripts/infra-monitoring.sh` (from the
|
||||
repository root). Paste its raw output verbatim into the report. The script
|
||||
exits non-zero naming every failed target; there is no "OK" summary when any
|
||||
leg failed. Port drift is caught by `scripts/test_infra_monitoring.sh` which
|
||||
asserts every probed port matches the documented value.
|
||||
|
||||
**PROBE SHAPE (per standing rules above):**
|
||||
- Every probe prints the target name + URL + HTTP code (or failure kind)
|
||||
- Retry once on connection failure at longer timeout
|
||||
- Every probe prints: `✅ <name>: alive` on success, or `🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)` on failure
|
||||
- PVE API failures include `(any-HTTP liveness, -k for self-signed)` to distinguish TLS vs connection
|
||||
- Retry once on connection failure at longer timeout (25s connect, 30s max)
|
||||
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
||||
- Report the actual probe command and its result, not a summary verdict
|
||||
- Report the actual probe output, not a summary verdict
|
||||
|
||||
```bash
|
||||
# Provenance — run first; paste the absolute path into the report
|
||||
@@ -269,6 +284,27 @@ code (or failure kind with retry details). Apply the standing probe rules: any
|
||||
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
||||
a failure.
|
||||
|
||||
### Docker Stats and PVE Exporter Ports
|
||||
|
||||
These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:
|
||||
|
||||
| Exporter | Port | Container | Metrics |
|
||||
|----------|------|-----------|---------|
|
||||
| **Docker Stats** | **9324** | harness-docker-stats | `docker_container_*` (per-container CPU/mem/network) |
|
||||
| **PVE Exporter** | **9221** | harness-pve-exporter | `pve_*` (5 cluster-level metrics) |
|
||||
|
||||
**IMPORTANT**: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (`builder_builds_*`, `containerd_build_info_*`).
|
||||
|
||||
```bash
|
||||
# Docker Stats (harness-docker-stats)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
|
||||
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||
|
||||
# PVE Exporter (harness-pve-exporter)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
|
||||
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||
```
|
||||
|
||||
### Phase 1: GPU Exporters
|
||||
|
||||
**NVIDIA (.8 and .110)**:
|
||||
|
||||
@@ -59,7 +59,7 @@ Before ANY update wave:
|
||||
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 112 (tanko, minipve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
||||
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
||||
@@ -208,19 +208,19 @@ mcp_servers:
|
||||
| Key | MCP Access |
|
||||
|-----|-----------|
|
||||
| Master key | ✅ Full — 90 tools (vault-injected) |
|
||||
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.99.1 |
|
||||
| Agent keys (mumuni, tanko, etc.) | ✅ Per-key grants supported (as of 2026-09-18 verification) |
|
||||
|
||||
### Known Limitations
|
||||
- Per-key MCP server grants not functional — only master key has access
|
||||
- ~~Per-key MCP server grants not functional — only master key has access~~ (resolved 2026-09-18: per-key grants now work)
|
||||
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
|
||||
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
|
||||
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
|
||||
|
||||
### Migration Path
|
||||
When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
||||
1. Grant agent keys `mcp_servers: ["ra_h_os"]`
|
||||
2. Update Hermes `mcp_servers.ra-h-os.url` from `http://192.168.68.65:3100/mcp` → `http://192.168.68.116:4000/mcp/`
|
||||
3. Add `headers: {x-litellm-api-key: "Bearer $LITELLM_API_KEY"}` to MCP config
|
||||
### Migration Path (COMPLETED 2026-09-18)
|
||||
Per-key MCP grants are now supported:
|
||||
1. ✅ Agent keys granted MCP access via `allowed_mcp_servers` field
|
||||
2. ✅ Hermes `mcp_servers.litellm.url` set to `https://litellm.sysloggh.net/mcp`
|
||||
3. ✅ `headers: {x-litellm-api-key: "Bearer <literal_key>"}` added to MCP config
|
||||
|
||||
## Security-Specific Updates
|
||||
|
||||
|
||||
@@ -71,6 +71,10 @@ description: >
|
||||
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
||||
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
||||
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
||||
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
|
||||
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
|
||||
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
|
||||
In LiteLLM Community both silently grant access to EVERY model, including cloud.
|
||||
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
||||
- Return the new key
|
||||
5. **If action == "rotate"**:
|
||||
@@ -87,6 +91,66 @@ description: >
|
||||
- Confirm key alias matches agent_name in LiteLLM key list
|
||||
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
||||
|
||||
## Cloud Provider Consolidation (2026-09-20)
|
||||
|
||||
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
|
||||
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
|
||||
|
||||
### Provider map (per-account namespacing)
|
||||
|
||||
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
|
||||
stay separate:
|
||||
|
||||
| Prefix | Upstream | Auth | Vault secret |
|
||||
|--------|----------|------|--------------|
|
||||
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
|
||||
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
|
||||
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
|
||||
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
|
||||
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
|
||||
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
|
||||
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
|
||||
|
||||
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
|
||||
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
|
||||
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
|
||||
|
||||
### Access tiers (MUST be enforced per key)
|
||||
|
||||
| Tier | Model names | Granted to |
|
||||
|------|-------------|------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
|
||||
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
|
||||
|
||||
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
|
||||
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
|
||||
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
|
||||
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
|
||||
|
||||
### Creating the cloud-enabled key
|
||||
|
||||
```bash
|
||||
# ALWAYS read the live roster first (key-scoped):
|
||||
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
|
||||
| jq -r '.data[].id'
|
||||
|
||||
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
|
||||
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
|
||||
```
|
||||
|
||||
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
|
||||
any `<prefix>/` cloud model.
|
||||
|
||||
### Adding a new cloud provider
|
||||
|
||||
1. Add the upstream key to Infisical `infrastructure/production/root`.
|
||||
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
|
||||
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
|
||||
3. Restart the `harness-litellm` container.
|
||||
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
|
||||
captain approval.
|
||||
5. Update this table and the access-tier section.
|
||||
|
||||
## Production Vault Access Process (canonical, 2026-07-17)
|
||||
|
||||
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
||||
@@ -144,8 +208,8 @@ through its agent wrapper.
|
||||
safety net for vault outage or token revocation. Must be kept in sync on rotation.
|
||||
Example:
|
||||
```bash
|
||||
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
|
||||
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
|
||||
MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY»
|
||||
MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY»
|
||||
```
|
||||
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
|
||||
```ini
|
||||
@@ -191,7 +255,7 @@ through its agent wrapper.
|
||||
### Tanko migration (COMPLETED 2026-07-17)
|
||||
|
||||
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
|
||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
|
||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-synthetic-tanko-example…`)
|
||||
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
|
||||
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
|
||||
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
|
||||
@@ -271,13 +335,13 @@ not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update ma
|
||||
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
||||
|
||||
**Key Storage:**
|
||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
|
||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»…`)
|
||||
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
||||
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
||||
(unlike fleet agents which require vault injection)
|
||||
|
||||
**Current Key (2026-09-01):**
|
||||
- **Prefix**: `sk-or-v1-0af3f3…`
|
||||
- **Prefix**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
- **Plan**: Paid (not free tier)
|
||||
- **Usage**: 0 (as of 2026-09-01)
|
||||
@@ -322,10 +386,10 @@ directly call OpenRouter via Python's requests library. Converting would require
|
||||
|
||||
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
|
||||
```bash
|
||||
# Read at runtime from the container's environment:
|
||||
# PRIMARY (proven, runs on CT 116 with no extra tooling):
|
||||
docker exec harness-litellm printenv LITELLM_MASTER_KEY
|
||||
# Or from Infisical vault (project=infrastructure env=prod) - NOTE: --plain is broken on CLI 0.43.110 (prints nothing):
|
||||
infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production | awk '$1=="LITELLM_MASTER_KEY"{print $NF}'
|
||||
# Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching)
|
||||
# The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it)
|
||||
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
|
||||
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
|
||||
```
|
||||
|
||||
@@ -19,7 +19,7 @@ description: >
|
||||
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
|
||||
- Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).
|
||||
---
|
||||
|
||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||
@@ -111,6 +111,13 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
||||
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
|
||||
|
||||
## Execution
|
||||
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||
execution — it only proves the agent read the result and reported it. The actual
|
||||
monitoring work happens in the host cron job.
|
||||
|
||||
|
||||
1. **Read parameters** — Use provided values or defaults
|
||||
|
||||
|
||||
@@ -116,7 +116,7 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (hardcoded key → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM (deprecated key, no live usage).
|
||||
|
||||
## Maintains
|
||||
|
||||
|
||||
+44
-6
@@ -6,13 +6,19 @@ name: memory-fixer
|
||||
description: >
|
||||
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
||||
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||
version: 2.1.0
|
||||
version: 2.2.0
|
||||
---
|
||||
---
|
||||
|
||||
# Memory Fixer
|
||||
|
||||
> **Canonical copy:** `/root/.hermes/contracts/memory-fixer-v3.md` (used by the `memory-fixer-daily` cron job). This file is the institutional record of the same contract. When the two diverge, treat the v3 source in `/root/.hermes/contracts/` as executable truth.
|
||||
> **Executable copy:** the `okyeame-memory-fixer` cron job on kagentz (`hermes cron list`) holds its instruction
|
||||
> set **inline in `~/.hermes/cron/jobs.json`** (`hermes cron edit <id> --prompt …`; there is no `--prompt-file`, and
|
||||
> `~/.hermes/cron/memory-fixer-prompt.md` is a synced draft, not the live instruction). This file is the institutional
|
||||
> record of the same contract; when the two diverge, the job prompt is what actually runs — diff it against this file
|
||||
> before claiming a prompt change landed.
|
||||
> ⚠️ Corrected 2026-09-26: the previous pointer (`/root/.hermes/contracts/memory-fixer-v3.md`) does not exist on
|
||||
> kagentz — no `/root` access from this container — and was verified unreachable, not merely stale.
|
||||
|
||||
## Purpose
|
||||
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever.
|
||||
@@ -125,15 +131,44 @@ updateNode(id, {
|
||||
|
||||
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
|
||||
|
||||
### 5. Duplicate-Node Detection (Level 1 — read-only, every run)
|
||||
|
||||
The graph's duplicate problem is rarely an agent mistyping a title: it is **recurring writers creating a new
|
||||
node per run instead of updating one**. This phase detects that class and reports it. It is read-only and
|
||||
**never merges**.
|
||||
|
||||
```bash
|
||||
python3 /home/hermes/.hermes/scripts/memory_dup_detect.py --json
|
||||
```
|
||||
Read-only, ~15s over the whole graph, exit 0. That script is the source of truth for the clustering logic —
|
||||
do not re-implement it in the prompt or hand-count "duplicates" from titles.
|
||||
|
||||
Consume each `items[]` entry's `verdict` field; do not invent your own:
|
||||
|
||||
| `verdict` | Meaning | Required action |
|
||||
|---|---|---|
|
||||
| `WRITER-DEFECT` (`run_family: true`) | ONE scheduled task writes a new node per run | Report the ids, the `agents` (the writer) and `span_days`. **Never merge** — each node is that run's audit record. If the family grew since the last report, say `UNFIXED` and name the writer. |
|
||||
| `SAFE-MERGE` | Bodies identical | Still requires an explicit `merge #A into #B` decision from Kwame. |
|
||||
| `HUMAN-DECISION` | Same subject, bodies differ | Propose **connect (an edge)**, never merge. |
|
||||
|
||||
- **Title overlap alone is not duplication.** Four distinct client workflows of one family (#357-#361) and two
|
||||
different machines' migrations (#1792/#1793) both score high on title tokens while their bodies sit 0.1-0.3
|
||||
apart. Confirm against body similarity before calling anything a duplicate.
|
||||
- Report clusters as **candidates for Kwame's decision**, never as established duplicates — a wrong auto-merge
|
||||
destroys distinct content irrecoverably.
|
||||
- Per-run history nodes are kept deliberately. Bulk-merging a run family destroys the audit trail the family exists for.
|
||||
|
||||
## Level 2 Escalations (Kwame Decision Required)
|
||||
|
||||
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
|
||||
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
|
||||
2. **Duplicate Nodes** — as detected by fix 5, by `verdict`, never by raw title overlap. `WRITER-DEFECT` is a writer fix (update one canonical node), not a merge decision; `SAFE-MERGE` and `HUMAN-DECISION` clusters are escalated for merge-or-connect.
|
||||
3. **Orphan Nodes >90 days old** — Archive or connect?
|
||||
|
||||
## Reporting Format
|
||||
|
||||
The fixer reports to Kwame via this Zulip DM:
|
||||
The fixer does **not** send anything. Under the single-egress model (2026-09-21) every report leaves the node
|
||||
through Mumuni's gate (`comms_drop.py` for the queue, `comms_gate.py` to release and read-back verify), so
|
||||
exit 0 means QUEUED, never delivered. A report body is written to a file and handed to the outbox helper:
|
||||
|
||||
```
|
||||
🦅 Memory Fixer — [HH:MM UTC]
|
||||
@@ -147,8 +182,10 @@ Stale nodes needing review (max 10):
|
||||
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
|
||||
...
|
||||
|
||||
Duplicates needing decision:
|
||||
1. [Node #AAA] vs [Node #BBB] — Same title
|
||||
Duplicate clusters (candidates — Kwame decides; the fixer never merges unilaterally):
|
||||
1. [WRITER-DEFECT] #AAA/#BBB/#CCC — writer <agent>, N nodes, span Nd (UNFIXED if it grew since the last report)
|
||||
2. [HUMAN-DECISION] #DDD/#EEE — same subject, bodies differ, SUGGEST: connect
|
||||
3. "none" when the scan returned no clusters
|
||||
|
||||
Orphans >90 days:
|
||||
1. [Node #EEE] Title — X days stale, orphaned
|
||||
@@ -195,6 +232,7 @@ The result must be 0 rows when all decisions are executed. Report what was done.
|
||||
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
||||
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
|
||||
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
||||
- **Duplicate scan ran:** every report carries the fix 5 block (`none` when there were no clusters). A report with no duplicate section means phase 5 was skipped — a silently skipped detection phase is the failure this phase exists to prevent.
|
||||
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
|
||||
|
||||
## Logging
|
||||
|
||||
+39
-15
@@ -2,16 +2,10 @@
|
||||
kind: responsibility
|
||||
name: pm2-self-heal
|
||||
description: >
|
||||
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
||||
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
||||
errored. Logs every action to the knowledge graph and alerts the owner via
|
||||
Zulip DM on failures.
|
||||
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
||||
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
||||
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
||||
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
||||
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
||||
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
||||
Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
|
||||
and auto-restarts any that are stopped or errored. Logs every action to
|
||||
Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary).
|
||||
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -22,6 +16,8 @@ description: >
|
||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||
- last_check: timestamp
|
||||
|
||||
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed.
|
||||
|
||||
|
||||
## Continuity
|
||||
|
||||
@@ -45,6 +41,9 @@ description: >
|
||||
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
||||
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
||||
spoton incident). Alerts include the restart count.
|
||||
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
|
||||
four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton
|
||||
reference above is historical context for the crash-loop guard, not a live process.
|
||||
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
||||
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
||||
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
||||
@@ -52,19 +51,44 @@ description: >
|
||||
and PM2 counter reset on 2026-06-28.
|
||||
|
||||
## Execution
|
||||
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||
execution — it only proves the agent read the result and reported it. The actual
|
||||
monitoring work happens in the host cron job.
|
||||
|
||||
|
||||
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
||||
2. **Check abiba-telegram**:
|
||||
2. **Check abiba-telegram** (safe to auto-restart):
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
|
||||
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
||||
- If status is "online" → pass, log restarts count
|
||||
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
6. **Wait 5 min** → repeat from step 1
|
||||
4. **Check gitea-runner**:
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
5. **Check zulip-watchdog**:
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
6. **Log results** — the durable per-run record is `/var/log/contract-runs/pm2-self-heal-<UTCstamp>.log` on CT 100, written by `scripts/contract-run.sh` from `/etc/cron.d/contract-runner` every 4 hours, with a firstmate inbox note raised on any non-zero exit.
|
||||
|
||||
**CORRECTED 2026-09-26 (relay-785):** this step previously required appending to
|
||||
`SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea and called it a "hard rule".
|
||||
That posting was **never implemented** — `scripts/pm2-self-heal.sh` contains no
|
||||
Gitea or git-push code — so the requirement was a coverage claim the executor did
|
||||
not honour, and `health-logs/pm2/` has held only its init commit since 2026-07-28.
|
||||
The claim is retired rather than implemented: the contract-runner's per-run logs
|
||||
plus its failure note already give a durable record and a working alarm, and a
|
||||
second posting path would add work without adding a signal. `health-logs/pm2/`
|
||||
is left as historical evidence, not as a live obligation.
|
||||
7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
8. **Wait 5 min** → repeat from step 1
|
||||
|
||||
## Example Output (when healthy)
|
||||
|
||||
|
||||
@@ -89,6 +89,48 @@ agent: abiba
|
||||
| minipve | 192.168.68.12 | PVE |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
||||
|
||||
## PBS GC (Proxmox Backup Server)
|
||||
|
||||
### Schedule
|
||||
Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**.
|
||||
|
||||
**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.
|
||||
|
||||
### What Actually Runs
|
||||
Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`.
|
||||
|
||||
**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC.
|
||||
|
||||
### Datastore Location
|
||||
- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free)
|
||||
- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume)
|
||||
|
||||
### Liveness Check
|
||||
The new `proxmox-monitor.sh` leg checks storepve-datastore GC health:
|
||||
- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json`
|
||||
- **FAILS** if `last-run-endtime` is older than 48 hours
|
||||
- Reports age in hours and pending-bytes
|
||||
|
||||
**All six verdict shapes** (exactly as emitted by the script):
|
||||
|
||||
1. **Healthy** (fresh GC, 0 B pending):
|
||||
`✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)`
|
||||
|
||||
2. **Stale** (GC ran >48h ago):
|
||||
`🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)`
|
||||
|
||||
3. **Probe-failed: empty read** (000/timeout/unreadable):
|
||||
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)`
|
||||
|
||||
4. **Probe-failed: unparseable** (non-empty but invalid JSON — the "command not found" case):
|
||||
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)`
|
||||
|
||||
5. **Never-run: datastore absent** (valid JSON but storepve-datastore not in list):
|
||||
`🔴 PBS GC: never-run (storepve-datastore not found in GC list)`
|
||||
|
||||
6. **Never-run: no endtime** (valid JSON with datastore present but last-run-endtime is null/0):
|
||||
`🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)`
|
||||
|
||||
## Operations
|
||||
|
||||
### view-dashboards
|
||||
|
||||
@@ -50,6 +50,11 @@ Changelog:
|
||||
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
|
||||
from her side. This script must not probe mumuni or .24 — the v2 changelog
|
||||
roster line was the last reference still placing her at .24 / CT100.
|
||||
v6 (2026-09-28): .8 GPU health probe now runs as `llmuser` instead of `root`.
|
||||
Root SSH to .8 was lost when the guest was rebuilt, so every .8 leg read as
|
||||
UNREACHABLE for a healthy host. llmuser owns llama-server and can read
|
||||
`systemctl is-active`, `systemctl show -p MainPID`, and the :8080 pid.
|
||||
.110 and .15 keep the default `root` user.
|
||||
"""
|
||||
|
||||
import subprocess, json, sys, os, time, re, io, contextlib
|
||||
@@ -70,7 +75,7 @@ PVE_NODES = {
|
||||
|
||||
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
||||
AGENTS = {
|
||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
|
||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "minipve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "hybrid"},
|
||||
# abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its
|
||||
# local env file (key_env below), not from the shared vault or .bashrc.
|
||||
# runtime=pi: abiba has run pi-only since the harness purge. There is no
|
||||
@@ -95,7 +100,7 @@ AGENTS = {
|
||||
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
|
||||
# .15 strixhalo (amdpve) -> strix-server.service (active)
|
||||
GPU_HOSTS = {
|
||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"},
|
||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service", "user": "llmuser"},
|
||||
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
|
||||
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
|
||||
}
|
||||
@@ -122,10 +127,8 @@ def _fail(key, agent_name=None):
|
||||
|
||||
|
||||
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
|
||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||
|
||||
# Fallback: if no env token, read the shared vault token file
|
||||
if not INFISICAL_TOKEN:
|
||||
# Fallback: read the shared vault token file
|
||||
_token_path = os.path.expanduser("~/.infisical-token")
|
||||
if os.path.isfile(_token_path):
|
||||
try:
|
||||
@@ -133,6 +136,7 @@ if not INFISICAL_TOKEN:
|
||||
INFISICAL_TOKEN = _f.read().strip()
|
||||
except (OSError, UnicodeDecodeError):
|
||||
pass
|
||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||
|
||||
# ── Helpers ──────────────────────────────────────────────────────────
|
||||
|
||||
@@ -148,6 +152,18 @@ def ssh(host, cmd, user="root"):
|
||||
except:
|
||||
return None
|
||||
|
||||
def get_user_home(user):
|
||||
"""Resolve the home directory for a user.
|
||||
|
||||
For 'root', returns '/root'. For any other user, returns '/home/<user>'.
|
||||
This is used to construct paths that reference a user's home directory
|
||||
(e.g., ~/.local/bin/hermes, ~/.hermes/config.yaml) instead of hardcoding /root/.
|
||||
"""
|
||||
if user == "root":
|
||||
return "/root"
|
||||
else:
|
||||
return f"/home/{user}"
|
||||
|
||||
def http_get(url, headers=None, timeout=5):
|
||||
"""Return HTTP status code as string."""
|
||||
try:
|
||||
@@ -301,13 +317,14 @@ def check_gpu_ports():
|
||||
host = gpu["host"]
|
||||
port = gpu["port"]
|
||||
svc = gpu["service"]
|
||||
user = gpu.get("user", "root") # default root, overridden per-host where needed
|
||||
|
||||
# `systemctl is-active` exits non-zero when the unit is inactive or
|
||||
# missing, which the ssh() helper would swallow as an SSH failure and
|
||||
# report as UNREACHABLE. `|| true` keeps the real state word so we can
|
||||
# tell "unit inactive" from "host unreachable".
|
||||
svc_status = ssh(host, f"systemctl is-active {svc} || true")
|
||||
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
|
||||
svc_status = ssh(host, f"systemctl is-active {svc} || true", user=user)
|
||||
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1", user=user)
|
||||
|
||||
if not svc_status:
|
||||
print(f" ❌ {label}: UNREACHABLE")
|
||||
@@ -318,14 +335,14 @@ def check_gpu_ports():
|
||||
print(f" ❌ {label}: PORT {port} NOT LISTENING (svc={svc_status})")
|
||||
FAIL.append(f"gpu-no-port:{label}")
|
||||
elif svc_status != "active":
|
||||
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2")
|
||||
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2", user=user)
|
||||
if svc_pid and port_owner != svc_pid:
|
||||
print(f" ❌ {label}: GHOST PROCESS — port owned by pid {port_owner}, svc pid {svc_pid} (svc={svc_status})")
|
||||
FAIL.append(f"gpu-ghost:{label}:{port_owner}")
|
||||
else:
|
||||
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
|
||||
else:
|
||||
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health")
|
||||
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health", user=user)
|
||||
if health and '"status":"ok"' in health:
|
||||
print(f" ✅ {label}: healthy (pid={port_owner})")
|
||||
elif health and '"status":"no slot available"' in health:
|
||||
@@ -377,10 +394,10 @@ def check_agents():
|
||||
ct = agent["ct"]
|
||||
report_only = agent.get("report_only", False)
|
||||
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
|
||||
# Tanko runs hybrid (DSH + Hermes) since 2026-08-27 — it runs both DSH and Hermes gateway.
|
||||
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
|
||||
# Non-Hermes runtimes have no gateway to probe. dsh = Tanko since
|
||||
# 2026-08-27; pi = abiba since the harness purge (.24 is pi-only).
|
||||
# Non-Hermes runtimes have no gateway to probe. dsh/pi-only skip the check;
|
||||
# hybrid runs both DSH and Hermes and is checked normally.
|
||||
if agent.get("runtime") in ("dsh", "pi"):
|
||||
is_dsh = agent.get("runtime") == "dsh"
|
||||
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
|
||||
@@ -501,7 +518,7 @@ def check_ct_liveness():
|
||||
def check_config_integrity():
|
||||
"""Verify agent config.yaml parses as valid YAML."""
|
||||
for name, agent in AGENTS.items():
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
|
||||
# DSH/pi-only runtimes have no Hermes config.yaml; hybrid has both.
|
||||
if agent.get("runtime") == "dsh":
|
||||
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
|
||||
continue
|
||||
@@ -514,11 +531,11 @@ def check_config_integrity():
|
||||
print(f" ⬜ {name}: cannot SSH — skip config check")
|
||||
continue
|
||||
|
||||
home = get_user_home(user)
|
||||
|
||||
# Check YAML parses
|
||||
yaml_ok = ssh(host,
|
||||
"python3 -c "
|
||||
'"import yaml; yaml.safe_load(open(\'/root/.hermes/config.yaml\')); print(\'OK\')" '
|
||||
"2>&1 || echo 'FAIL'",
|
||||
f"python3 -c \"import yaml; yaml.safe_load(open('{home}/.hermes/config.yaml')); print('OK')\" 2>&1 || echo 'FAIL'",
|
||||
user=user)
|
||||
if not yaml_ok:
|
||||
print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)")
|
||||
@@ -557,7 +574,7 @@ def _infisical_invocation_paths(wrapper_body):
|
||||
def check_wrapper_integrity():
|
||||
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
||||
for name, agent in AGENTS.items():
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
|
||||
# DSH/pi-only runtimes have no hermes CLI wrapper; hybrid has both.
|
||||
if agent.get("runtime") == "dsh":
|
||||
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
|
||||
continue
|
||||
@@ -571,7 +588,8 @@ def check_wrapper_integrity():
|
||||
continue
|
||||
|
||||
# Check wrapper exists
|
||||
wrapper = ssh(host, "ls -la /root/.local/bin/hermes 2>/dev/null", user=user)
|
||||
home = get_user_home(user)
|
||||
wrapper = ssh(host, f"ls -la {home}/.local/bin/hermes 2>/dev/null", user=user)
|
||||
if not wrapper:
|
||||
# Check alternate wrapper locations
|
||||
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
|
||||
@@ -594,7 +612,7 @@ def check_wrapper_integrity():
|
||||
# a removed path (litellm-api-keys.prose.md documents
|
||||
# `rm -f /usr/local/bin/infisical`) must neither produce a dangling path
|
||||
# nor trigger the PATH check — it is not an invocation.
|
||||
wrapper_body = ssh(host, "cat /root/.local/bin/hermes 2>/dev/null", user=user) or ""
|
||||
wrapper_body = ssh(host, f"cat {home}/.local/bin/hermes 2>/dev/null", user=user) or ""
|
||||
wrapper_code = "\n".join(line.split("#", 1)[0] for line in wrapper_body.splitlines())
|
||||
invoked_paths = _infisical_invocation_paths(wrapper_body)
|
||||
if "infisical" in wrapper_code:
|
||||
@@ -628,24 +646,35 @@ def check_wrapper_integrity():
|
||||
else:
|
||||
print(f" ℹ️ {name}: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK")
|
||||
|
||||
# Check hermes-real exists
|
||||
# Check that the wrapper's target resolves. The fleet's wrappers do NOT
|
||||
# all use a hermes-real indirection — some exec the venv module directly.
|
||||
# Verify the wrapper actually points to something runnable.
|
||||
hermes_real = ssh(host,
|
||||
"ls -la /root/.local/bin/hermes-real 2>/dev/null || echo MISS",
|
||||
f"ls -la {home}/.local/bin/hermes-real 2>/dev/null || echo MISS",
|
||||
user=user)
|
||||
if not hermes_real or hermes_real.strip() == "MISS":
|
||||
# Check venv path
|
||||
hermes_real = ssh(host,
|
||||
"ls -la /usr/local/lib/hermes-agent/venv/bin/hermes 2>/dev/null || echo MISS",
|
||||
if hermes_real and hermes_real.strip() != "MISS":
|
||||
print(f" ✅ {name}: wrapper shape: hermes-real at {home}/.local/bin/hermes-real")
|
||||
else:
|
||||
# Try the venv under home
|
||||
venv_home = ssh(host,
|
||||
f"test -x {home}/.hermes/hermes-agent/venv/bin/python && echo OK || echo MISS",
|
||||
user=user)
|
||||
if not hermes_real or hermes_real.strip() == "MISS":
|
||||
print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)")
|
||||
_fail(f"wrapper-no-hermes-real:{name}", name)
|
||||
if venv_home and venv_home.strip().splitlines()[-1] == "OK":
|
||||
print(f" ✅ {name}: wrapper shape: direct venv exec ({home}/.hermes/hermes-agent/venv/bin/python)")
|
||||
else:
|
||||
print(f" ✅ {name}: hermes-real at alt path")
|
||||
# Try the system-wide venv
|
||||
venv_sys = ssh(host,
|
||||
"test -x /usr/local/lib/hermes-agent/venv/bin/python && echo OK || echo MISS",
|
||||
user=user)
|
||||
if venv_sys and venv_sys.strip().splitlines()[-1] == "OK":
|
||||
print(f" ✅ {name}: wrapper shape: system venv (/usr/local/lib/hermes-agent/venv/bin/python)")
|
||||
else:
|
||||
print(f" ❌ {name}: wrapper target NOT RESOLVABLE (no hermes-real, no venv)")
|
||||
_fail(f"wrapper-no-hermes-real:{name}", name)
|
||||
|
||||
# Check the .env file has the key
|
||||
env_has_key = ssh(host,
|
||||
"grep -c 'LITELLM_API_KEY' /root/.hermes/.env 2>/dev/null || echo 0",
|
||||
f"grep -c 'LITELLM_API_KEY' {home}/.hermes/.env 2>/dev/null || echo 0",
|
||||
user=user)
|
||||
if env_has_key and env_has_key.strip() not in ("", "0"):
|
||||
print(f" ✅ {name}: wrapper + .env key present")
|
||||
|
||||
Executable
+225
@@ -0,0 +1,225 @@
|
||||
#!/bin/bash
|
||||
# contract-run.sh — Deterministic contract execution from machine scheduler
|
||||
#
|
||||
# Takes a contract name, resolves its script, runs it with timeout,
|
||||
# logs output to $CONTRACT_RUN_LOG_DIR (default: /var/log/contract-runs/),
|
||||
# and alerts on failure.
|
||||
#
|
||||
# Environment:
|
||||
# CONTRACT_RUN_LOG_DIR Override the log directory (default: /var/log/contract-runs)
|
||||
#
|
||||
# Usage: bash scripts/contract-run.sh <contract-name>
|
||||
#
|
||||
# Contract names map to scripts as follows:
|
||||
# infrastructure-monitoring -> scripts/infra-monitoring.sh
|
||||
# proxmox-monitor -> scripts/proxmox-monitor.sh
|
||||
# zulip-health -> scripts/zulip-monitor.sh
|
||||
# agent-health-check -> scripts/agent-health-check.py
|
||||
# litellm-health -> scripts/litellm-health-check.py
|
||||
# disk-gc-threat-response -> scripts/disk-gc-scan.py
|
||||
# pm2-self-heal -> scripts/pm2-self-heal.sh
|
||||
# search-stack-visibility -> scripts/search-stack-check.py
|
||||
# health-log-freshness -> scripts/health-log-freshness.py
|
||||
#
|
||||
# Execution copy: every contract pins the clone this script lives in (see
|
||||
# docs/contract-execution-pinning.md). Before a contract runs, this wrapper
|
||||
# proves the script it is about to execute byte-matches origin/master:
|
||||
# CONTRACT_REVISION_PREFLIGHT=warn (default) log a refusal, still report
|
||||
# CONTRACT_REVISION_PREFLIGHT=enforce refuse to report on a mismatch
|
||||
# CONTRACT_REVISION_PREFLIGHT=off skip the check entirely
|
||||
# A refusal names its class: cannot-verify:fetch-failed|ref-unresolvable,
|
||||
# or mismatch:content|path-absent|detached-head|clone-ahead.
|
||||
#
|
||||
# Exit codes:
|
||||
# 0 = contract passed
|
||||
# 1 = contract failed (alert sent)
|
||||
# 2 = probe failed (script missing, timeout, etc.)
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
CONTRACT_NAME="$1"
|
||||
SCRIPTS_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
LOG_DIR="${CONTRACT_RUN_LOG_DIR:-/var/log/contract-runs}"
|
||||
TIMESTAMP=$(date -u '+%Y%m%d-%H%M%S')
|
||||
LOG_FILE="${LOG_DIR}/${CONTRACT_NAME}-${TIMESTAMP}.log"
|
||||
|
||||
# Ensure log directory exists
|
||||
mkdir -p "$LOG_DIR"
|
||||
|
||||
# Map contract name to script path
|
||||
case "$CONTRACT_NAME" in
|
||||
infrastructure-monitoring)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/infra-monitoring.sh"
|
||||
INTERPRETER="bash"
|
||||
;;
|
||||
proxmox-monitor)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/proxmox-monitor.sh"
|
||||
INTERPRETER="bash"
|
||||
;;
|
||||
zulip-health)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/zulip-monitor.sh"
|
||||
INTERPRETER="bash"
|
||||
;;
|
||||
agent-health-check)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/agent-health-check.py"
|
||||
INTERPRETER="python3"
|
||||
;;
|
||||
litellm-health)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/litellm-health-check.py"
|
||||
INTERPRETER="python3"
|
||||
;;
|
||||
pm2-self-heal)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/pm2-self-heal.sh"
|
||||
INTERPRETER="bash"
|
||||
;;
|
||||
disk-gc-threat-response)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/disk-gc-scan.py"
|
||||
INTERPRETER="python3"
|
||||
;;
|
||||
search-stack-visibility)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/search-stack-check.py"
|
||||
INTERPRETER="python3"
|
||||
;;
|
||||
health-log-freshness)
|
||||
SCRIPT_PATH="${SCRIPTS_DIR}/health-log-freshness.py"
|
||||
INTERPRETER="python3"
|
||||
;;
|
||||
*)
|
||||
echo "Unknown contract: $CONTRACT_NAME" | tee -a "$LOG_FILE"
|
||||
# Send alert for unknown contract
|
||||
ALERT_MSG="🔴 Contract $CONTRACT_NAME: unknown contract name. Log: $LOG_FILE"
|
||||
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
||||
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
||||
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||
-d "type=private" \
|
||||
-d "to=9" \
|
||||
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
||||
fi
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
|
||||
# Check if script exists
|
||||
if [ ! -f "$SCRIPT_PATH" ]; then
|
||||
echo "Script not found: $SCRIPT_PATH" | tee -a "$LOG_FILE"
|
||||
# Send alert for missing script
|
||||
ALERT_MSG="🔴 Contract $CONTRACT_NAME: script not found at $SCRIPT_PATH. Log: $LOG_FILE"
|
||||
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
||||
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
||||
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||
-d "type=private" \
|
||||
-d "to=9" \
|
||||
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
||||
fi
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# Run the script with timeout and capture output
|
||||
echo "=== Contract: $CONTRACT_NAME ===" | tee "$LOG_FILE"
|
||||
echo "Started: $(date -u '+%Y-%m-%d %H:%M:%S UTC')" | tee -a "$LOG_FILE"
|
||||
echo "Script: $SCRIPT_PATH" | tee -a "$LOG_FILE"
|
||||
echo "" | tee -a "$LOG_FILE"
|
||||
|
||||
# ── Revision preflight ───────────────────────────────────────────────────────
|
||||
# A verdict is only meaningful if it came from the merged copy. This reports
|
||||
# whether the executing copy matches, and distinguishes "could not check" from
|
||||
# "this copy is wrong" so an operator can tell them apart.
|
||||
# See docs/contract-execution-pinning.md.
|
||||
#
|
||||
# Default is WARN, not enforce: the guard gates every scheduled contract, and a
|
||||
# legitimate state (branch mid-review, detached HEAD, briefly offline) would
|
||||
# otherwise turn the whole fleet's monitoring into withheld verdicts. The
|
||||
# criteria for flipping the default to enforce are written down in the doc.
|
||||
REVISION_PREFLIGHT_MODE="${CONTRACT_REVISION_PREFLIGHT:-warn}"
|
||||
REPO_ROOT="$(cd "${SCRIPTS_DIR}/.." && pwd)"
|
||||
if [ "$REVISION_PREFLIGHT_MODE" != "off" ] && [ -x "${SCRIPTS_DIR}/revision-preflight.sh" ]; then
|
||||
PREFLIGHT_OUT="$(mktemp)"
|
||||
if "${SCRIPTS_DIR}/revision-preflight.sh" "$SCRIPT_PATH" "$REPO_ROOT" >"$PREFLIGHT_OUT" 2>&1; then
|
||||
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
|
||||
else
|
||||
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
|
||||
PREFLIGHT_REASON="$(grep -m1 '^REASON=' "$PREFLIGHT_OUT" | cut -d= -f2-)"
|
||||
[ -n "$PREFLIGHT_REASON" ] || PREFLIGHT_REASON="unclassified"
|
||||
if [ "$REVISION_PREFLIGHT_MODE" = "enforce" ]; then
|
||||
echo "🚫 VERDICT WITHHELD: $PREFLIGHT_REASON" | tee -a "$LOG_FILE"
|
||||
ALERT_MSG="🔴 Contract $CONTRACT_NAME: revision preflight REFUSED ($PREFLIGHT_REASON) — verdict withheld. Log: $LOG_FILE"
|
||||
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
||||
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
||||
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||
-d "type=private" \
|
||||
-d "to=9" \
|
||||
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
||||
fi
|
||||
rm -f "$PREFLIGHT_OUT"
|
||||
exit 2
|
||||
fi
|
||||
echo "⚠️ revision preflight: $PREFLIGHT_REASON — continuing because CONTRACT_REVISION_PREFLIGHT=$REVISION_PREFLIGHT_MODE" | tee -a "$LOG_FILE"
|
||||
fi
|
||||
rm -f "$PREFLIGHT_OUT"
|
||||
fi
|
||||
|
||||
# Use timeout to prevent hangs (10 minutes default)
|
||||
TIMEOUT=600
|
||||
timeout "$TIMEOUT" $INTERPRETER "$SCRIPT_PATH" 2>&1 | tee -a "$LOG_FILE"
|
||||
EXIT_CODE=${PIPESTATUS[0]}
|
||||
|
||||
# If timeout killed the process, EXIT_CODE will be 124
|
||||
if [ $EXIT_CODE -eq 124 ]; then
|
||||
echo "⏰ TIMEOUT: script exceeded ${TIMEOUT}s limit" | tee -a "$LOG_FILE"
|
||||
fi
|
||||
|
||||
echo "" | tee -a "$LOG_FILE"
|
||||
if [ $EXIT_CODE -eq 0 ]; then
|
||||
echo "✅ VERDICT: PASS" | tee -a "$LOG_FILE"
|
||||
exit 0
|
||||
else
|
||||
echo "🔴 VERDICT: FAIL (exit code $EXIT_CODE)" | tee -a "$LOG_FILE"
|
||||
|
||||
# Send alert (Zulip DM to user 9 + stream agent-hub topic alerts-infra)
|
||||
# Using the same alert path as other monitors
|
||||
ALERT_MSG="🔴 Contract $CONTRACT_NAME failed (exit $EXIT_CODE). Log: $LOG_FILE"
|
||||
ALERT_SENT=false
|
||||
|
||||
# Take credentials from environment (ZULIP_API_KEY required)
|
||||
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
||||
|
||||
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
||||
# DM to user 9
|
||||
DM_EXIT=0
|
||||
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||
-d "type=private" \
|
||||
-d "to=9" \
|
||||
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || DM_EXIT=$?
|
||||
|
||||
# Stream agent-hub topic alerts-infra
|
||||
STREAM_EXIT=0
|
||||
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||
-d "type=stream" \
|
||||
-d "to=agent-hub" \
|
||||
-d "topic=alerts-infra" \
|
||||
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || STREAM_EXIT=$?
|
||||
|
||||
if [ $DM_EXIT -eq 0 ] || [ $STREAM_EXIT -eq 0 ]; then
|
||||
ALERT_SENT=true
|
||||
else
|
||||
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT FAILURE: DM exit=$DM_EXIT, stream exit=$STREAM_EXIT" >> "$LOG_FILE"
|
||||
fi
|
||||
else
|
||||
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT SKIPPED: no ZULIP_API_KEY or curl" >> "$LOG_FILE"
|
||||
fi
|
||||
|
||||
exit 1
|
||||
fi
|
||||
+230
-40
@@ -16,14 +16,42 @@ from email.mime.text import MIMEText
|
||||
from email.mime.multipart import MIMEMultipart
|
||||
|
||||
PVE = "https://192.168.68.12:8006"
|
||||
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
|
||||
# The HTTP header prefix is a protocol constant, not a credential. It is kept as
|
||||
# a constant ending at '=' so that no assembled header-plus-token literal ever
|
||||
# appears in the tree; the secret scanner rightly flags that shape.
|
||||
PVE_AUTH_HEADER = "PVEAPIToken="
|
||||
|
||||
|
||||
def pve_auth():
|
||||
"""PVE API auth header, resolved at call time from the injected environment.
|
||||
|
||||
The token is injected by ``infisical run --env=prod`` as ``PVE_TOKEN``
|
||||
(format ``user@realm!tokenid=secret``). It must never be hardcoded: a
|
||||
placeholder literal authenticates as nobody, which is how this probe
|
||||
reported zero nodes while still exiting 0. Raise loudly instead.
|
||||
"""
|
||||
token = os.environ.get("PVE_TOKEN")
|
||||
if not token:
|
||||
raise RuntimeError("PVE_TOKEN is not set (run under `infisical run --env=prod`)")
|
||||
return f"Authorization: {PVE_AUTH_HEADER}{token}"
|
||||
|
||||
# ── Shared credentials —─
|
||||
|
||||
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_KEY = "cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_KEY}"
|
||||
# Note: /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential).
|
||||
# No Zulip API key is required for this call. If a future leg genuinely needs abiba-bot's key,
|
||||
# it must prove it with a 200 from /api/v1/users/me as abiba-bot and label itself degraded when it cannot.
|
||||
# Never fall back to the vault's shared ZULIP_API_KEY.
|
||||
ZULIP_AUTH = None
|
||||
DEGRADED_LEGS = []
|
||||
|
||||
# Probe failures are different from degraded legs. A missing credential is an
|
||||
# expected, survivable state (stays exit 0). A probe that cannot reach the API
|
||||
# means the report has NO data for that section, which is a monitoring loss and
|
||||
# must exit non-zero so it cannot pass unnoticed.
|
||||
PROBE_FAILURES = []
|
||||
|
||||
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||
LITELLM_BACKEND = "192.168.68.116"
|
||||
@@ -44,9 +72,13 @@ TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
|
||||
# ── Helpers ──
|
||||
|
||||
def pve_get(path):
|
||||
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list)."""
|
||||
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list).
|
||||
|
||||
A missing PVE_TOKEN is caught here and reported as ``None`` so the caller
|
||||
records a probe failure; it must not escape as an unhandled exception.
|
||||
"""
|
||||
try:
|
||||
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{pve_auth()}"'
|
||||
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
|
||||
if r.returncode != 0:
|
||||
return None
|
||||
@@ -114,6 +146,7 @@ def collect():
|
||||
report["node_count"] = 0
|
||||
report["nodes_online"] = 0
|
||||
report["pve_probe_status"] = "unreachable"
|
||||
PROBE_FAILURES.append("proxmox: node list unreachable (PVE_TOKEN missing or API down)")
|
||||
else:
|
||||
report["nodes"] = {n["node"]: {
|
||||
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||
@@ -133,6 +166,7 @@ def collect():
|
||||
if resources is None:
|
||||
vms = []
|
||||
report["resources_probe_status"] = "unreachable"
|
||||
PROBE_FAILURES.append("proxmox: cluster resources unreachable")
|
||||
else:
|
||||
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||
report["resources_probe_status"] = "ok"
|
||||
@@ -156,7 +190,7 @@ def collect():
|
||||
# ── Storage ──
|
||||
storages = pve_get("/api2/json/nodes/storepve/storage")
|
||||
report["storage"] = []
|
||||
for s in storages:
|
||||
for s in (storages or []):
|
||||
total = s.get("total",0) or 1
|
||||
used = s.get("used",0)
|
||||
pct = used/total*100
|
||||
@@ -209,7 +243,7 @@ def collect():
|
||||
("Pulse", "https://pulse.sysloggh.net"),
|
||||
("Proxmox", "https://192.168.68.12:8006"),
|
||||
("SearXNG", "http://192.168.68.7:8888"),
|
||||
("Firecrawl", "http://192.168.68.7:3002/health"),
|
||||
("Firecrawl", "http://192.168.68.7:3002/"), # Firecrawl serves no /health - the root is its liveness endpoint
|
||||
]
|
||||
report["endpoints"] = []
|
||||
for name, url in endpoints:
|
||||
@@ -355,6 +389,28 @@ def collect():
|
||||
|
||||
# ── HTML Dashboard ──
|
||||
|
||||
def classify_endpoint(code):
|
||||
"""Classify an endpoint probe per the fleet's probe policy.
|
||||
|
||||
Codified 2026-09-14 in the monitoring contracts: ANY HTTP status proves the
|
||||
service answered, so the service is ALIVE - 200/301/302/401/403/404 alike.
|
||||
Only a failed CONNECTION (000 / timeout / refused) is a failed probe. A 404
|
||||
from a wrong path is not a service fault and must not render as one.
|
||||
|
||||
This replaces a string comparison that was wrong in both directions
|
||||
(`ep["code"] >= "400"`): it rendered 301 as red, 404 as yellow, and a real
|
||||
500 as yellow. 5xx is kept as its own "server error" signal rather than
|
||||
being merged with 4xx.
|
||||
"""
|
||||
if not code or code == "000":
|
||||
return "red", "no connection"
|
||||
if code.startswith("5"):
|
||||
return "yellow", "server error"
|
||||
if code.startswith(("2", "3", "4")):
|
||||
return "green", "alive"
|
||||
return "yellow", f"unexpected {code}"
|
||||
|
||||
|
||||
def build_html(r):
|
||||
issues = []
|
||||
|
||||
@@ -590,7 +646,7 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']
|
||||
# ── Network Endpoints ──
|
||||
html += '<div class="card"><h2>🌐 Network Endpoints</h2><table><tr><th>Service</th><th>Status</th></tr>'
|
||||
for ep in r["endpoints"]:
|
||||
color = "green" if ep["code"] in ("200","302","401") else ("yellow" if ep["code"] >= "400" else "red")
|
||||
color = classify_endpoint(ep["code"])[0]
|
||||
html += f'<tr><td>{ep["name"]}</td><td class="{color}">HTTP {ep["code"]}</td></tr>'
|
||||
html += '</table></div>'
|
||||
|
||||
@@ -654,60 +710,188 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']
|
||||
return html
|
||||
|
||||
|
||||
# ── Send Email ──
|
||||
# ── Delivery: Zulip DM carrying the report as an HTML ATTACHMENT ──
|
||||
#
|
||||
# Captain's decision, clarified 2026-09-26: the report is sent as an HTML FILE,
|
||||
# i.e. an attachment - NOT HTML rendered in the message body, and NOT a Markdown
|
||||
# translation of it. So the styled dashboard is built exactly as before, uploaded
|
||||
# through Zulip's file-upload API, and the message body stays short: subject,
|
||||
# top-line status, and a pointer to the attachment.
|
||||
#
|
||||
# This removes the Google dependency entirely (no SMTP, no EMAIL_PASSWORD).
|
||||
# The 10,000-character message cap does not apply: it bounds message TEXT only,
|
||||
# and the report travels as a file.
|
||||
|
||||
def send_email(html_content, subject_prefix=""):
|
||||
FROM = "abiba@sysloggh.com"
|
||||
TO = "jerome@sysloggh.com"
|
||||
SUBJECT = f"{subject_prefix}{'🏗️ Infrastructure Report — ' + DATE_STR}"
|
||||
|
||||
msg = MIMEMultipart("alternative")
|
||||
msg["From"] = FROM
|
||||
msg["To"] = TO
|
||||
msg["Subject"] = SUBJECT
|
||||
msg.attach(MIMEText("Infrastructure report in HTML format — enable images to view.", "plain"))
|
||||
msg.attach(MIMEText(html_content, "html"))
|
||||
|
||||
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||
ZULIP_BOT_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||
CAPTAIN_USER_ID = 9
|
||||
ZULIP_KEY_FILE = "/root/.pi/agent/extensions/zulip/.env"
|
||||
REPORT_ARTIFACT_DIR = "/var/log/daily-infra-report"
|
||||
|
||||
|
||||
def zulip_key():
|
||||
"""abiba-bot's Zulip key, from the env or the on-host 600 file."""
|
||||
key = os.environ.get("ABIBA_ZULIP_API_KEY")
|
||||
if key:
|
||||
return key.strip()
|
||||
try:
|
||||
EMAIL_PASSWORD = "rgbuomwcydxwbszd"
|
||||
GMAIL_EMAIL = "jtabiri@gmail.com"
|
||||
|
||||
server = smtplib.SMTP("smtp.gmail.com", 587)
|
||||
server.starttls()
|
||||
server.login(GMAIL_EMAIL, EMAIL_PASSWORD)
|
||||
server.sendmail(FROM, [TO], msg.as_string())
|
||||
server.quit()
|
||||
return True, "✅ Email sent to jerome@sysloggh.com"
|
||||
except Exception as e:
|
||||
return False, f"❌ Email failed: {e}"
|
||||
with open(ZULIP_KEY_FILE) as fh:
|
||||
for line in fh:
|
||||
if line.startswith("ABIBA_ZULIP_API_KEY="):
|
||||
return line.split("=", 1)[1].strip()
|
||||
except OSError:
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def build_summary(r, filename, test=False):
|
||||
"""Short Markdown body: subject, top-line status, pointer to the attachment.
|
||||
|
||||
Deliberately NOT a reproduction of the report - the attachment is the report.
|
||||
"""
|
||||
nodes = f"{r.get('nodes_online', 0)}/{r.get('node_count', 0)} nodes online"
|
||||
guests = f"{r.get('running_vms', 0)}/{r.get('total_vms', 0)} guests running"
|
||||
lines = [
|
||||
("\U0001F9EA **TEST — **" if test else "") + "\U0001F3D7\uFE0F **Infrastructure Report — " + DATE_STR + "**",
|
||||
f"**{nodes}** \u00b7 **{guests}** \u00b7 generated {TIME_STR}",
|
||||
]
|
||||
problems = []
|
||||
if r.get("pve_probe_status") != "ok":
|
||||
problems.append(f"\u274c Proxmox probe: {r.get('pve_probe_status')}")
|
||||
if r.get("resources_probe_status") != "ok":
|
||||
problems.append(f"\u274c Resources probe: {r.get('resources_probe_status')}")
|
||||
lit = r.get("litellm", {}) or {}
|
||||
checks = lit.get("checks", []) or []
|
||||
if checks:
|
||||
passed = sum(1 for c in checks if c.get("status") == "pass")
|
||||
if passed != len(checks):
|
||||
problems.append(f"\u274c LiteLLM: {passed}/{len(checks)} checks pass")
|
||||
if not (r.get("zulip_ext", {}) or {}).get("connected"):
|
||||
problems.append("\u274c Zulip extension: not connected")
|
||||
for leg in DEGRADED_LEGS:
|
||||
problems.append(f"\u26a0\uFE0F degraded: {leg}")
|
||||
|
||||
lines.append("\n".join(problems) if problems else "\u2705 All monitored services healthy")
|
||||
lines.append(f"\U0001F4CE **Full report attached:** `{filename}`")
|
||||
return "\n\n".join(lines)
|
||||
|
||||
|
||||
def _curl(args, timeout=60):
|
||||
r = subprocess.run(["curl", "-s", "-m", str(timeout)] + args,
|
||||
capture_output=True, text=True)
|
||||
try:
|
||||
return json.loads(r.stdout or "{}"), r.stdout
|
||||
except json.JSONDecodeError:
|
||||
return {}, r.stdout
|
||||
|
||||
|
||||
def _curl_json(args, timeout=90):
|
||||
r = subprocess.run(["curl", "-s", "-m", str(timeout)] + args,
|
||||
capture_output=True, text=True)
|
||||
try:
|
||||
return json.loads(r.stdout or "{}"), r.stdout
|
||||
except json.JSONDecodeError:
|
||||
return {}, r.stdout
|
||||
|
||||
|
||||
def send_zulip(html_content, report, test=False):
|
||||
"""Upload the styled HTML and post a short pointer to the captain's DM.
|
||||
|
||||
Returns (ok, message). On ANY failure the report body is also printed to
|
||||
stdout and persisted to disk, so a delivery failure can never swallow the
|
||||
content - the defect this folds in.
|
||||
"""
|
||||
os.makedirs(REPORT_ARTIFACT_DIR, exist_ok=True)
|
||||
stamp = NOW.strftime("%Y%m%d-%H%M%S")
|
||||
filename = f"infra-report-{stamp}.html"
|
||||
html_path = os.path.join(REPORT_ARTIFACT_DIR, filename)
|
||||
try:
|
||||
with open(html_path, "w") as fh:
|
||||
fh.write(html_content)
|
||||
except OSError as e:
|
||||
print(f" \u26a0\uFE0F could not persist report artifact: {e}", file=sys.stderr)
|
||||
|
||||
key = zulip_key()
|
||||
if not key:
|
||||
print(html_content) # never swallow the content
|
||||
return False, ("\u274c Delivery FAILED: no Zulip credential "
|
||||
"(ABIBA_ZULIP_API_KEY unset and "
|
||||
f"{ZULIP_KEY_FILE} unreadable). Report persisted to {html_path}")
|
||||
|
||||
auth = ["-u", f"{ZULIP_BOT_EMAIL}:{key}"]
|
||||
|
||||
# 1. Upload the report as a file.
|
||||
up, up_raw = _curl_json(auth + [
|
||||
"-X", "POST", f"{ZULIP_SITE}/api/v1/user_uploads",
|
||||
"-F", f"file=@{html_path};type=text/html",
|
||||
])
|
||||
if up.get("result") != "success" or not up.get("uri"):
|
||||
print(html_content)
|
||||
return False, (f"\u274c Delivery FAILED at upload: {up.get('msg') or up_raw[:160]} "
|
||||
f"(report persisted to {html_path})")
|
||||
|
||||
uri = up["uri"]
|
||||
size = os.path.getsize(html_path)
|
||||
|
||||
# 2. Post a short message pointing at it.
|
||||
body = build_summary(report, filename, test=test)
|
||||
link = f"[{filename}]({uri})"
|
||||
body = body.replace(f"`{filename}`", link)
|
||||
payload, raw = _curl_json(auth + [
|
||||
"-X", "POST", f"{ZULIP_SITE}/api/v1/messages",
|
||||
"-d", "type=private",
|
||||
"-d", f"to=[{CAPTAIN_USER_ID}]",
|
||||
"--data-urlencode", f"content={body}",
|
||||
])
|
||||
if payload.get("result") == "success":
|
||||
return True, (f"\u2705 Delivered to Zulip DM (user {CAPTAIN_USER_ID}), "
|
||||
f"message id {payload.get('id')}, attachment {size} bytes at {uri}")
|
||||
|
||||
print(html_content)
|
||||
return False, (f"\u274c Delivery FAILED at message post: {payload.get('msg') or raw[:160]} "
|
||||
f"(uploaded {uri}; report persisted to {html_path})")
|
||||
|
||||
|
||||
# ── Main ──
|
||||
|
||||
if __name__ == "__main__":
|
||||
is_test = "--test-email" in sys.argv
|
||||
is_test = ("--test-email" in sys.argv) or ("--test-zulip" in sys.argv)
|
||||
|
||||
print(f"{'🧪 TEST MODE' if is_test else '📊'} Collecting infrastructure data...")
|
||||
report = collect()
|
||||
|
||||
if "--json" in sys.argv:
|
||||
print(json.dumps(report, indent=2, default=str))
|
||||
if PROBE_FAILURES:
|
||||
for leg in PROBE_FAILURES:
|
||||
print(f"PROBE FAILURE: {leg}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
sys.exit(0)
|
||||
|
||||
print(" Building dashboard...")
|
||||
html = build_html(report)
|
||||
|
||||
print(f" report ready: {len(html)} chars of HTML (delivered as a file attachment)")
|
||||
|
||||
if is_test:
|
||||
prefix = "🧪 TEST — "
|
||||
print(" Sending test email...")
|
||||
print(" Sending TEST message to the captain's Zulip DM...")
|
||||
else:
|
||||
prefix = ""
|
||||
print(" Sending email...")
|
||||
|
||||
ok, msg = send_email(html, subject_prefix=prefix)
|
||||
print(" Sending to the captain's Zulip DM...")
|
||||
|
||||
ok, msg = send_zulip(html, report, test=is_test)
|
||||
print(f" {msg}")
|
||||
|
||||
# Show summary
|
||||
if DEGRADED_LEGS:
|
||||
print(f"\n⚠️ Degraded legs ({len(DEGRADED_LEGS)}):")
|
||||
for leg in DEGRADED_LEGS:
|
||||
print(f" - {leg}")
|
||||
else:
|
||||
print("\n✅ All legs fully credentialed")
|
||||
|
||||
# A failed send must exit non-zero; a degraded leg (no credential) must stay exit 0
|
||||
if not ok:
|
||||
sys.exit(1)
|
||||
|
||||
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
|
||||
print(f"\n📋 Summary:")
|
||||
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
|
||||
@@ -718,3 +902,9 @@ if __name__ == "__main__":
|
||||
for k,v in report.get('agents',{}).items():
|
||||
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
||||
print(f" Agents: {', '.join(agent_parts)}")
|
||||
|
||||
if PROBE_FAILURES:
|
||||
print(f"\n❌ Probe failures ({len(PROBE_FAILURES)}):")
|
||||
for leg in PROBE_FAILURES:
|
||||
print(f" - {leg}")
|
||||
sys.exit(1)
|
||||
|
||||
+238
-8
@@ -101,8 +101,6 @@ GUESTS: list[Guest] = [
|
||||
# amdpve (192.168.68.15)
|
||||
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
|
||||
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
|
||||
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
|
||||
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
|
||||
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
|
||||
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
|
||||
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
|
||||
@@ -110,6 +108,8 @@ GUESTS: list[Guest] = [
|
||||
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
|
||||
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
|
||||
# minipve (192.168.68.12)
|
||||
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="minipve",
|
||||
access_method="pct-run", probe_target="tanko (CT 112, minipve)"),
|
||||
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
|
||||
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
|
||||
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
|
||||
@@ -153,6 +153,25 @@ GPU_HOSTS = [
|
||||
CONNECT_TIMEOUT = 5
|
||||
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
|
||||
|
||||
# Host filesystem thresholds (from contract)
|
||||
HOST_THRESHOLDS = {
|
||||
"WARN": 85,
|
||||
"AMBER": 90,
|
||||
"RED": 95,
|
||||
}
|
||||
|
||||
# State file path (absolute, so execution context doesn't matter)
|
||||
STATE_FILE = pathlib.Path(__file__).resolve().parent.parent / "state" / "host-disk-bands.json"
|
||||
|
||||
# PVE nodes to probe for host filesystems
|
||||
HOST_NODES = [
|
||||
{"hostname": "acerpve", "ip": "192.168.68.9"},
|
||||
{"hostname": "amdpve", "ip": "192.168.68.15"},
|
||||
{"hostname": "storepve", "ip": "192.168.68.6"},
|
||||
{"hostname": "minipve", "ip": "192.168.68.12"},
|
||||
{"hostname": "ocupve", "ip": "192.168.68.5"},
|
||||
]
|
||||
|
||||
|
||||
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
|
||||
"""Run a command and return (exit_code, stdout, stderr)."""
|
||||
@@ -298,6 +317,151 @@ def scan_fleet() -> list[dict]:
|
||||
return results
|
||||
|
||||
|
||||
def classify_band(usage_pct: float) -> str:
|
||||
"""Classify a percentage into a band."""
|
||||
if usage_pct >= HOST_THRESHOLDS["RED"]:
|
||||
return "HOST-RED"
|
||||
elif usage_pct >= HOST_THRESHOLDS["AMBER"]:
|
||||
return "HOST-AMBER"
|
||||
elif usage_pct >= HOST_THRESHOLDS["WARN"]:
|
||||
return "HOST-WARN"
|
||||
else:
|
||||
return "GREEN"
|
||||
|
||||
|
||||
def probe_host_filesystems() -> tuple[list[dict], dict[str, str]]:
|
||||
"""Probe host filesystems on all PVE nodes.
|
||||
|
||||
Returns:
|
||||
- List of host filesystem results
|
||||
- Dict of volume_key -> current_band (for state file)
|
||||
"""
|
||||
results = []
|
||||
current_bands = {}
|
||||
|
||||
for node in HOST_NODES:
|
||||
ip = node["ip"]
|
||||
hostname = node["hostname"]
|
||||
|
||||
# Probe df for host filesystems
|
||||
probe_cmd = f'ssh {SSH_OPTS} root@{ip} "df -hP / /media/* tank 2>/dev/null | tail -n +2"'
|
||||
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
|
||||
|
||||
if exit_code != 0:
|
||||
results.append({
|
||||
"target": f"{hostname} ({ip})",
|
||||
"hostname": hostname,
|
||||
"ip": ip,
|
||||
"reachable": False,
|
||||
"volumes": [],
|
||||
"probe_cmd": probe_cmd,
|
||||
"failure_kind": "ssh-error",
|
||||
})
|
||||
continue
|
||||
|
||||
# Parse df output and classify each volume
|
||||
volumes = []
|
||||
for line in stdout.splitlines():
|
||||
parts = line.split()
|
||||
if len(parts) < 6:
|
||||
continue
|
||||
|
||||
dev, size, used, avail, pct_str, mount = parts[:6]
|
||||
pct = float(pct_str.rstrip("%"))
|
||||
band = classify_band(pct)
|
||||
|
||||
# Volume type classification
|
||||
if mount.startswith("/media/"):
|
||||
vol_type = "media"
|
||||
elif mount == "/" or "pve" in dev:
|
||||
vol_type = "host-root"
|
||||
elif mount == "tank" or "tank" in mount:
|
||||
vol_type = "pbs-datastore"
|
||||
else:
|
||||
vol_type = "other"
|
||||
|
||||
# Volume key for state file (host/volume)
|
||||
volume_key = f"{hostname}/{mount}"
|
||||
current_bands[volume_key] = band
|
||||
|
||||
volumes.append({
|
||||
"mount": mount,
|
||||
"device": dev,
|
||||
"size": size,
|
||||
"used": used,
|
||||
"avail": avail,
|
||||
"pct": pct,
|
||||
"band": band,
|
||||
"type": vol_type,
|
||||
})
|
||||
|
||||
results.append({
|
||||
"target": f"{hostname} ({ip})",
|
||||
"hostname": hostname,
|
||||
"ip": ip,
|
||||
"reachable": True,
|
||||
"volumes": volumes,
|
||||
"probe_cmd": probe_cmd,
|
||||
"failure_kind": None,
|
||||
})
|
||||
|
||||
return results, current_bands
|
||||
|
||||
|
||||
def read_state_file() -> Optional[dict[str, str]]:
|
||||
"""Read the state file if it exists."""
|
||||
if not STATE_FILE.exists():
|
||||
return None
|
||||
try:
|
||||
with open(STATE_FILE) as f:
|
||||
return json.load(f)
|
||||
except (json.JSONDecodeError, IOError) as e:
|
||||
print(f"⚠️ State file exists but unreadable: {e}", file=sys.stderr)
|
||||
return {}
|
||||
|
||||
|
||||
def write_state_file(bands: dict[str, str]) -> None:
|
||||
"""Write the state file."""
|
||||
STATE_FILE.parent.mkdir(parents=True, exist_ok=True)
|
||||
try:
|
||||
with open(STATE_FILE, "w") as f:
|
||||
json.dump(bands, f, indent=2)
|
||||
except IOError as e:
|
||||
print(f"⚠️ State file write failed: {e}", file=sys.stderr)
|
||||
|
||||
|
||||
def detect_transitions(current_bands: dict[str, str], prior_bands: Optional[dict[str, str]]) -> list[dict]:
|
||||
"""Detect band transitions (current vs. prior)."""
|
||||
if prior_bands is None:
|
||||
# First run — no transitions, just establish baseline
|
||||
return []
|
||||
|
||||
transitions = []
|
||||
# Check for volumes that moved to a higher band (escalation)
|
||||
for volume, current_band in current_bands.items():
|
||||
prior_band = prior_bands.get(volume, "GREEN")
|
||||
|
||||
# Band ordering: GREEN < HOST-WARN < HOST-AMBER < HOST-RED
|
||||
band_order = {"GREEN": 0, "HOST-WARN": 1, "HOST-AMBER": 2, "HOST-RED": 3}
|
||||
|
||||
if band_order[current_band] > band_order[prior_band]:
|
||||
transitions.append({
|
||||
"type": "escalation",
|
||||
"volume": volume,
|
||||
"from": prior_band,
|
||||
"to": current_band,
|
||||
})
|
||||
elif band_order[current_band] < band_order[prior_band]:
|
||||
transitions.append({
|
||||
"type": "recovery",
|
||||
"volume": volume,
|
||||
"from": prior_band,
|
||||
"to": current_band,
|
||||
})
|
||||
|
||||
return transitions
|
||||
|
||||
|
||||
def render_results(results: list[dict]) -> str:
|
||||
"""Render scan results in human-readable format."""
|
||||
lines = []
|
||||
@@ -316,22 +480,88 @@ def render_results(results: list[dict]) -> str:
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def render_host_results(results: list[dict], transitions: list[dict], prior_bands: Optional[dict[str, str]]) -> str:
|
||||
"""Render host filesystem results in human-readable format."""
|
||||
lines = []
|
||||
lines.append("")
|
||||
lines.append("=== Host Filesystem Bands ===")
|
||||
lines.append("")
|
||||
|
||||
# Render transitions first (they're the actionable alerts)
|
||||
if prior_bands is None:
|
||||
lines.append(" (first run — recording baseline, no alerts)")
|
||||
elif not transitions:
|
||||
lines.append(" (no band changes since last scan)")
|
||||
else:
|
||||
for t in transitions:
|
||||
volume, from_band, to_band = t["volume"], t["from"], t["to"]
|
||||
if t["type"] == "escalation":
|
||||
lines.append(f" ⚠️ {volume}: {from_band} → {to_band} (ESCALATION)")
|
||||
else:
|
||||
lines.append(f" ✅ {volume}: {from_band} → {to_band} (RECOVERY)")
|
||||
|
||||
# Render all volumes with their bands
|
||||
lines.append("")
|
||||
for node_result in results:
|
||||
if not node_result["reachable"]:
|
||||
lines.append(f" ❌ {node_result['target']}: UNREACHABLE ({node_result['failure_kind']})")
|
||||
continue
|
||||
|
||||
lines.append(f" {node_result['target']}:")
|
||||
for vol in node_result["volumes"]:
|
||||
lines.append(f" {vol['mount']} ({vol['type']}): {vol['pct']}% ({vol['used']}/{vol['size']}, {vol['avail']} free) -> {vol['band']}")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
import argparse
|
||||
|
||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
|
||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests and host filesystems.")
|
||||
ap.add_argument("--json", action="store_true", help="machine-readable output")
|
||||
ap.add_argument("--hosts-only", action="store_true", help="scan host filesystems only")
|
||||
ap.add_argument("--guests-only", action="store_true", help="scan guests only (skip host filesystems)")
|
||||
args = ap.parse_args()
|
||||
|
||||
results = scan_fleet()
|
||||
# Scan guests (unless --hosts-only)
|
||||
guest_results = []
|
||||
if not args.hosts_only:
|
||||
guest_results = scan_fleet()
|
||||
|
||||
# Scan host filesystems (unless --guests-only)
|
||||
host_results = []
|
||||
current_bands = {}
|
||||
if not args.guests_only:
|
||||
host_results, current_bands = probe_host_filesystems()
|
||||
|
||||
# Read prior state and detect transitions
|
||||
prior_bands = read_state_file()
|
||||
transitions = detect_transitions(current_bands, prior_bands)
|
||||
|
||||
# Write new state
|
||||
write_state_file(current_bands)
|
||||
else:
|
||||
prior_bands = None
|
||||
transitions = []
|
||||
|
||||
if args.json:
|
||||
print(json.dumps(results, indent=2))
|
||||
# JSON output
|
||||
output = {
|
||||
"guests": guest_results,
|
||||
"hosts": host_results,
|
||||
"transitions": transitions,
|
||||
"prior_bands": prior_bands,
|
||||
}
|
||||
print(json.dumps(output, indent=2))
|
||||
else:
|
||||
print(render_results(results))
|
||||
# Human-readable output
|
||||
if guest_results:
|
||||
print(render_results(guest_results))
|
||||
|
||||
if host_results:
|
||||
print(render_host_results(host_results, transitions, prior_bands))
|
||||
|
||||
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
|
||||
# (a probe error means the probe itself failed, not just that the guest was unreachable)
|
||||
# Exit 0 if all probed (reachable or not), 1 if any probe error
|
||||
return 0
|
||||
|
||||
|
||||
|
||||
Executable
+227
@@ -0,0 +1,227 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Health-log freshness watchdog (dead-man's-switch).
|
||||
|
||||
Absence of logs raises no alarm. On 2026-09-26 the `gpu/` directory of
|
||||
SyslogSolution/health-logs went silent for 12 days unnoticed, because the only
|
||||
thing that would have noticed was the job that had stopped. This check lives on
|
||||
a DIFFERENT host (CT 100) from the producers, so a producer host that is dead,
|
||||
or a schedule that was deleted, still raises an alarm.
|
||||
|
||||
It asks the Gitea API for the newest commit touching each watched directory and
|
||||
fails when that commit is older than the directory's threshold.
|
||||
|
||||
Usage:
|
||||
health-log-freshness.py # check every watched directory
|
||||
health-log-freshness.py --json # machine-readable output
|
||||
|
||||
Exit: 0 = all fresh, 1 = at least one stale or unreachable, 2 = check could not run.
|
||||
|
||||
Environment:
|
||||
GITEA_URL default https://git.sysloggh.net
|
||||
GITEA_TOKEN API token (falls back to GITEA_PAT, then to the token
|
||||
embedded in ~/.git-credentials for that host)
|
||||
HEALTH_LOG_MAX_AGE_H override thresholds, e.g. HEALTH_LOG_MAX_AGE_GPU=12
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from datetime import datetime, timezone
|
||||
|
||||
REPO = "SyslogSolution/health-logs"
|
||||
GITEA_URL = os.environ.get("GITEA_URL", "https://git.sysloggh.net").rstrip("/")
|
||||
|
||||
# dir -> (threshold_hours, why)
|
||||
# Thresholds are sized for the producer cadence plus one missed run:
|
||||
# gpu/ runs every 6h -> 12h tolerates one miss, catches a second
|
||||
# litellm/ runs every 6h -> 18h (it is proven healthy; a looser bound avoids
|
||||
# noise while still catching a real stop)
|
||||
WATCHED: dict[str, tuple[float, str]] = {
|
||||
"gpu": (12.0, "gpu-self-heal.py, CT116 cron 2 */6 * * *"),
|
||||
"litellm": (18.0, "litellm-health-check.sh, CT116 cron 0 */6 * * *"),
|
||||
}
|
||||
|
||||
|
||||
def _threshold(directory: str, default: float) -> float:
|
||||
"""Allow HEALTH_LOG_MAX_AGE_<DIR> to override a threshold.
|
||||
|
||||
Documented override; used both operationally (tighten a bound while
|
||||
investigating) and in tests (force a stale verdict deterministically).
|
||||
"""
|
||||
raw = os.environ.get(f"HEALTH_LOG_MAX_AGE_{directory.upper()}")
|
||||
if raw is None:
|
||||
return default
|
||||
try:
|
||||
return float(raw)
|
||||
except ValueError:
|
||||
print(f"WARN: ignoring non-numeric HEALTH_LOG_MAX_AGE_{directory.upper()}={raw!r}")
|
||||
return default
|
||||
|
||||
# pm2/ is deliberately NOT watched: pm2-self-heal.prose.md claimed a health-logs
|
||||
# posting that its executor never implemented, and the claim was retired on
|
||||
# 2026-09-26 in favour of the contract-runner's per-run logs. The directory is
|
||||
# left as historical evidence, not as a live obligation.
|
||||
|
||||
|
||||
def _auth_candidates() -> list[str]:
|
||||
"""Authorization header values to try, in order.
|
||||
|
||||
Gitea accepts either an API token (``token <tok>``) or HTTP basic auth, and
|
||||
the value held under ``GITEA_TOKEN`` is not reliably an API token - on this
|
||||
host it is the git account's *password*, which sent as a token returns HTTP
|
||||
401. So return every candidate and let the caller use the first that works,
|
||||
rather than guessing and failing.
|
||||
"""
|
||||
candidates: list[str] = []
|
||||
for var in ("GITEA_TOKEN", "GITEA_PAT"):
|
||||
if os.environ.get(var):
|
||||
candidates.append(f"token {os.environ[var]}")
|
||||
candidates.append(
|
||||
"Basic "
|
||||
+ base64.b64encode(
|
||||
f"{_git_user()}:{os.environ[var]}".encode()
|
||||
).decode()
|
||||
)
|
||||
# ~/.git-credentials may hold this Gitea under several hostnames (the public
|
||||
# name and the internal IP both appear in this fleet), so accept any of them
|
||||
# rather than filtering on the configured host - that filter produced zero
|
||||
# candidates whenever GITEA_URL pointed at the internal address.
|
||||
try:
|
||||
with open(os.path.expanduser("~/.git-credentials")) as fh:
|
||||
for line in fh:
|
||||
m = re.match(r"https://([^:]+):([^@]+)@", line.strip())
|
||||
if m:
|
||||
raw = f"{m.group(1)}:{m.group(2)}".encode()
|
||||
candidates.append("Basic " + base64.b64encode(raw).decode())
|
||||
except OSError:
|
||||
pass
|
||||
seen, out = set(), []
|
||||
for c in candidates:
|
||||
if c not in seen:
|
||||
seen.add(c)
|
||||
out.append(c)
|
||||
return out
|
||||
|
||||
|
||||
def _git_user() -> str:
|
||||
return os.environ.get("GITEA_USER", "abiba-bot")
|
||||
|
||||
|
||||
def newest_commit_iso(directory: str, auth: list[str] | None) -> tuple[str | None, str]:
|
||||
"""Return (iso_timestamp, detail) for the newest commit touching `directory`."""
|
||||
url = (
|
||||
f"{GITEA_URL}/api/v1/repos/{REPO}/commits"
|
||||
f"?path={directory}&limit=1&stat=false"
|
||||
)
|
||||
req = urllib.request.Request(url, headers={"Accept": "application/json"})
|
||||
headers = list(auth or [])
|
||||
last = "no credential"
|
||||
for i, hdr in enumerate(headers or [None]):
|
||||
r = urllib.request.Request(url, headers={"Accept": "application/json"})
|
||||
if hdr:
|
||||
r.add_header("Authorization", hdr)
|
||||
try:
|
||||
with urllib.request.urlopen(r, timeout=20) as resp:
|
||||
data = json.loads(resp.read().decode("utf-8", "replace"))
|
||||
break
|
||||
except urllib.error.HTTPError as exc:
|
||||
last = f"HTTP {exc.code}"
|
||||
if exc.code not in (401, 403):
|
||||
return None, last
|
||||
except Exception as exc: # noqa: BLE001
|
||||
return None, repr(exc)
|
||||
else:
|
||||
return None, last
|
||||
|
||||
if not data:
|
||||
return None, "no commits"
|
||||
commit = data[0].get("commit", {})
|
||||
when = (
|
||||
(commit.get("committer") or {}).get("date")
|
||||
or (commit.get("author") or {}).get("date")
|
||||
)
|
||||
msg = (commit.get("message") or "").splitlines()[0][:60]
|
||||
return when, msg
|
||||
|
||||
|
||||
def main() -> int:
|
||||
as_json = "--json" in sys.argv
|
||||
auth = _auth_candidates()
|
||||
now = datetime.now(timezone.utc)
|
||||
failures: list[str] = []
|
||||
report: dict[str, dict] = {}
|
||||
|
||||
if not auth:
|
||||
print("FAIL: no Gitea credential available (GITEA_TOKEN/GITEA_PAT/~/.git-credentials)")
|
||||
return 2
|
||||
for directory, (default_age_h, why) in WATCHED.items():
|
||||
max_age_h = _threshold(directory, default_age_h)
|
||||
when, detail = newest_commit_iso(directory, auth)
|
||||
entry: dict = {"directory": directory, "producer": why, "max_age_h": max_age_h}
|
||||
|
||||
if when is None:
|
||||
entry.update(ok=False, reason=f"could not read newest commit: {detail}")
|
||||
failures.append(
|
||||
f"health-logs/{directory}/ could not be read ({detail}) — "
|
||||
f"a directory that cannot be read is indistinguishable from one that stopped"
|
||||
)
|
||||
else:
|
||||
try:
|
||||
ts = datetime.fromisoformat(when.replace("Z", "+00:00"))
|
||||
except ValueError:
|
||||
entry.update(ok=False, reason=f"unparseable timestamp {when!r}")
|
||||
failures.append(f"health-logs/{directory}/ timestamp unparseable: {when!r}")
|
||||
report[directory] = entry
|
||||
continue
|
||||
age_h = (now - ts).total_seconds() / 3600.0
|
||||
stale = age_h > max_age_h
|
||||
entry.update(
|
||||
ok=not stale,
|
||||
newest=when,
|
||||
age_h=round(age_h, 2),
|
||||
newest_commit=detail,
|
||||
)
|
||||
if stale:
|
||||
entry["reason"] = f"stale: {age_h:.1f}h > {max_age_h}h"
|
||||
failures.append(
|
||||
f"health-logs/{directory}/ is STALE: newest entry {when} "
|
||||
f"({age_h:.1f}h old, limit {max_age_h}h) from {why}"
|
||||
)
|
||||
report[directory] = entry
|
||||
|
||||
if as_json:
|
||||
print(json.dumps({"checked": report, "failures": failures}, indent=2))
|
||||
return 1 if failures else 0
|
||||
|
||||
print("Health-log freshness — dead-man's-switch")
|
||||
print("=" * 66)
|
||||
for directory, entry in report.items():
|
||||
mark = "✅" if entry.get("ok") else "❌"
|
||||
if entry.get("newest"):
|
||||
print(
|
||||
f"{mark} health-logs/{directory}/ newest {entry['newest']} "
|
||||
f"({entry['age_h']}h old, limit {entry['max_age_h']}h)"
|
||||
)
|
||||
print(f" last commit: {entry.get('newest_commit')}")
|
||||
else:
|
||||
print(f"{mark} health-logs/{directory}/ {entry.get('reason')}")
|
||||
print(f" producer: {entry['producer']}")
|
||||
|
||||
print("=" * 66)
|
||||
if failures:
|
||||
print("VERDICT: FAIL")
|
||||
for f in failures:
|
||||
print(f" - {f}")
|
||||
return 1
|
||||
print("VERDICT: PASS — every watched health-log directory is advancing")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Executable
+234
@@ -0,0 +1,234 @@
|
||||
#!/bin/bash
|
||||
# infrastructure-monitoring.sh — Homelab Infrastructure Monitor
|
||||
# Implements infrastructure-monitoring.prose.md (check-health section)
|
||||
#
|
||||
# Legs: Grafana, Prometheus, LiteLLM, PVE API (5 nodes), GPU exporters,
|
||||
# Docker Stats, PVE Exporter
|
||||
#
|
||||
# Design:
|
||||
# - Every target, port, path, and expected status is defined in code
|
||||
# - Liveness rule: any HTTP status = ALIVE for auth-gated/redirect endpoints;
|
||||
# only connection failures (000/timeout) = probe-failed
|
||||
# - Bare-200 rule: expected status must match exactly (200); anything else = alert
|
||||
# - PVE API uses -k flag (self-signed certs), probes /api2/json/version
|
||||
# - Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116, probed via SSH
|
||||
# with one retry at longer timeout (25s connect, 30s max) to distinguish
|
||||
# transient timeout from host-down
|
||||
# - Non-zero exit naming every failed target; no "OK" summary when any leg failed
|
||||
#
|
||||
# Output shape per leg:
|
||||
# ✅ <name>: alive
|
||||
# 🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)
|
||||
#
|
||||
# Failure kinds: timeout | refused | tls | unexpected:<code> (printed in the failure line)
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
# ── Configuration (documented in infrastructure-monitoring.prose.md) ────────
|
||||
# Change these in ONE place; test_infra_monitoring.sh asserts against these.
|
||||
|
||||
GRAFANA_HOST="192.168.68.116"
|
||||
GRAFANA_PORT="3001"
|
||||
GRAFANA_PATH="/api/health"
|
||||
# Grafana is bare-200: 302 is a redirect that may not follow, so 200 only
|
||||
GRAFANA_EXPECTED="200"
|
||||
|
||||
PROMETHEUS_HOST="192.168.68.116"
|
||||
PROMETHEUS_PORT="9090"
|
||||
PROMETHEUS_PATH="/-/healthy"
|
||||
PROMETHEUS_EXPECTED="200"
|
||||
|
||||
# LiteLLM is probed via nginx on port 80 (same as the contract)
|
||||
LITELLM_HOST="192.168.68.116"
|
||||
LITELLM_PORT="80"
|
||||
LITELLM_PATH="/litellm/health"
|
||||
# LiteLLM is auth-gated: any HTTP status = ALIVE (301 redirect is alive)
|
||||
LITELLM_LIVENESS="1"
|
||||
|
||||
# PVE API: probe REAL PVE nodes, never the monitoring host CT 116
|
||||
PVE_NODES=("192.168.68.9" "192.168.68.12" "192.168.68.6" "192.168.68.15" "192.168.68.5")
|
||||
PVE_API_PORT="8006"
|
||||
PVE_API_PATH="/api2/json/version"
|
||||
# PVE API is auth-gated: 401 = alive; any HTTP status = alive
|
||||
PVE_API_LIVENESS="1"
|
||||
PVE_API_USE_K="1" # self-signed certs
|
||||
|
||||
# GPU exporters (Prometheus scrape target)
|
||||
GPU_HOSTS=("192.168.68.8" "192.168.68.110" "192.168.68.15")
|
||||
GPU_PORT="9400"
|
||||
GPU_PATH="/metrics"
|
||||
GPU_EXPECTED="200"
|
||||
|
||||
# Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116
|
||||
DOCKER_STATS_PORT="9324" # harness-docker-stats (docker_container_* metrics)
|
||||
PVE_EXPORTER_PORT="9221" # harness-pve-exporter (5 pve_* metrics)
|
||||
CT116_SSH_HOST="192.168.68.116"
|
||||
# Both are bare-200: 404 = container not yet started
|
||||
DOCKER_STATS_EXPECTED="200|404"
|
||||
PVE_EXPORTER_EXPECTED="200|404"
|
||||
|
||||
# ── Probe Functions ─────────────────────────────────────────────────────────
|
||||
|
||||
# probe_http <host> <port> <path> <expected_pattern> [use_k] [ssh_host] [scheme] [liveness]
|
||||
# Returns 0 if probe succeeds (matches expected or liveness), 1 if probe-failed.
|
||||
# Prints the result line.
|
||||
#
|
||||
# FIX C1: The kind value is computed and printed in the failure line.
|
||||
# FIX C2: SSH retry logic is in the first attempt branch (not unreachable).
|
||||
|
||||
LAST_KIND=""
|
||||
probe_http() {
|
||||
local host="$1" port="$2" path="$3" expected="$4"
|
||||
local use_k="${5:-}" ssh_host="${6:-}" scheme="${7:-http}" liveness="${8:-0}"
|
||||
local url="${scheme}://${host}:${port}${path}"
|
||||
local code="" kind=""
|
||||
LAST_KIND=""
|
||||
|
||||
# Single invocation that captures both output and status
|
||||
if [ -n "$ssh_host" ]; then
|
||||
out=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||
rc=$?
|
||||
else
|
||||
out=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} "$url" 2>/dev/null)
|
||||
rc=$?
|
||||
fi
|
||||
code=$(printf '%s' "$out" | tr -d '[:space:]')
|
||||
|
||||
# Classify failure kind and retry if needed
|
||||
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||
# Distinguish timeout from TLS error from refused
|
||||
case "$rc" in
|
||||
35|51|58|59|60|77|83) kind="tls" ;;
|
||||
*) kind="timeout" ;;
|
||||
esac
|
||||
# Retry once at longer timeout (25s connect, 30s max)
|
||||
if [ -n "$ssh_host" ]; then
|
||||
code=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||
rc=$?
|
||||
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||
else
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} "$url" 2>/dev/null)
|
||||
rc=$?
|
||||
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||
# On retry, classify: still 000 = keep existing kind (or timeout if empty), unexpected status = refused
|
||||
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||
[ -z "$kind" ] && kind="timeout"
|
||||
elif ! echo "$code" | grep -qE "^(${expected})$"; then
|
||||
kind="refused"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# Check result
|
||||
if [ -n "$code" ] && [ "$code" != "000" ]; then
|
||||
if [ "$liveness" = "1" ]; then
|
||||
# Any HTTP status = ALIVE for auth-gated/redirect endpoints
|
||||
return 0
|
||||
else
|
||||
# Bare-200 or specific expected pattern
|
||||
if echo "$code" | grep -qE "^(${expected})$"; then
|
||||
return 0
|
||||
else
|
||||
kind="unexpected:$code"
|
||||
LAST_KIND="$kind"
|
||||
return 1
|
||||
fi
|
||||
fi
|
||||
else
|
||||
[ -z "$kind" ] && kind="refused"
|
||||
LAST_KIND="$kind"
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
|
||||
# ── Main ────────────────────────────────────────────────────────────────────
|
||||
|
||||
FAILED=()
|
||||
FAILED_KIND=()
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
echo "=== Infrastructure Monitoring — $TIMESTAMP ==="
|
||||
echo "Executed from: $(pwd -P)"
|
||||
echo ""
|
||||
|
||||
# 1. Grafana (CT 116 :3001 /api/health) — bare-200
|
||||
if probe_http "$GRAFANA_HOST" "$GRAFANA_PORT" "$GRAFANA_PATH" "$GRAFANA_EXPECTED"; then
|
||||
echo " ✅ Grafana: alive"
|
||||
else
|
||||
echo " 🔴 Grafana: probe-failed: ${GRAFANA_HOST}:${GRAFANA_PORT} (expected ${GRAFANA_EXPECTED}) (<${LAST_KIND}>)"
|
||||
FAILED+=("grafana")
|
||||
fi
|
||||
|
||||
# 2. Prometheus (CT 116 :9090 /-/healthy) — bare-200
|
||||
if probe_http "$PROMETHEUS_HOST" "$PROMETHEUS_PORT" "$PROMETHEUS_PATH" "$PROMETHEUS_EXPECTED"; then
|
||||
echo " ✅ Prometheus: alive"
|
||||
else
|
||||
echo " 🔴 Prometheus: probe-failed: ${PROMETHEUS_HOST}:${PROMETHEUS_PORT} (expected ${PROMETHEUS_EXPECTED}) (<${LAST_KIND}>)"
|
||||
FAILED+=("prometheus")
|
||||
fi
|
||||
|
||||
# 3. LiteLLM (CT 116 :80/litellm/health via nginx) — liveness (any HTTP = alive)
|
||||
if probe_http "$LITELLM_HOST" "$LITELLM_PORT" "$LITELLM_PATH" "" "" "" "http" "$LITELLM_LIVENESS"; then
|
||||
echo " ✅ LiteLLM: alive"
|
||||
else
|
||||
echo " 🔴 LiteLLM: probe-failed: ${LITELLM_HOST}:${LITELLM_PORT}${LITELLM_PATH} (any-HTTP liveness) (<${LAST_KIND}>)"
|
||||
FAILED+=("litellm")
|
||||
fi
|
||||
|
||||
# 4. PVE API (5 real nodes :8006 /api2/json/version, -k, liveness)
|
||||
PVE_FAILED=()
|
||||
for node in "${PVE_NODES[@]}"; do
|
||||
if probe_http "$node" "$PVE_API_PORT" "$PVE_API_PATH" "" "$PVE_API_USE_K" "" "https" "$PVE_API_LIVENESS"; then
|
||||
echo " ✅ PVE API ${node}: alive"
|
||||
else
|
||||
echo " 🔴 PVE API ${node}: probe-failed: ${node}:${PVE_API_PORT} (any-HTTP liveness, -k for self-signed) (<${LAST_KIND}>)"
|
||||
PVE_FAILED+=("$node")
|
||||
fi
|
||||
done
|
||||
if [ ${#PVE_FAILED[@]} -gt 0 ]; then
|
||||
FAILED+=("pve-api: ${PVE_FAILED[*]}")
|
||||
fi
|
||||
|
||||
# 5. GPU exporters (:9400/metrics) — bare-200
|
||||
GPU_FAILED=()
|
||||
for host in "${GPU_HOSTS[@]}"; do
|
||||
if probe_http "$host" "$GPU_PORT" "$GPU_PATH" "$GPU_EXPECTED"; then
|
||||
echo " ✅ GPU exporter ${host}: alive"
|
||||
else
|
||||
echo " 🔴 GPU exporter ${host}: probe-failed: ${host}:${GPU_PORT} (expected 200) (<${LAST_KIND}>)"
|
||||
GPU_FAILED+=("$host")
|
||||
fi
|
||||
done
|
||||
if [ ${#GPU_FAILED[@]} -gt 0 ]; then
|
||||
FAILED+=("gpu-exporters: ${GPU_FAILED[*]}")
|
||||
fi
|
||||
|
||||
# 6. Docker Stats (CT 116 :9324, 127.0.0.1 via SSH) — 200|404
|
||||
if probe_http "127.0.0.1" "$DOCKER_STATS_PORT" "/" "$DOCKER_STATS_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||
echo " ✅ Docker Stats: alive"
|
||||
else
|
||||
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:${DOCKER_STATS_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||
FAILED+=("docker-stats")
|
||||
fi
|
||||
|
||||
# 7. PVE Exporter (CT 116 :9221, 127.0.0.1 via SSH) — 200|404
|
||||
if probe_http "127.0.0.1" "$PVE_EXPORTER_PORT" "/" "$PVE_EXPORTER_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||
echo " ✅ PVE Exporter: alive"
|
||||
else
|
||||
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:${PVE_EXPORTER_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||
FAILED+=("pve-exporter")
|
||||
fi
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
|
||||
echo ""
|
||||
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||
echo " ✅ All legs OK"
|
||||
exit 0
|
||||
else
|
||||
for f in "${FAILED[@]}"; do
|
||||
echo " 🔴 FAILED: $f"
|
||||
done
|
||||
exit 1
|
||||
fi
|
||||
+156
-36
@@ -35,7 +35,12 @@ def run_command(cmd, timeout=15):
|
||||
return 1, "", str(e)
|
||||
|
||||
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
|
||||
"""Probe HTTP endpoint and return status code"""
|
||||
"""Probe HTTP endpoint and return (status_code, failure_kind)
|
||||
|
||||
Returns:
|
||||
(code, None) if successful or HTTP response received
|
||||
(000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc.
|
||||
"""
|
||||
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
|
||||
if method == "POST":
|
||||
cmd += " -X POST"
|
||||
@@ -47,15 +52,64 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
||||
cmd += " -L"
|
||||
cmd += " '" + url + "'"
|
||||
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
if rc != 0 and "TIMEOUT" not in stderr:
|
||||
return 000 # Connection failed
|
||||
try:
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
if rc != 0:
|
||||
# Check if this is a timeout from run_command (rc=1, stderr="TIMEOUT")
|
||||
if rc == 1 and stderr == "TIMEOUT":
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
# Otherwise, determine failure kind from curl exit code
|
||||
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
||||
elif rc == 28:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
elif rc == 7:
|
||||
return (000, "connection refused")
|
||||
elif rc == 6:
|
||||
return (000, "dns failure")
|
||||
elif rc == 35:
|
||||
return (000, "ssl error")
|
||||
elif rc == 52:
|
||||
return (000, "empty response")
|
||||
else:
|
||||
return (000, "curl exit " + str(rc))
|
||||
return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response")
|
||||
except subprocess.TimeoutExpired:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
|
||||
def check_host_health(host_ip):
|
||||
"""Check if the GPU host's llama-chat-api health endpoint is reachable
|
||||
|
||||
return int(stdout) if stdout.isdigit() else 000
|
||||
Returns: (healthy: bool, detail: str)
|
||||
"""
|
||||
code, _ = probe_http("http://" + host_ip + ":8080/health", timeout=10)
|
||||
if code == 200:
|
||||
return True, "host healthy (200)"
|
||||
elif code == 000:
|
||||
return False, "host unreachable (timeout or refused)"
|
||||
else:
|
||||
return False, "host unhealthy (HTTP " + str(code) + ")"
|
||||
|
||||
|
||||
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
||||
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
||||
cmd = "curl -s -m " + str(timeout)
|
||||
if method == "POST":
|
||||
cmd += " -X POST"
|
||||
if bearer_token:
|
||||
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
|
||||
if data:
|
||||
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
|
||||
cmd += " '" + url + "'"
|
||||
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
# Return first 200 chars, single line
|
||||
body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else ""
|
||||
return body
|
||||
|
||||
|
||||
def check_liveliness():
|
||||
"""Step 1: Liveliness probe"""
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
|
||||
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
|
||||
|
||||
def check_containers():
|
||||
@@ -78,33 +132,86 @@ def check_model_probes():
|
||||
|
||||
results = []
|
||||
|
||||
# Host health mapping: model -> host IP
|
||||
model_hosts = {
|
||||
"gpu-dense": "192.168.68.8", # RTX 3090
|
||||
"gpu-vision": "192.168.68.110", # RTX 5070
|
||||
"strix-moe": "192.168.68.15" # Strix Halo
|
||||
}
|
||||
|
||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||
# Single-host aliases: 30s timeout
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
host_ip = model_hosts[model]
|
||||
# Single-host aliases: 30s initial timeout, retry once at 90s on failure
|
||||
# Worst-case prefill ~76s, so 90s retry ensures we cover it
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
|
||||
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
first_kind = None
|
||||
if code == 000 and failure_kind:
|
||||
first_kind = failure_kind
|
||||
time.sleep(1)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=90)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Both attempts failed - check host health to distinguish busy from down
|
||||
host_healthy, host_detail = check_host_health(host_ip)
|
||||
if host_healthy:
|
||||
results.append((model, False, "busy (completion timed out after retry; " + host_detail + ")"))
|
||||
else:
|
||||
# Host unreachable - report both kinds
|
||||
if first_kind:
|
||||
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
|
||||
else:
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
|
||||
elif code == 200:
|
||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
elif code in (401, 403):
|
||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||
timeout=10)
|
||||
alias = "monitor-20260813"
|
||||
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||
else:
|
||||
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
|
||||
# Pool alias (syslog-auto): 60s timeout, retry once on 000
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
if code == 000:
|
||||
if code == 000 and failure_kind:
|
||||
# Retry once with same timeout
|
||||
time.sleep(1)
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
if code == 000 and failure_kind:
|
||||
results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)"))
|
||||
elif code in (401, 403):
|
||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||
timeout=10)
|
||||
alias = "monitor-20260813"
|
||||
results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||
else:
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
else:
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
|
||||
return results
|
||||
|
||||
@@ -131,33 +238,33 @@ def check_admin_key_list():
|
||||
# Try to parse the response
|
||||
try:
|
||||
data = json.loads(stdout)
|
||||
# Response is a dict with "keys" field
|
||||
# Response is a dict with "keys" (paginated list) and "total_count" fields
|
||||
if isinstance(data, dict) and "keys" in data:
|
||||
key_count = len(data["keys"])
|
||||
key_count = data.get("total_count", len(data["keys"]))
|
||||
elif isinstance(data, list):
|
||||
key_count = len(data)
|
||||
else:
|
||||
key_count = 0
|
||||
if key_count == 0:
|
||||
return "Admin Key List", False, "admin-call-failed (empty response)"
|
||||
return "Admin Key List", True, str(key_count) + " keys"
|
||||
return "Admin Key List", True, str(key_count) + " total (" + str(len(data.get("keys", []) if isinstance(data, dict) else data)) + " on page 1)" if isinstance(data, dict) else str(key_count) + " total"
|
||||
except Exception as e:
|
||||
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
|
||||
|
||||
def check_github_status():
|
||||
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
|
||||
code = probe_http("https://status.github.com/api/status.json", timeout=15)
|
||||
code, _ = probe_http("https://status.github.com/api/status.json", timeout=15)
|
||||
# GitHub status API returns 301 redirect, which is expected behavior
|
||||
return "GitHub Status", code == 301, str(code)
|
||||
|
||||
def check_prometheus():
|
||||
"""Step 4: Prometheus health"""
|
||||
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
|
||||
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
|
||||
|
||||
def check_grafana():
|
||||
"""Step 9: Grafana health"""
|
||||
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
|
||||
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
|
||||
|
||||
def check_docker_stats():
|
||||
@@ -181,6 +288,7 @@ def main():
|
||||
print("")
|
||||
|
||||
all_pass = True
|
||||
degraded = [] # Track degraded (busy) checks
|
||||
|
||||
# Run all checks
|
||||
checks = [
|
||||
@@ -200,9 +308,15 @@ def main():
|
||||
# Model probes
|
||||
model_results = check_model_probes()
|
||||
for name, passed, detail in model_results:
|
||||
status = "✅" if passed else "❌"
|
||||
# Check if this is a busy (degraded) verdict
|
||||
if not passed and detail.startswith("busy "):
|
||||
status = "⚠️"
|
||||
degraded.append(name)
|
||||
else:
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
if not passed:
|
||||
# Only set all_pass=False for real failures (not busy)
|
||||
if not passed and not detail.startswith("busy "):
|
||||
all_pass = False
|
||||
|
||||
# Admin key list
|
||||
@@ -228,10 +342,16 @@ def main():
|
||||
|
||||
print("")
|
||||
if all_pass:
|
||||
print("✅ All checks passed")
|
||||
if degraded:
|
||||
print("✅ All checks passed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||
else:
|
||||
print("✅ All checks passed")
|
||||
return 0
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
if degraded:
|
||||
print("❌ Some checks failed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
return 1
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
+1
-1
@@ -12,7 +12,6 @@ set -euo pipefail
|
||||
declare -A CT_NODES=(
|
||||
# amdpve (192.168.68.15)
|
||||
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
|
||||
[112]=amdpve # tanko
|
||||
[113]=amdpve # baggy
|
||||
[115]=amdpve # scottdenya
|
||||
[120]=amdpve # adguard2 (added 2026-09-12)
|
||||
@@ -21,6 +20,7 @@ declare -A CT_NODES=(
|
||||
[102]=minipve # adguard (was acerpve)
|
||||
[104]=minipve # authentik
|
||||
[110]=minipve # gitea
|
||||
[112]=minipve # tanko (was amdpve — migrated 2026-09-27; live per pvesh)
|
||||
[116]=minipve # syslog-api
|
||||
[119]=minipve # infisical-vault
|
||||
# storepve (192.168.68.6)
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
#!/bin/bash
|
||||
# pm2-self-heal — hourly PM2 process check
|
||||
# Part of the pm2-self-heal prose contract
|
||||
# Alerts via Telegram (abiba-zulip decommissioned 2026-07-04)
|
||||
# Alerts via Telegram (primary) and Zulip DM (secondary, abiba-zulip restored 2026-08-03)
|
||||
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||
|
||||
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
||||
@@ -46,9 +46,75 @@ if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
|
||||
fi
|
||||
fi
|
||||
|
||||
# Check abiba-zulip (live Zulip bridge, heartbeating)
|
||||
ZULIP_LINE=$(echo "$STATUS" | grep "abiba-zulip")
|
||||
ZULIP_STATUS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
ZULIP_RESTARTS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$ZULIP_STATUS" != "online" ]; then
|
||||
pm2 restart abiba-zulip > /dev/null 2>&1
|
||||
sleep 3
|
||||
ZULIP_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-zulip")
|
||||
ZULIP_STATUS2=$(echo "$ZULIP_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$ZULIP_STATUS2" = "online" ]; then
|
||||
msg="⚠️ abiba-zulip was **$ZULIP_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 abiba-zulip **failed restart** (was $ZULIP_STATUS, still $ZULIP_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$ZULIP_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ abiba-zulip has **$ZULIP_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Check gitea-runner
|
||||
GITEA_LINE=$(echo "$STATUS" | grep "gitea-runner")
|
||||
GITEA_STATUS=$(echo "$GITEA_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
GITEA_RESTARTS=$(echo "$GITEA_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$GITEA_STATUS" != "online" ]; then
|
||||
pm2 restart gitea-runner > /dev/null 2>&1
|
||||
sleep 3
|
||||
GITEA_LINE2=$(pm2 status --no-color 2>/dev/null | grep "gitea-runner")
|
||||
GITEA_STATUS2=$(echo "$GITEA_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$GITEA_STATUS2" = "online" ]; then
|
||||
msg="⚠️ gitea-runner was **$GITEA_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 gitea-runner **failed restart** (was $GITEA_STATUS, still $GITEA_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$GITEA_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ gitea-runner has **$GITEA_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Check zulip-watchdog
|
||||
WATCHDOG_LINE=$(echo "$STATUS" | grep "zulip-watchdog")
|
||||
WATCHDOG_STATUS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
WATCHDOG_RESTARTS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$WATCHDOG_STATUS" != "online" ]; then
|
||||
pm2 restart zulip-watchdog > /dev/null 2>&1
|
||||
sleep 3
|
||||
WATCHDOG_LINE2=$(pm2 status --no-color 2>/dev/null | grep "zulip-watchdog")
|
||||
WATCHDOG_STATUS2=$(echo "$WATCHDOG_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$WATCHDOG_STATUS2" = "online" ]; then
|
||||
msg="⚠️ zulip-watchdog was **$WATCHDOG_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 zulip-watchdog **failed restart** (was $WATCHDOG_STATUS, still $WATCHDOG_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$WATCHDOG_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ zulip-watchdog has **$WATCHDOG_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Log check
|
||||
{
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS alerts=${ALERTS:+yes}"
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS zulip=$ZULIP_STATUS gitea=$GITEA_STATUS watchdog=$WATCHDOG_STATUS alerts=${ALERTS:+yes}"
|
||||
[ -n "$ALERTS" ] && echo "$ALERTS"
|
||||
} >> "$LOG"
|
||||
|
||||
|
||||
@@ -47,8 +47,8 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
|
||||
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
|
||||
|
||||
**Proxmox Cluster "Tabiri" (5 nodes):**
|
||||
- amdpve (192.168.68.15): kagentz, tanko, baggy, scottdenya, adguard2
|
||||
- minipve (192.168.68.12): abiba, adguard, authentik, gitea, syslog-api, infisical-vault
|
||||
- amdpve (192.168.68.15): kagentz, baggy, scottdenya, adguard2
|
||||
- minipve (192.168.68.12): abiba, tanko, adguard, authentik, gitea, syslog-api, infisical-vault
|
||||
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
|
||||
- acerpve (192.168.68.9): llm-gpu
|
||||
- ocupve (192.168.68.5): ocu-llm
|
||||
|
||||
+16
-1
@@ -135,7 +135,22 @@ fi
|
||||
|
||||
echo " Cross-contract: $WARNINGS total warnings across all checks"
|
||||
|
||||
# ── 4. Summary ──
|
||||
# ── 4. Committed-credential scan ──
|
||||
# The 2026-09-17 purge removed six live credentials that had sat in .md prose
|
||||
# and scripts for weeks. This step makes that class of commit FAIL the gate
|
||||
# instead of printing a warning. Patterns: scripts/secret-patterns.tsv.
|
||||
# Only deliberate synthetic examples may be listed in scripts/secret-allowlist.tsv,
|
||||
# each with a reason. Run `bash scripts/secret-scan.sh --staged` before committing.
|
||||
echo ""
|
||||
echo "── 4. Secret scan (committed credentials) ──"
|
||||
if bash scripts/secret-scan.sh; then
|
||||
echo " ✅ No committed credentials"
|
||||
else
|
||||
echo " ❌ COMMITTED CREDENTIAL DETECTED"
|
||||
FAILED=1
|
||||
fi
|
||||
|
||||
# ── 5. Summary ──
|
||||
echo ""
|
||||
echo "═══════════════════════════════════"
|
||||
if [ $FAILED -eq 1 ]; then
|
||||
|
||||
Executable
+166
@@ -0,0 +1,166 @@
|
||||
#!/bin/bash
|
||||
# proxmox-monitor.sh — Proxmox Cluster + Docker Monitoring Health Check
|
||||
# Implements proxmox-monitor.prose.md (check-health section)
|
||||
#
|
||||
# Legs: Prometheus, Grafana, Docker Stats Exporter, PVE Exporter, PBS GC
|
||||
# All legs must return 200 for healthy status.
|
||||
#
|
||||
# Run: bash scripts/proxmox-monitor.sh
|
||||
# Exits 0 if all probes pass, 1 if any fails.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
CT116_HOST="192.168.68.116"
|
||||
FAILED=()
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
|
||||
echo "=== Proxmox Monitor — $TIMESTAMP ==="
|
||||
echo "Executed from: $(pwd -P)"
|
||||
echo ""
|
||||
|
||||
# 1. Prometheus health (bound to 0.0.0.0:9090 on .116)
|
||||
PROM_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:9090/-/healthy 2>/dev/null)
|
||||
PROM_CODE=$(printf '%s' "$PROM_CODE" | tr -d '[:space:]')
|
||||
[ -n "$PROM_CODE" ] || PROM_CODE="000"
|
||||
|
||||
if [ "$PROM_CODE" = "200" ]; then
|
||||
echo " ✅ Prometheus: alive"
|
||||
else
|
||||
echo " 🔴 Prometheus: probe-failed: ${CT116_HOST}:9090 (expected 200, got ${PROM_CODE})"
|
||||
FAILED+=("prometheus")
|
||||
fi
|
||||
|
||||
# 2. Grafana health (bound to 0.0.0.0:3001 on .116)
|
||||
GRAF_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:3001/api/health 2>/dev/null)
|
||||
GRAF_CODE=$(printf '%s' "$GRAF_CODE" | tr -d '[:space:]')
|
||||
[ -n "$GRAF_CODE" ] || GRAF_CODE="000"
|
||||
|
||||
if [ "$GRAF_CODE" = "200" ]; then
|
||||
echo " ✅ Grafana: alive"
|
||||
else
|
||||
echo " 🔴 Grafana: probe-failed: ${CT116_HOST}:3001 (expected 200, got ${GRAF_CODE})"
|
||||
FAILED+=("grafana")
|
||||
fi
|
||||
|
||||
# 3. Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — probe from .116 localhost)
|
||||
DOCKER_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics" 2>/dev/null)
|
||||
DOCKER_CODE=$(printf '%s' "$DOCKER_CODE" | tr -d '[:space:]')
|
||||
[ -n "$DOCKER_CODE" ] || DOCKER_CODE="000"
|
||||
|
||||
if [ "$DOCKER_CODE" = "200" ]; then
|
||||
echo " ✅ Docker Stats: alive"
|
||||
else
|
||||
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:9324 (expected 200, got ${DOCKER_CODE})"
|
||||
FAILED+=("docker-stats")
|
||||
fi
|
||||
|
||||
# 4. PVE Exporter (bound to 127.0.0.1:9221 on .116 — probe from .116 localhost)
|
||||
PVE_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics" 2>/dev/null)
|
||||
PVE_CODE=$(printf '%s' "$PVE_CODE" | tr -d '[:space:]')
|
||||
[ -n "$PVE_CODE" ] || PVE_CODE="000"
|
||||
|
||||
if [ "$PVE_CODE" = "200" ]; then
|
||||
echo " ✅ PVE Exporter: alive"
|
||||
else
|
||||
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:9221 (expected 200, got ${PVE_CODE})"
|
||||
FAILED+=("pve-exporter")
|
||||
fi
|
||||
|
||||
# 5. PBS GC liveness (storepve-datastore GC must have run within 48h)
|
||||
# Use absolute path for pct to avoid PATH issues in non-interactive ssh
|
||||
PBS_GC_OUTPUT=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192.168.68.6 \
|
||||
"/sbin/pct exec 107 -- /sbin/proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null)
|
||||
PBS_GC_OUTPUT=$(printf '%s' "$PBS_GC_OUTPUT" | tr -d '[:space:]')
|
||||
[ -n "$PBS_GC_OUTPUT" ] || PBS_GC_OUTPUT="000"
|
||||
|
||||
if [ "$PBS_GC_OUTPUT" = "000" ]; then
|
||||
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
# Parse the JSON to get storepve-datastore's state with four distinct outcomes:
|
||||
# 1. probe-failed: non-zero ssh status / empty / unparseable JSON
|
||||
# 2. running: collection in progress (last-run-endtime absent or 0, but upid present)
|
||||
# 3. stale: no completed run within 48h
|
||||
# 4. healthy: completed within 48h
|
||||
PBS_GC_RESULT=$(echo "$PBS_GC_OUTPUT" | python3 -c "
|
||||
import sys, json
|
||||
try:
|
||||
data = json.load(sys.stdin)
|
||||
for store in data:
|
||||
if store['store'] == 'storepve-datastore':
|
||||
endtime = store.get('last-run-endtime')
|
||||
upid = store.get('upid')
|
||||
pending = store.get('pending-bytes', 0)
|
||||
|
||||
# State 2: Running (collection in progress) — last-run-endtime absent while run is in progress
|
||||
if (endtime is None or endtime == 0) and upid is not None:
|
||||
print(f'running|{pending}')
|
||||
break
|
||||
|
||||
# State 3: No completed run (never-run or stale)
|
||||
if endtime is None or endtime == 0:
|
||||
print(f'no-completed-run|{pending}')
|
||||
break
|
||||
|
||||
# States 3 & 4: Completed (has endtime)
|
||||
print(f'completed|{endtime}|{pending}')
|
||||
break
|
||||
else:
|
||||
print(f'absent|0')
|
||||
except json.JSONDecodeError:
|
||||
print(f'unparseable|0')
|
||||
" 2>/dev/null)
|
||||
|
||||
# Parse the state|endtime|pending format
|
||||
PBS_GC_STATE=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1)
|
||||
|
||||
if [ "$PBS_GC_STATE" = "unparseable" ]; then
|
||||
# State 1: probe-failed (unparseable JSON)
|
||||
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)"
|
||||
FAILED+=("pbs-gc")
|
||||
elif [ "$PBS_GC_STATE" = "absent" ]; then
|
||||
# State 1: probe-failed (store not found)
|
||||
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (storepve-datastore not found)"
|
||||
FAILED+=("pbs-gc")
|
||||
elif [ "$PBS_GC_STATE" = "running" ]; then
|
||||
# State 2: collection in progress — do NOT fail
|
||||
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||
echo " ⏳ PBS GC: running (started: in-progress, pending-bytes: ${PENDING_BYTES} B)"
|
||||
elif [ "$PBS_GC_STATE" = "no-completed-run" ]; then
|
||||
# State 3: no completed run within 48h
|
||||
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||
echo " 🔴 PBS GC: no completed run within 48h (pending-bytes: ${PENDING_BYTES} B)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
# States 3 & 4: completed (has endtime)
|
||||
LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f3)
|
||||
|
||||
# Convert epoch to age in hours
|
||||
NOW_EPOCH=$(date -u +%s)
|
||||
AGE_HOURS=$(( (NOW_EPOCH - LAST_RUN_ENDTIME) / 3600 ))
|
||||
|
||||
if [ $AGE_HOURS -gt 48 ]; then
|
||||
# State 3: stale (no completed run within 48h)
|
||||
echo " 🔴 PBS GC: stale — last completed run was ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
# State 4: healthy (completed within 48h)
|
||||
echo " ✅ PBS GC: healthy — last completed run ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||
echo " ✅ All legs OK"
|
||||
exit 0
|
||||
else
|
||||
for f in "${FAILED[@]}"; do
|
||||
echo " 🔴 FAILED: $f"
|
||||
done
|
||||
exit 1
|
||||
fi
|
||||
Executable
+209
@@ -0,0 +1,209 @@
|
||||
#!/usr/bin/env bash
|
||||
# revision-preflight.sh — prove the copy a contract is about to execute is the
|
||||
# copy that is merged.
|
||||
#
|
||||
# Usage:
|
||||
# revision-preflight.sh [options] <script-path> <clone-path>
|
||||
#
|
||||
# Options:
|
||||
# --ref <ref> Ref to compare against (default: origin/master)
|
||||
# --no-fetch Do not refresh the ref first (see FRESHNESS)
|
||||
# --fetch-timeout <secs> Bound the default fetch (default: 20; 0 = no bound)
|
||||
# --quiet Print nothing on success
|
||||
# -h, --help Show this help
|
||||
#
|
||||
# Exit codes:
|
||||
# 0 the executing script byte-matches <ref>:<repo-relative-path>
|
||||
# 1 the copy is NOT the merged one -> REASON=mismatch:<class>
|
||||
# 2 the check could not be performed -> REASON=cannot-verify:<class>
|
||||
#
|
||||
# Every non-zero exit prints one machine-readable line
|
||||
# REASON=<class>
|
||||
# followed by the human explanation. The two top-level classes are deliberately
|
||||
# distinct: "I could not check" is a different situation from "this copy is
|
||||
# wrong", and an operator must never have to guess which they are looking at.
|
||||
#
|
||||
# cannot-verify:fetch-failed the remote could not be reached (or timed out)
|
||||
# cannot-verify:ref-unresolvable <ref> does not exist in the clone
|
||||
# mismatch:path-absent the script does not exist in <ref>
|
||||
# mismatch:content the script differs from <ref>
|
||||
# mismatch:detached-head the clone is on a detached HEAD
|
||||
# mismatch:clone-ahead local HEAD is ahead of <ref> (mid-review?)
|
||||
#
|
||||
# FRESHNESS
|
||||
# A guard is only as good as the ref it compares against. On 2026-09-25 a
|
||||
# stale local origin/master made an ancestry check on this fleet report
|
||||
# "unlanded work" for a branch that had in fact merged, and it would equally
|
||||
# have passed a stale script as current. So by default this guard FETCHES the
|
||||
# remote before comparing, bounded by --fetch-timeout so a hung remote cannot
|
||||
# block a scheduled contract. With --no-fetch it compares against whatever the
|
||||
# local ref points at and says so out loud; it never silently assumes
|
||||
# freshness.
|
||||
#
|
||||
# FAIL CLOSED
|
||||
# An unresolvable path or ref is a FAILURE, never a warning. "Cannot verify"
|
||||
# is precisely the state a stale or hand-edited copy produces, so treating it
|
||||
# as success would defeat the guard. The original draft did exactly that: it
|
||||
# resolved the master revision with
|
||||
# `git show origin/master:$(basename "$SCRIPT")`, which drops the scripts/
|
||||
# prefix, queries the repo root, fails, and exited 0 — passing a script that
|
||||
# exists in no revision at all.
|
||||
#
|
||||
# WHICH CLONE
|
||||
# Pass the clone the contract is actually executing from. See
|
||||
# docs/contract-execution-pinning.md for which clone each contract pins.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
REF="origin/master"
|
||||
FETCH=1
|
||||
QUIET=0
|
||||
FETCH_TIMEOUT="${REVISION_PREFLIGHT_FETCH_TIMEOUT:-20}"
|
||||
|
||||
usage() {
|
||||
sed -n '2,55p' "$0" | sed 's/^# \{0,1\}//'
|
||||
}
|
||||
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--ref)
|
||||
[[ $# -ge 2 ]] || { echo "revision-preflight: --ref needs a value" >&2; exit 2; }
|
||||
REF="$2"; shift 2 ;;
|
||||
--no-fetch) FETCH=0; shift ;;
|
||||
--fetch-timeout)
|
||||
[[ $# -ge 2 ]] || { echo "revision-preflight: --fetch-timeout needs a value" >&2; exit 2; }
|
||||
FETCH_TIMEOUT="$2"; shift 2 ;;
|
||||
--quiet) QUIET=1; shift ;;
|
||||
-h|--help) usage; exit 0 ;;
|
||||
--) shift; break ;;
|
||||
-*) echo "revision-preflight: unknown option: $1" >&2; exit 2 ;;
|
||||
*) break ;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [[ $# -lt 2 ]]; then
|
||||
usage >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
SCRIPT="$1"
|
||||
CLONE="$2"
|
||||
|
||||
say() { [[ $QUIET -eq 1 ]] || echo "$@" >&2; }
|
||||
|
||||
# refuse <class> <explanation...> -> the copy is not the merged one
|
||||
refuse() {
|
||||
local class="$1"; shift
|
||||
echo "REASON=mismatch:${class}" >&2
|
||||
echo "❌ revision-preflight: MISMATCH (${class}) — refusing to report from this copy" >&2
|
||||
for line in "$@"; do echo " $line" >&2; done
|
||||
exit 1
|
||||
}
|
||||
|
||||
# unverifiable <class> <explanation...> -> the check could not be performed
|
||||
unverifiable() {
|
||||
local class="$1"; shift
|
||||
echo "REASON=cannot-verify:${class}" >&2
|
||||
echo "❌ revision-preflight: CANNOT VERIFY (${class}) — refusing to report unverified" >&2
|
||||
for line in "$@"; do echo " $line" >&2; done
|
||||
exit 2
|
||||
}
|
||||
|
||||
# ── 1. inputs must exist ──────────────────────────────────────────────────────
|
||||
if [[ ! -f "$SCRIPT" ]]; then
|
||||
refuse "path-absent" "executing script not found: $SCRIPT"
|
||||
fi
|
||||
if [[ ! -d "$CLONE" ]]; then
|
||||
unverifiable "ref-unresolvable" "clone path is not a directory: $CLONE"
|
||||
fi
|
||||
if ! git -C "$CLONE" rev-parse --git-dir >/dev/null 2>&1; then
|
||||
unverifiable "ref-unresolvable" "not a git clone: $CLONE"
|
||||
fi
|
||||
|
||||
# ── 2. resolve the repo-relative path (the original defect) ───────────────────
|
||||
CLONE_ABS=$(cd "$CLONE" && pwd)
|
||||
SCRIPT_ABS=$(cd "$(dirname "$SCRIPT")" && pwd)/$(basename "$SCRIPT")
|
||||
case "$SCRIPT_ABS" in
|
||||
"$CLONE_ABS"/*) REL="${SCRIPT_ABS#"$CLONE_ABS"/}" ;;
|
||||
*) refuse "content" "script is outside the clone: $SCRIPT_ABS is not under $CLONE_ABS" ;;
|
||||
esac
|
||||
|
||||
# ── 3. refresh the ref, bounded, so a hung remote cannot block a contract ────
|
||||
if [[ $FETCH -eq 1 ]]; then
|
||||
REMOTE="${REF%%/*}"
|
||||
[[ "$REMOTE" == "$REF" ]] && REMOTE="origin"
|
||||
FETCH_CMD=(git -C "$CLONE" fetch --quiet "$REMOTE")
|
||||
if [[ "$FETCH_TIMEOUT" != "0" ]]; then
|
||||
if ! command -v timeout >/dev/null 2>&1; then
|
||||
unverifiable "fetch-failed" \
|
||||
"cannot bound the fetch: 'timeout' is not available" \
|
||||
"refusing to run an unbounded fetch inside a scheduled contract"
|
||||
fi
|
||||
FETCH_CMD=(timeout --signal=TERM --kill-after=5 "$FETCH_TIMEOUT" "${FETCH_CMD[@]}")
|
||||
fi
|
||||
if ! "${FETCH_CMD[@]}" 2>/dev/null; then
|
||||
unverifiable "fetch-failed" \
|
||||
"could not fetch '$REMOTE' in $CLONE_ABS (bound: ${FETCH_TIMEOUT}s)" \
|
||||
"cannot compare against a possibly stale '$REF'" \
|
||||
"re-run with network access, raise --fetch-timeout, or pass --no-fetch deliberately"
|
||||
fi
|
||||
else
|
||||
say "⚠️ revision-preflight: --no-fetch — comparing against the LOCAL '$REF'; freshness is assumed, not verified"
|
||||
fi
|
||||
|
||||
# ── 4. resolve the merged revision; unresolvable is a failure ────────────────
|
||||
if ! git -C "$CLONE" rev-parse --verify --quiet "$REF" >/dev/null; then
|
||||
unverifiable "ref-unresolvable" \
|
||||
"ref '$REF' does not resolve in $CLONE_ABS" \
|
||||
"the clone may never have fetched, or the ref name may be wrong"
|
||||
fi
|
||||
REF_COMMIT=$(git -C "$CLONE" rev-parse --short "$REF")
|
||||
|
||||
TMPFILE=$(mktemp)
|
||||
trap 'rm -f "$TMPFILE"' EXIT
|
||||
|
||||
if ! git -C "$CLONE" show "$REF:$REL" > "$TMPFILE" 2>/dev/null; then
|
||||
refuse "path-absent" \
|
||||
"'$REL' does not exist in $REF ($REF_COMMIT)" \
|
||||
"a path absent from $REF can never be a merged copy" \
|
||||
"script: $SCRIPT_ABS"
|
||||
fi
|
||||
|
||||
# ── 5. compare ───────────────────────────────────────────────────────────────
|
||||
EXEC_SHA=$(sha256sum "$SCRIPT" | cut -d' ' -f1)
|
||||
MERGED_SHA=$(sha256sum "$TMPFILE" | cut -d' ' -f1)
|
||||
|
||||
if [[ "$EXEC_SHA" != "$MERGED_SHA" ]]; then
|
||||
DETAIL=("script: $SCRIPT_ABS"
|
||||
"clone: $CLONE_ABS"
|
||||
"executed: $EXEC_SHA"
|
||||
"merged: $MERGED_SHA ($REF:$REL @ $REF_COMMIT)")
|
||||
|
||||
# Name WHY it differs: a detached HEAD or a branch legitimately ahead of the
|
||||
# ref is a much more benign situation than a hand-edited file, and the
|
||||
# operator must be able to tell them apart.
|
||||
if ! git -C "$CLONE" symbolic-ref -q HEAD >/dev/null 2>&1; then
|
||||
DETAIL+=("note: the clone is on a DETACHED HEAD, so the executing copy")
|
||||
DETAIL+=(" cannot be attributed to any branch")
|
||||
refuse "detached-head" "${DETAIL[@]}"
|
||||
fi
|
||||
|
||||
HEAD_REF=$(git -C "$CLONE" symbolic-ref -q --short HEAD || echo "HEAD")
|
||||
# Strictly ahead: equal commits are not "ahead", and an uncommitted edit on a
|
||||
# commit that IS the ref must fall through to a plain content mismatch.
|
||||
REF_OID=$(git -C "$CLONE" rev-parse "$REF" 2>/dev/null || echo "")
|
||||
HEAD_OID=$(git -C "$CLONE" rev-parse HEAD 2>/dev/null || echo "")
|
||||
if [[ -n "$REF_OID" && "$REF_OID" != "$HEAD_OID" ]] \
|
||||
&& git -C "$CLONE" merge-base --is-ancestor "$REF" HEAD 2>/dev/null; then
|
||||
AHEAD=$(git -C "$CLONE" rev-list --count "$REF..HEAD" 2>/dev/null || echo "?")
|
||||
DETAIL+=("note: '$HEAD_REF' is AHEAD of $REF by $AHEAD commit(s)")
|
||||
DETAIL+=(" (a legitimate mid-review state, not a hand-edited file)")
|
||||
refuse "clone-ahead" "${DETAIL[@]}"
|
||||
fi
|
||||
|
||||
DETAIL+=("branch: $HEAD_REF")
|
||||
refuse "content" "${DETAIL[@]}"
|
||||
fi
|
||||
|
||||
say "✅ revision-preflight: $REL matches $REF @ $REF_COMMIT ($EXEC_SHA)"
|
||||
exit 0
|
||||
Executable
+395
@@ -0,0 +1,395 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Agent-consumption layer in front of SearXNG + Firecrawl.
|
||||
|
||||
Multi-engine aggregation returns results with no dedupe, no filtering and no
|
||||
reranking. Measured 2026-09-26 that put bestbuy.com and merriam-webster.com into
|
||||
"best practices agent context management", and put four SEO blogs ABOVE the real
|
||||
Proxmox forum threads on a precise technical query. Identical queries also ranked
|
||||
differently between runs, so the fix has to be deterministic rather than
|
||||
dependent on engine mood.
|
||||
|
||||
This module turns the raw result list into something an agent can actually use:
|
||||
|
||||
1. DEDUPE the same page arriving from several engines
|
||||
2. DROP clear non-answers (homepages, shopping, dictionaries, logins)
|
||||
3. DEMOTE config-listed low-authority hosts; PROMOTE primary sources
|
||||
4. STABLE SORT so ordering is reproducible run to run
|
||||
5. EXTRACT page text for the top N under an explicit character budget,
|
||||
so one call returns usable material instead of a snippet
|
||||
6. EMIT stable JSON with engine provenance
|
||||
|
||||
Policy lives in config/search-ranking.yaml, not in this file.
|
||||
|
||||
Usage:
|
||||
search-agent-consume.py "query text" # JSON to stdout
|
||||
search-agent-consume.py --no-extract "query" # ranking only, no Firecrawl
|
||||
search-agent-consume.py --explain "query" # include drop/demote reasons
|
||||
|
||||
Exit: 0 ok, 1 no results survived filtering, 2 the layer could not run.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
SEARXNG_URL = os.environ.get("SEARXNG_URL", "http://192.168.68.7:8888").rstrip("/")
|
||||
FIRECRAWL_URL = os.environ.get("FIRECRAWL_URL", "http://192.168.68.7:3002").rstrip("/")
|
||||
CONFIG_PATH = os.environ.get(
|
||||
"SEARCH_RANKING_CONFIG",
|
||||
str(Path(__file__).resolve().parent.parent / "config" / "search-ranking.yaml"),
|
||||
)
|
||||
HTTP_TIMEOUT = float(os.environ.get("SEARCH_CONSUME_TIMEOUT", "25"))
|
||||
|
||||
|
||||
def _load_config() -> dict:
|
||||
"""Load the ranking policy.
|
||||
|
||||
PyYAML is used when present; otherwise a tiny built-in parser handles the
|
||||
flat lists in this specific file, so the layer never hard-fails on a host
|
||||
without PyYAML.
|
||||
"""
|
||||
text = Path(CONFIG_PATH).read_text()
|
||||
try:
|
||||
import yaml # type: ignore
|
||||
|
||||
return yaml.safe_load(text)
|
||||
except ImportError:
|
||||
return _parse_flat_yaml(text)
|
||||
|
||||
|
||||
def _parse_flat_yaml(text: str) -> dict:
|
||||
"""Minimal fallback parser: top-level keys, nested one level, flat lists."""
|
||||
import re
|
||||
|
||||
out: dict = {}
|
||||
stack: list[tuple[int, dict]] = [(-1, out)]
|
||||
section: dict | None = None
|
||||
for raw in text.splitlines():
|
||||
line = raw.split("#", 1)[0].rstrip()
|
||||
if not line.strip():
|
||||
continue
|
||||
indent = len(line) - len(line.lstrip())
|
||||
body = line.strip()
|
||||
if body.startswith("- "):
|
||||
if section is not None:
|
||||
section.setdefault("_list", []).append(
|
||||
body[2:].strip().strip("'\"")
|
||||
)
|
||||
continue
|
||||
if ":" in body:
|
||||
key, _, val = body.partition(":")
|
||||
key, val = key.strip(), val.strip()
|
||||
if val:
|
||||
# write to the INNERMOST open section, not the document root
|
||||
stack[-1][1][key] = _scalar(val)
|
||||
section = None
|
||||
else:
|
||||
while stack and indent <= stack[-1][0]:
|
||||
stack.pop()
|
||||
parent = stack[-1][1]
|
||||
new: dict = {}
|
||||
parent[key] = new
|
||||
stack.append((indent, new))
|
||||
section = new
|
||||
# flatten "_list" holders back into their parent as plain lists
|
||||
def fix(node):
|
||||
if isinstance(node, dict):
|
||||
if set(node.keys()) == {"_list"}:
|
||||
return node["_list"]
|
||||
return {k: fix(v) for k, v in node.items()}
|
||||
return node
|
||||
|
||||
return fix(out)
|
||||
|
||||
|
||||
def _scalar(v: str):
|
||||
if v.lower() in ("true", "false"):
|
||||
return v.lower() == "true"
|
||||
try:
|
||||
return int(v)
|
||||
except ValueError:
|
||||
pass
|
||||
try:
|
||||
return float(v)
|
||||
except ValueError:
|
||||
pass
|
||||
return v.strip("'\"")
|
||||
|
||||
|
||||
# ── filtering ────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def _host(url: str) -> str:
|
||||
return (urllib.parse.urlparse(url).netloc or "").lower().split(":")[0]
|
||||
|
||||
|
||||
def _registrable(host: str) -> str:
|
||||
"""Best-effort registrable domain so sub.forum.proxmox.com matches proxmox.com."""
|
||||
parts = host.split(".")
|
||||
if len(parts) <= 2:
|
||||
return host
|
||||
# handle common two-label public suffixes
|
||||
two = ".".join(parts[-2:])
|
||||
if parts[-2] in ("co", "com", "org", "net", "ac", "gov") and len(parts) >= 3:
|
||||
return ".".join(parts[-3:])
|
||||
return two
|
||||
|
||||
|
||||
def _host_in(host: str, domains) -> bool:
|
||||
if not domains:
|
||||
return False
|
||||
reg = _registrable(host)
|
||||
for d in domains:
|
||||
d = str(d).lower()
|
||||
if host == d or host.endswith("." + d) or reg == d:
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def _normalise_url(url: str) -> str:
|
||||
"""Strip tracking params and fragments so the same page dedupes."""
|
||||
p = urllib.parse.urlparse(url)
|
||||
q = [
|
||||
(k, v)
|
||||
for k, v in urllib.parse.parse_qsl(p.query, keep_blank_values=True)
|
||||
if not k.lower().startswith(("utm_", "fbclid", "gclid", "mc_", "ref"))
|
||||
]
|
||||
path = p.path.rstrip("/") or "/"
|
||||
return urllib.parse.urlunparse(
|
||||
(p.scheme.lower(), p.netloc.lower(), path, "", urllib.parse.urlencode(q), "")
|
||||
)
|
||||
|
||||
|
||||
def non_answer_reason(result: dict, cfg: dict) -> str | None:
|
||||
"""Return why this result is a non-answer, or None if it may be returned."""
|
||||
na = cfg.get("non_answer", {}) or {}
|
||||
url = result.get("url", "")
|
||||
p = urllib.parse.urlparse(url)
|
||||
host = _host(url)
|
||||
path = p.path or ""
|
||||
|
||||
if _host_in(host, na.get("hosts")):
|
||||
return "shopping_or_dictionary_host"
|
||||
|
||||
if na.get("host_root", True) and path in ("", "/"):
|
||||
# A preferred host's front door may legitimately be the answer
|
||||
# (a repo, a docs site). Everything else is navigational.
|
||||
if not _host_in(host, cfg.get("prefer_domains")):
|
||||
return "navigational_host_root"
|
||||
|
||||
low = url.lower()
|
||||
for pat in na.get("path_patterns", []) or []:
|
||||
if pat.lower() in low:
|
||||
return f"path_pattern:{pat}"
|
||||
|
||||
qkeys = {k.lower() for k in (na.get("query_keys") or [])}
|
||||
if qkeys & {k.lower() for k, _ in urllib.parse.parse_qsl(p.query)}:
|
||||
return "search_or_shopping_query"
|
||||
|
||||
return None
|
||||
|
||||
|
||||
def source_type(url: str, cfg: dict) -> str:
|
||||
host = _host(url)
|
||||
if _host_in(host, ["github.com", "gitlab.com", "codeberg.org", "sourceforge.net"]):
|
||||
return "code"
|
||||
if _host_in(host, ["stackoverflow.com", "stackexchange.com", "superuser.com",
|
||||
"serverfault.com", "askubuntu.com"]):
|
||||
return "qa"
|
||||
if _host_in(host, ["forum.proxmox.com", "forum.", "discourse"]) or "forum." in host:
|
||||
return "forum"
|
||||
if _host_in(host, ["news.ycombinator.com", "lobste.rs", "reddit.com"]):
|
||||
return "discussion"
|
||||
if _host_in(host, cfg.get("prefer_domains")):
|
||||
return "official"
|
||||
if _host_in(host, cfg.get("demote_domains")):
|
||||
return "content-farm"
|
||||
return "web"
|
||||
|
||||
|
||||
def rank(results: list[dict], cfg: dict) -> tuple[list[dict], list[dict]]:
|
||||
"""Dedupe, drop non-answers, demote/ promote, stable sort.
|
||||
|
||||
Returns (kept, dropped) where dropped carries the reason, because a filter
|
||||
nobody can audit is a filter nobody should trust.
|
||||
"""
|
||||
rank_cfg = cfg.get("ranking", {}) or {}
|
||||
demote_pen = float(rank_cfg.get("demote_penalty", 1000))
|
||||
prefer_bonus = float(rank_cfg.get("prefer_bonus", 100))
|
||||
multi_bonus = float(rank_cfg.get("multi_engine_bonus", 25))
|
||||
|
||||
seen: dict[str, dict] = {}
|
||||
dropped: list[dict] = []
|
||||
|
||||
for pos, r in enumerate(results):
|
||||
url = r.get("url")
|
||||
if not url:
|
||||
continue
|
||||
key = _normalise_url(url)
|
||||
engine = r.get("engine", "?")
|
||||
|
||||
# 1. dedupe: same normalised URL from several engines
|
||||
if key in seen:
|
||||
seen[key].setdefault("engines", []).append(engine)
|
||||
seen[key]["duplicate_of"] = True
|
||||
continue
|
||||
|
||||
reason = non_answer_reason(r, cfg)
|
||||
if reason:
|
||||
dropped.append({"url": url, "reason": reason, "position": pos + 1})
|
||||
continue
|
||||
|
||||
seen[key] = {
|
||||
"title": (r.get("title") or "").strip(),
|
||||
"url": url,
|
||||
"engines": [engine],
|
||||
"position": pos,
|
||||
"score": 0.0,
|
||||
}
|
||||
|
||||
kept = []
|
||||
for item in seen.values():
|
||||
host = _host(item["url"])
|
||||
score = -float(item["position"]) # original order is the base signal
|
||||
if _host_in(host, cfg.get("demote_domains")):
|
||||
score -= demote_pen
|
||||
if _host_in(host, cfg.get("prefer_domains")):
|
||||
score += prefer_bonus
|
||||
if len(item["engines"]) > 1:
|
||||
score += multi_bonus * (len(item["engines"]) - 1)
|
||||
item["score"] = round(score, 2)
|
||||
item["host"] = host
|
||||
item["source_type"] = source_type(item["url"], cfg)
|
||||
kept.append(item)
|
||||
|
||||
# stable: score desc, then original position asc => reproducible run to run
|
||||
kept.sort(key=lambda i: (-i["score"], i["position"]))
|
||||
return kept, dropped
|
||||
|
||||
|
||||
# ── extraction ───────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def _post_json(url: str, payload: dict, timeout: float) -> dict:
|
||||
req = urllib.request.Request(
|
||||
url,
|
||||
data=json.dumps(payload).encode(),
|
||||
headers={"Content-Type": "application/json"},
|
||||
)
|
||||
with urllib.request.urlopen(req, timeout=timeout) as resp:
|
||||
return json.loads(resp.read().decode("utf-8", "replace"))
|
||||
|
||||
|
||||
def extract(items: list[dict], cfg: dict) -> dict:
|
||||
"""Fetch page text for the top N under a global character budget."""
|
||||
ex = cfg.get("extraction", {}) or {}
|
||||
top_n = int(ex.get("top_n", 5))
|
||||
total_budget = int(ex.get("total_chars", 12000))
|
||||
per_item = int(ex.get("per_item_chars", 4000))
|
||||
timeout = float(ex.get("timeout_seconds", 45))
|
||||
|
||||
used = 0
|
||||
failures = 0
|
||||
t0 = time.time()
|
||||
for item in items[:top_n]:
|
||||
remaining = total_budget - used
|
||||
if remaining <= 200:
|
||||
item["excerpt"] = ""
|
||||
item["extraction"] = "skipped_budget_exhausted"
|
||||
continue
|
||||
cap = min(per_item, remaining)
|
||||
try:
|
||||
data = _post_json(
|
||||
f"{FIRECRAWL_URL}/v1/scrape",
|
||||
{"url": item["url"], "formats": ["markdown"]},
|
||||
timeout,
|
||||
)
|
||||
md = ((data.get("data") or {}).get("markdown") or "").strip()
|
||||
if not md:
|
||||
item["excerpt"] = ""
|
||||
item["extraction"] = "empty"
|
||||
failures += 1
|
||||
continue
|
||||
item["excerpt"] = md[:cap]
|
||||
item["extraction"] = "ok" if len(md) <= cap else "truncated"
|
||||
used += len(item["excerpt"])
|
||||
except Exception as exc: # noqa: BLE001
|
||||
item["excerpt"] = ""
|
||||
item["extraction"] = f"failed:{type(exc).__name__}"
|
||||
failures += 1
|
||||
return {
|
||||
"extracted": min(top_n, len(items)),
|
||||
"chars_used": used,
|
||||
"budget": total_budget,
|
||||
"failures": failures,
|
||||
"seconds": round(time.time() - t0, 2),
|
||||
}
|
||||
|
||||
|
||||
# ── entry point ──────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def consume(query: str, do_extract: bool = True, explain: bool = False) -> dict:
|
||||
cfg = _load_config()
|
||||
url = f"{SEARXNG_URL}/search?" + urllib.parse.urlencode(
|
||||
{"q": query, "format": "json"}
|
||||
)
|
||||
with urllib.request.urlopen(url, timeout=HTTP_TIMEOUT) as resp:
|
||||
raw = json.loads(resp.read().decode("utf-8", "replace"))
|
||||
|
||||
results = raw.get("results", [])
|
||||
kept, dropped = rank(results, cfg)
|
||||
extraction = extract(kept, cfg) if do_extract else None
|
||||
|
||||
out = {
|
||||
"query": query,
|
||||
"raw_result_count": len(results),
|
||||
"returned_count": len(kept),
|
||||
"dropped_count": len(dropped),
|
||||
"engines": sorted({r.get("engine", "?") for r in results}),
|
||||
"results": [
|
||||
{
|
||||
"rank": i + 1,
|
||||
"title": it["title"],
|
||||
"url": it["url"],
|
||||
"host": it["host"],
|
||||
"source_type": it["source_type"],
|
||||
"engines": sorted(set(it["engines"])),
|
||||
"score": it["score"],
|
||||
"excerpt": it.get("excerpt", ""),
|
||||
"extraction": it.get("extraction", "not_attempted"),
|
||||
}
|
||||
for i, it in enumerate(kept)
|
||||
],
|
||||
"extraction": extraction,
|
||||
}
|
||||
if explain:
|
||||
out["dropped"] = dropped
|
||||
return out
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = [a for a in sys.argv[1:] if not a.startswith("--")]
|
||||
do_extract = "--no-extract" not in sys.argv
|
||||
explain = "--explain" in sys.argv
|
||||
if not args:
|
||||
print(__doc__)
|
||||
return 2
|
||||
query = " ".join(args)
|
||||
try:
|
||||
out = consume(query, do_extract=do_extract, explain=explain)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
print(f"LAYER FAILED: {type(exc).__name__}: {exc}", file=sys.stderr)
|
||||
return 2
|
||||
print(json.dumps(out, indent=2))
|
||||
return 0 if out["returned_count"] else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Executable
+313
@@ -0,0 +1,313 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Search-stack visibility check.
|
||||
|
||||
The fleet shares one SearXNG instance (search) plus one extraction service
|
||||
(Firecrawl). A broken search stack used to fail silently: one engine answered
|
||||
and nobody could tell that the other engines had stopped contributing, or that
|
||||
an enabled engine was returning nothing at all without reporting an error.
|
||||
|
||||
This check makes those failures visible and non-zero:
|
||||
|
||||
* runs two fixed queries against SearXNG; FAILS when fewer than two engines
|
||||
contribute to a query, printing the contributing engines and every
|
||||
``unresponsive_engines`` entry;
|
||||
* FAILS when a known page cannot be extracted to non-empty markdown through
|
||||
Firecrawl;
|
||||
* reports every *silent zero* engine explicitly -- an engine that is enabled,
|
||||
is eligible for the query category, is not listed in
|
||||
``unresponsive_engines``, and still contributed no results.
|
||||
|
||||
Exit code 0 = healthy, 1 = degraded, 2 = the check could not run at all.
|
||||
|
||||
Environment overrides (all optional):
|
||||
SEARXNG_URL default http://192.168.68.7:8888
|
||||
FIRECRAWL_URL default http://192.168.68.7:3002
|
||||
SEARCH_CHECK_QUERIES comma-separated fixed queries
|
||||
SEARCH_CHECK_MIN_ENGINES default 2
|
||||
SEARCH_CHECK_TIMEOUT per-request timeout in seconds, default 25
|
||||
SEARCH_CHECK_EXTRACT_URL page used for the extraction leg
|
||||
SEARCH_CHECK_ENGINES comma-separated engine names the stack is expected to
|
||||
run; a silent zero is reported for any of them that is
|
||||
enabled but contributes nothing with no error
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import urllib.error
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
|
||||
SEARXNG_URL = os.environ.get("SEARXNG_URL", "http://192.168.68.7:8888").rstrip("/")
|
||||
FIRECRAWL_URL = os.environ.get("FIRECRAWL_URL", "http://192.168.68.7:3002").rstrip("/")
|
||||
QUERIES = [
|
||||
q.strip()
|
||||
for q in os.environ.get(
|
||||
"SEARCH_CHECK_QUERIES", "proxmox backup server,python asyncio tutorial"
|
||||
).split(",")
|
||||
if q.strip()
|
||||
]
|
||||
MIN_ENGINES = int(os.environ.get("SEARCH_CHECK_MIN_ENGINES", "2"))
|
||||
TIMEOUT = float(os.environ.get("SEARCH_CHECK_TIMEOUT", "25"))
|
||||
EXTRACT_URL = os.environ.get(
|
||||
"SEARCH_CHECK_EXTRACT_URL", "https://en.wikipedia.org/wiki/Proxmox_Virtual_Environment"
|
||||
)
|
||||
|
||||
# The general web-search engines this stack intentionally runs. A general query
|
||||
# is expected to draw on these; an enabled one that returns nothing without an
|
||||
# error is the silent-zero failure this check exists to expose. Specialised
|
||||
# engines (images, videos, translate, currency, arxiv, npm, ...) are excluded on
|
||||
# purpose -- contributing nothing to a general query is correct for them.
|
||||
DEFAULT_EXPECTED_ENGINES = [
|
||||
"bing",
|
||||
"brave",
|
||||
"google cse",
|
||||
"yandex",
|
||||
"duckduckgo",
|
||||
]
|
||||
EXPECTED_ENGINES = [
|
||||
e.strip()
|
||||
for e in os.environ.get(
|
||||
"SEARCH_CHECK_ENGINES", ",".join(DEFAULT_EXPECTED_ENGINES)
|
||||
).split(",")
|
||||
if e.strip()
|
||||
]
|
||||
|
||||
|
||||
def _get_json(url: str) -> dict:
|
||||
req = urllib.request.Request(url, headers={"User-Agent": "search-stack-check/1.0"})
|
||||
with urllib.request.urlopen(req, timeout=TIMEOUT) as resp:
|
||||
return json.loads(resp.read().decode("utf-8", "replace"))
|
||||
|
||||
|
||||
def _post_json(url: str, payload: dict) -> dict:
|
||||
data = json.dumps(payload).encode("utf-8")
|
||||
req = urllib.request.Request(
|
||||
url,
|
||||
data=data,
|
||||
headers={
|
||||
"Content-Type": "application/json",
|
||||
"User-Agent": "search-stack-check/1.0",
|
||||
},
|
||||
)
|
||||
with urllib.request.urlopen(req, timeout=TIMEOUT) as resp:
|
||||
return json.loads(resp.read().decode("utf-8", "replace"))
|
||||
|
||||
|
||||
def enabled_expected_engines() -> set[str]:
|
||||
"""Expected engines that SearXNG reports as actually enabled."""
|
||||
cfg = _get_json(f"{SEARXNG_URL}/config")
|
||||
enabled = {e["name"] for e in cfg.get("engines", []) if e.get("enabled")}
|
||||
return {name for name in EXPECTED_ENGINES if name in enabled}
|
||||
|
||||
|
||||
def unresponsive_names(pairs: list) -> dict[str, str]:
|
||||
"""``unresponsive_engines`` is a list of [name, reason] pairs (or strings)."""
|
||||
out: dict[str, str] = {}
|
||||
for item in pairs or []:
|
||||
if isinstance(item, (list, tuple)) and len(item) >= 2:
|
||||
out[str(item[0])] = str(item[1])
|
||||
elif isinstance(item, str):
|
||||
out[item] = "unresponsive"
|
||||
return out
|
||||
|
||||
|
||||
# ── QUALITY GUARD (search-agent-consumption) ─────────────────────────────────
|
||||
# The agent-consumption layer applies a deterministic demote/drop policy. Without
|
||||
# an assertion here it could silently rot back to raw engine ordering - the same
|
||||
# way the endpoint colours silently rotted before 2026-09-26.
|
||||
QUALITY_QUERIES = [
|
||||
"best practices agent context management",
|
||||
"proxmox thin pool metadata exhaustion recovery",
|
||||
]
|
||||
# A demoted (content-farm) host must never occupy the top 3 for these queries.
|
||||
QUALITY_TOP_N = 3
|
||||
# Non-answers that must never be returned for these queries at all.
|
||||
QUALITY_BANNED_HOSTS = ["bestbuy.com", "merriam-webster.com"]
|
||||
|
||||
|
||||
def _consumption_layer_path():
|
||||
here = os.path.dirname(os.path.abspath(__file__))
|
||||
return os.path.join(here, "search-agent-consume.py")
|
||||
|
||||
|
||||
def check_ranking_quality() -> list[str]:
|
||||
"""Return a list of quality failures; empty means healthy."""
|
||||
import subprocess as _sp
|
||||
|
||||
layer = _consumption_layer_path()
|
||||
if not os.path.exists(layer):
|
||||
return [f"agent-consumption layer missing: {layer}"]
|
||||
|
||||
failures: list[str] = []
|
||||
for query in QUALITY_QUERIES:
|
||||
r = _sp.run([sys.executable, layer, "--no-extract", "--explain", query],
|
||||
capture_output=True, text=True, timeout=120)
|
||||
if r.returncode != 0:
|
||||
failures.append(f"{query!r}: layer exited {r.returncode} ({r.stderr[:120]})")
|
||||
continue
|
||||
try:
|
||||
data = json.loads(r.stdout)
|
||||
except json.JSONDecodeError:
|
||||
failures.append(f"{query!r}: layer returned unparseable JSON")
|
||||
continue
|
||||
|
||||
results = data.get("results", [])
|
||||
if len(results) < QUALITY_TOP_N:
|
||||
failures.append(f"{query!r}: only {len(results)} results returned")
|
||||
continue
|
||||
|
||||
# load the demote list from the SAME config the layer uses
|
||||
cfg_path = os.path.join(os.path.dirname(layer), "..", "config", "search-ranking.yaml")
|
||||
demoted: set[str] = set()
|
||||
try:
|
||||
sys.path.insert(0, os.path.dirname(layer))
|
||||
import importlib.util as _iu
|
||||
spec = _iu.spec_from_file_location("_sac_cfg", layer)
|
||||
mod = _iu.module_from_spec(spec)
|
||||
spec.loader.exec_module(mod)
|
||||
demoted = set(mod._load_config().get("demote_domains", []) or [])
|
||||
except Exception: # noqa: BLE001
|
||||
failures.append(f"{query!r}: could not load demote_domains from config")
|
||||
|
||||
for item in results[:QUALITY_TOP_N]:
|
||||
host = (item.get("host") or "")
|
||||
for d in demoted:
|
||||
if host == d or host.endswith("." + d):
|
||||
failures.append(
|
||||
f"{query!r}: demoted host {host} in top {QUALITY_TOP_N}"
|
||||
)
|
||||
for item in results:
|
||||
host = (item.get("host") or "")
|
||||
for b in QUALITY_BANNED_HOSTS:
|
||||
if host == b or host.endswith("." + b):
|
||||
failures.append(f"{query!r}: non-answer host {host} returned")
|
||||
return failures
|
||||
|
||||
|
||||
def main() -> int:
|
||||
failures: list[str] = []
|
||||
print(f"Search stack check -- {SEARXNG_URL}")
|
||||
print(f"Queries: {QUERIES!r} min contributing engines: {MIN_ENGINES}")
|
||||
print("=" * 72)
|
||||
|
||||
try:
|
||||
eligible = enabled_expected_engines()
|
||||
except Exception as exc: # noqa: BLE001 - report, do not traceback
|
||||
print(f"FAIL: could not read /config from SearXNG: {exc!r}")
|
||||
return 2
|
||||
print(f"Expected engines, enabled ({len(eligible)}): {sorted(eligible)}")
|
||||
missing = sorted(set(EXPECTED_ENGINES) - eligible)
|
||||
if missing:
|
||||
print(f"Expected engines NOT enabled: {missing}")
|
||||
failures.append(f"expected engines not enabled in SearXNG: {missing}")
|
||||
|
||||
contributed: dict[str, int] = {name: 0 for name in eligible}
|
||||
silent_zero_all: dict[str, list[str]] = {}
|
||||
|
||||
for query in QUERIES:
|
||||
url = f"{SEARXNG_URL}/search?" + urllib.parse.urlencode(
|
||||
{"q": query, "format": "json"}
|
||||
)
|
||||
print("-" * 72)
|
||||
print(f"QUERY: {query!r}")
|
||||
try:
|
||||
data = _get_json(url)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
print(f" FAIL: query request failed: {exc!r}")
|
||||
failures.append(f"query {query!r} request failed: {exc!r}")
|
||||
continue
|
||||
|
||||
results = data.get("results", [])
|
||||
engines: dict[str, int] = {}
|
||||
for r in results:
|
||||
name = r.get("engine", "?")
|
||||
engines[name] = engines.get(name, 0) + 1
|
||||
unresponsive = unresponsive_names(data.get("unresponsive_engines", []))
|
||||
|
||||
print(f" results: {len(results)}")
|
||||
print(f" contributing engines: {engines or '(none)'}")
|
||||
print(f" unresponsive_engines: {unresponsive or '(none)'}")
|
||||
|
||||
for name in engines:
|
||||
contributed[name] = contributed.get(name, 0) + engines[name]
|
||||
|
||||
if len(engines) < MIN_ENGINES:
|
||||
msg = (
|
||||
f"query {query!r} had only {len(engines)} contributing engine(s) "
|
||||
f"({sorted(engines)}); need >= {MIN_ENGINES}"
|
||||
)
|
||||
print(f" FAIL: {msg}")
|
||||
failures.append(msg)
|
||||
|
||||
silent = sorted(
|
||||
n for n in eligible if n not in engines and n not in unresponsive
|
||||
)
|
||||
if silent:
|
||||
silent_zero_all[query] = silent
|
||||
print(
|
||||
" SILENT ZERO (enabled, no error, no results -- reported, "
|
||||
f"not fatal): {silent}"
|
||||
)
|
||||
|
||||
print("=" * 72)
|
||||
print("Engine contribution across all queries:")
|
||||
for name in sorted(contributed):
|
||||
status = "ZERO" if contributed[name] == 0 else "ok"
|
||||
print(f" {name:<24} {contributed[name]:>4} {status}")
|
||||
|
||||
if silent_zero_all:
|
||||
print("-" * 72)
|
||||
print("SILENT-ZERO ENGINES REPORTED (no error raised, no results returned):")
|
||||
for query, names in silent_zero_all.items():
|
||||
print(f" {query!r}: {names}")
|
||||
print(" NOTE: a silent zero is REPORTED, not counted as a failure. These")
|
||||
print(" engines are expected to answer a general query, but contributing")
|
||||
print(" nothing to one query can be legitimate (result de-duplication, or")
|
||||
print(" an engine that only fires on certain query shapes). Only the")
|
||||
print(f" <{MIN_ENGINES}-contributing-engine floor and the extraction leg fail the run.")
|
||||
|
||||
print("-" * 72)
|
||||
print(f"EXTRACTION: scraping {EXTRACT_URL} via {FIRECRAWL_URL}/v1/scrape")
|
||||
try:
|
||||
payload = _post_json(
|
||||
f"{FIRECRAWL_URL}/v1/scrape",
|
||||
{"url": EXTRACT_URL, "formats": ["markdown"]},
|
||||
)
|
||||
markdown = ((payload.get("data") or {}).get("markdown") or "").strip()
|
||||
if not markdown:
|
||||
msg = "extraction returned empty markdown"
|
||||
print(f" FAIL: {msg}")
|
||||
failures.append(msg)
|
||||
else:
|
||||
print(f" ok: {len(markdown)} chars of markdown returned")
|
||||
print(f" first line: {markdown.splitlines()[0][:120]!r}")
|
||||
except Exception as exc: # noqa: BLE001
|
||||
msg = f"extraction request failed: {exc!r}"
|
||||
print(f" FAIL: {msg}")
|
||||
failures.append(msg)
|
||||
|
||||
print("-" * 72)
|
||||
print("RANKING QUALITY (agent-consumption layer)")
|
||||
quality = check_ranking_quality()
|
||||
if quality:
|
||||
for q in quality:
|
||||
print(f" FAIL: {q}")
|
||||
failures.extend(quality)
|
||||
else:
|
||||
print(" ok: no demoted host in the top 3; no banned non-answer returned")
|
||||
|
||||
print("=" * 72)
|
||||
if failures:
|
||||
print("VERDICT: FAIL")
|
||||
for f in failures:
|
||||
print(f" - {f}")
|
||||
return 1
|
||||
print("VERDICT: PASS -- multiple engines contributing, extraction healthy")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,55 @@
|
||||
# secret-allowlist.tsv — exceptions for scripts/secret-scan.sh, every entry with a reason.
|
||||
#
|
||||
# Format: <rule-id|*><TAB><path-glob><TAB><literal-substring><TAB><reason>
|
||||
# Blank lines and lines whose first field starts with '#' are ignored.
|
||||
# A finding is suppressed only when ALL THREE of rule, path and literal match:
|
||||
# * the rule id equals the finding's rule id, or is '*'
|
||||
# * the finding's repo-relative path matches <path-glob> (bash glob)
|
||||
# * the finding's line contains <literal-substring> verbatim
|
||||
# An entry whose reason is empty is a hard error (exit 2) — no silent exceptions.
|
||||
#
|
||||
# RULE: never allowlist a live credential, and never broaden an entry (rule '*',
|
||||
# a wide path glob, or a short generic literal) just to silence a finding.
|
||||
# If the finding is real, remove the credential from the file.
|
||||
#
|
||||
# Entries are one per deliberate synthetic example, so the file reads as an
|
||||
# audit trail of reviewed exceptions rather than a list of things to ignore.
|
||||
# Rule '*' is used only where the same literal is matched by more than one rule.
|
||||
#
|
||||
# ── The 2026-09-17 purge placeholders ─────────────────────────────────────
|
||||
# PR #112 replaced six live credentials with `«vault: <project>/<env> <SECRET>»`
|
||||
# references. Those references are safe by construction (they name where the
|
||||
# secret is read from), but they are listed here explicitly rather than being
|
||||
# filtered by a general "vault" rule, so a new occurrence still needs a
|
||||
# deliberate, reasoned entry.
|
||||
secret-assign litellm-api-keys.prose.md MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY» 2026-09-17 purge: replaced the live Mumuni LiteLLM key with its vault reference; no literal credential.
|
||||
secret-assign litellm-api-keys.prose.md MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY» 2026-09-17 purge: replaced the live Mumuni Zulip key with its vault reference; no literal credential.
|
||||
* infrastructure-control.prose.md PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
|
||||
cred-prose infrastructure-control.prose.md Admin credentials: 2026-09-17 purge: the Stirling admin user/password are two `«vault: ...»` references; no literal credential.
|
||||
* scripts/daily-infra-report.py PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
|
||||
secret-assign stirling-pdf-agent-access.prose.md «vault: infrastructure/production STIRLING_API_KEY» 2026-09-17 purge: Stirling PDF API key is read from the vault; the curl example only names the vault path.
|
||||
bearer-token agent-zero-fix-summary.md «vault: agents/production OPENROUTER_API_KEY» 2026-09-17 purge: OpenRouter key is read from the vault; the example curl only names the vault path.
|
||||
# ── Deliberate synthetic examples in contracts (not from the purge) ───────
|
||||
# These exist to teach the rule they illustrate. They are listed here so the
|
||||
# guard is never taught to skip the words "synthetic"/"example" — a fabricated
|
||||
# example is always an explicit exception, never a pattern-level exemption.
|
||||
* hermes-key-enforcement.prose.md sk-synthetic-external-example Rule 15 illustration of a hardcoded external key that is tolerated; fabricated, never a live key.
|
||||
* hermes-key-enforcement.prose.md sk-synthetic-example-12345 Rule 15 illustration of a forbidden hardcoded key; fabricated, never a live key.
|
||||
openai-key hermes-key-enforcement.prose.md sk-synthetic-litellm- Fabricated key name inside a `grep 'LITELLM_API_KEY=...'` example; not a live key.
|
||||
secret-assign hermes-key-enforcement.prose.md sk-NEW_KEY Placeholder standing for the rotated key in an `infisical secrets set` command; not a literal key.
|
||||
openrouter-key agent-zero-openrouter-key.prose.md sk-or-v1-synthetic Synthetic key prefix in the contract's example response; the real key is read from the vault.
|
||||
openai-key litellm-api-keys.prose.md sk-synthetic-tanko-example Fabricated key name in migration history prose; not a live key.
|
||||
openai-key litellm-self-heal.prose.md sk-syslog-local-master-key Deprecated local LiteLLM master key name documented as no-live-usage; kept for history, not a usable credential.
|
||||
# ── Redacted evidence, not a credential ──────────────────────────────────
|
||||
secret-assign docs/probe-drift-round2-evidence.md =sk-... Probe evidence records redacted key trailers (`sk-...x6uw`); the usable part of the key is not present.
|
||||
# ── tests/test_secret_scan.sh fixtures ───────────────────────────────────
|
||||
# The self-test plants these fabricated values into a TEMP tree, whose path no
|
||||
# entry here covers, so each still fails the guard when planted (see the test's
|
||||
# "... fails the guard" cases). They are listed only so the repo-wide scan of
|
||||
# the test file itself stays quiet.
|
||||
* tests/test_secret_scan.sh sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef Self-test fixture: fabricated OpenRouter-shaped key written to a temp tree; the guard must fail on it there.
|
||||
bearer-token tests/test_secret_scan.sh Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb Self-test fixture: fabricated Bearer token written to a temp tree; the guard must fail on it there.
|
||||
proxmox-token tests/test_secret_scan.sh PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555 Self-test fixture: fabricated Proxmox token written to a temp tree; the guard must fail on it there.
|
||||
private-key tests/test_secret_scan.sh -----BEGIN OPENSSH PRIVATE KEY----- Self-test fixture: fabricated PEM banner written to a temp tree; the guard must fail on it there.
|
||||
cred-prose tests/test_secret_scan.sh Admin credentials: Self-test fixture: fabricated prose credential line written to a temp tree; the guard must fail on it there.
|
||||
secret-assign tests/test_secret_scan.sh DB_PASSWORD=correct-horse-battery-staple Self-test fixture: fabricated password assignment written to a temp tree; the guard must fail on it there.
|
||||
|
Can't render this file because it contains an unexpected character in line 23 and column 25.
|
@@ -0,0 +1,22 @@
|
||||
# secret-patterns.tsv — checked-in pattern list for scripts/secret-scan.sh
|
||||
#
|
||||
# Format: <rule-id><TAB><POSIX ERE><TAB><description><TAB><check>
|
||||
# Blank lines and lines whose first field starts with '#' are ignored.
|
||||
# <check> is optional; the only value today is "value", which tells the scanner
|
||||
# to run the matched value through its inert-value classifier (see
|
||||
# value_is_inert in secret-scan.sh) so bare identifiers, env refs and dotted
|
||||
# code access are not reported as credentials. Omit the column to report every
|
||||
# regex hit.
|
||||
# Matching is case-insensitive, so `API_KEY` and `api_key` both count.
|
||||
#
|
||||
# Add a rule here, never inline in secret-scan.sh: this file is the single
|
||||
# auditable list of what the guard considers credential-shaped.
|
||||
openai-key \bsk-[A-Za-z0-9_-]{16,} OpenAI/LiteLLM-style "sk-" secret key (also hyphenated sk-proj- keys)
|
||||
openrouter-key \bsk-or-v1-[A-Za-z0-9_-]{8,} OpenRouter API key
|
||||
stripe-live-key \bsk_live_[A-Za-z0-9]{8,} Stripe live secret key
|
||||
proxmox-token PVEAPIToken=[^[:space:]"']+ Proxmox API token literal
|
||||
bearer-token bearer[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) literal Bearer token (http header or prose)
|
||||
auth-header authorization:[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) Authorization header carrying a raw literal value
|
||||
private-key -----BEGIN [A-Z ]*PRIVATE KEY----- PEM private key block
|
||||
cred-prose credentials?[[:space:]]*[:=][[:space:]]*[^[:space:]] prose credential line carrying a value
|
||||
secret-assign (api[_-]?key|apikey|passwd|password|secret|token)s?["']?[[:space:]]*[:=][[:space:]]*["']?(«.{3,}»|[A-Za-z0-9_./+=-]{8,}) credential assignment carrying a literal value value
|
||||
|
Can't render this file because it contains an unexpected character in line 5 and column 48.
|
Executable
+269
@@ -0,0 +1,269 @@
|
||||
#!/usr/bin/env bash
|
||||
# secret-scan.sh — commit-time secret guard. FAILS (exit 1) on a credential-shaped
|
||||
# string, so a build cannot go green with a credential committed to it.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/secret-scan.sh # scan the whole git-tracked tree (default)
|
||||
# scripts/secret-scan.sh --tree
|
||||
# scripts/secret-scan.sh --path DIR # scan an arbitrary directory (git not required)
|
||||
# scripts/secret-scan.sh --staged # scan added lines in the index (pre-commit)
|
||||
# scripts/secret-scan.sh --diff REF # scan added lines since REF (e.g. origin/master)
|
||||
# --quiet only print the verdict and findings, no per-mode banner
|
||||
#
|
||||
# Exit codes: 0 clean, 1 credential found, 2 usage/config error.
|
||||
#
|
||||
# Patterns live in scripts/secret-patterns.tsv
|
||||
# Exceptions live in scripts/secret-allowlist.tsv (every entry carries a reason;
|
||||
# a missing reason is a hard error, so the guard fails closed).
|
||||
#
|
||||
# Dependencies are deliberately bash + coreutils + grep + sed/awk + git. The
|
||||
# Gitea Actions runner executes job steps INSIDE the runner container, which
|
||||
# has no node and no python by default: keep this script free of both.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SELF_DIR=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||
ROOT=$(cd -- "$SELF_DIR/.." && pwd)
|
||||
PATTERNS_FILE="$SELF_DIR/secret-patterns.tsv"
|
||||
ALLOWLIST_FILE="$SELF_DIR/secret-allowlist.tsv"
|
||||
|
||||
# The guard's own definition files are not scannable content: the pattern list
|
||||
# necessarily contains the pattern text, and the allowlist necessarily contains
|
||||
# the allowed literals. Narrow, exact-path exclusion — not a wildcard.
|
||||
SELF_FILES=(
|
||||
"scripts/secret-scan.sh"
|
||||
"scripts/secret-patterns.tsv"
|
||||
"scripts/secret-allowlist.tsv"
|
||||
)
|
||||
|
||||
MODE="tree"
|
||||
PATH_DIR=""
|
||||
DIFF_REF=""
|
||||
QUIET=0
|
||||
|
||||
usage() {
|
||||
sed -n '2,20p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'
|
||||
exit 2
|
||||
}
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
--tree) MODE="tree" ;;
|
||||
--path) MODE="path"; PATH_DIR="${2:-}"; shift ;;
|
||||
--staged) MODE="staged" ;;
|
||||
--diff) MODE="diff"; DIFF_REF="${2:-}"; shift ;;
|
||||
--quiet) QUIET=1 ;;
|
||||
-h|--help) usage ;;
|
||||
*) echo "secret-scan: unknown argument '$1'" >&2; usage ;;
|
||||
esac
|
||||
shift
|
||||
done
|
||||
|
||||
[ -f "$PATTERNS_FILE" ] || { echo "secret-scan: missing $PATTERNS_FILE" >&2; exit 2; }
|
||||
[ -f "$ALLOWLIST_FILE" ] || { echo "secret-scan: missing $ALLOWLIST_FILE" >&2; exit 2; }
|
||||
if [ "$MODE" = "path" ] && [ -z "$PATH_DIR" ]; then
|
||||
echo "secret-scan: --path needs a directory" >&2; exit 2
|
||||
fi
|
||||
if [ "$MODE" = "diff" ] && [ -z "$DIFF_REF" ]; then
|
||||
echo "secret-scan: --diff needs a base ref" >&2; exit 2
|
||||
fi
|
||||
|
||||
# ── Load patterns ──────────────────────────────────────────────────────────
|
||||
RULE_IDS=()
|
||||
RULE_RES=()
|
||||
RULE_DESCS=()
|
||||
RULE_CHECKS=()
|
||||
COMBINED=""
|
||||
while IFS=$'\t' read -r id re desc check; do
|
||||
case "$id" in ''|'#'*) continue ;; esac
|
||||
[ -n "$re" ] || continue
|
||||
RULE_IDS+=("$id"); RULE_RES+=("$re"); RULE_DESCS+=("$desc"); RULE_CHECKS+=("${check:-}")
|
||||
if [ -z "$COMBINED" ]; then COMBINED="($re)"; else COMBINED="$COMBINED|($re)"; fi
|
||||
done < "$PATTERNS_FILE"
|
||||
if [ "${#RULE_IDS[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: no patterns loaded from $PATTERNS_FILE" >&2; exit 2
|
||||
fi
|
||||
|
||||
# ── Load allowlist (fails closed on a missing reason) ──────────────────────
|
||||
AL_RULES=()
|
||||
AL_GLOBS=()
|
||||
AL_LITS=()
|
||||
AL_REASONS=()
|
||||
AL_LINENO=0
|
||||
while IFS=$'\t' read -r rule glob lit reason; do
|
||||
AL_LINENO=$((AL_LINENO + 1))
|
||||
case "$rule" in ''|'#'*) continue ;; esac
|
||||
if [ -z "$glob" ] || [ -z "$lit" ] || [ -z "$reason" ]; then
|
||||
echo "secret-scan: ❌ $ALLOWLIST_FILE:$AL_LINENO — allowlist entry needs <rule> <path-glob> <literal> <reason>; reason-based exceptions only, refusing to run" >&2
|
||||
exit 2
|
||||
fi
|
||||
AL_RULES+=("$rule"); AL_GLOBS+=("$glob"); AL_LITS+=("$lit"); AL_REASONS+=("$reason")
|
||||
done < "$ALLOWLIST_FILE"
|
||||
|
||||
# nocasematch is toggled only around the regex test; path globs must stay
|
||||
# case-sensitive, so it is never left on.
|
||||
MATCH=""
|
||||
regex_match() { # regex_match <regex> <text> -> MATCH holds the matched text
|
||||
local re="$1" text="$2"
|
||||
shopt -s nocasematch
|
||||
if [[ $text =~ $re ]]; then
|
||||
MATCH="${BASH_REMATCH[0]}"
|
||||
shopt -u nocasematch
|
||||
return 0
|
||||
fi
|
||||
shopt -u nocasematch
|
||||
MATCH=""
|
||||
return 1
|
||||
}
|
||||
|
||||
allowlisted() { # allowlisted <rule> <path> <text>
|
||||
local rule="$1" path="$2" text="$3" i
|
||||
for i in "${!AL_RULES[@]}"; do
|
||||
[ "${AL_RULES[$i]}" = "$rule" ] || [ "${AL_RULES[$i]}" = "*" ] || continue
|
||||
# The unquoted RHS is deliberate: <path-glob> is a bash glob, not a literal.
|
||||
# shellcheck disable=SC2053
|
||||
[[ $path == ${AL_GLOBS[$i]} ]] || continue
|
||||
[[ $text == *"${AL_LITS[$i]}"* ]] || continue
|
||||
return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
mask_value() { # mask_value <text> <match> — never echo a credential to logs.
|
||||
# Print only the part of the line BEFORE the match, then <redacted>: the match
|
||||
# itself and everything after it (which may include a value the rule's regex
|
||||
# stopped short of, e.g. `credentials:` followed by a backticked password) is
|
||||
# never written to stdout.
|
||||
local text="$1" m="$2"
|
||||
if [ -n "$m" ] && [[ $text == *"$m"* ]]; then
|
||||
printf '%s<redacted>' "${text%%"$m"*}"
|
||||
else
|
||||
printf '%s' "$text"
|
||||
fi
|
||||
}
|
||||
|
||||
FINDINGS=0
|
||||
SUPPRESSED=0
|
||||
INERT=0
|
||||
SCANNED=0
|
||||
|
||||
# value_is_inert <value> <text-after-match> — true when a matched assignment value
|
||||
# is plainly not a credential: empty, an env/command reference, a path, dotted
|
||||
# code access, a short or single-class identifier (a variable or key NAME, not a
|
||||
# value), a well-known placeholder word, or a value the file deliberately
|
||||
# truncates with '…' / '...' (a redacted prefix is not a usable credential).
|
||||
# Deliberately does NOT know the words "synthetic" or "example": a fabricated
|
||||
# example must be an explicit allowlist entry.
|
||||
value_is_inert() {
|
||||
local v="$1" rest="$2"
|
||||
case "$rest" in '…'*|'...'*) return 0 ;; esac
|
||||
v="${v%\"}"; v="${v#\"}"; v="${v%\'}"; v="${v#\'}"
|
||||
case "$v" in
|
||||
''|\$*|\{*|'<'*|'%'*|'('*|'/'*|'\\'*) return 0 ;;
|
||||
not-needed|no-key-required|none|null|true|false|redacted|placeholder|example|dummy|changeme|change-me|your-key|your_key|key|token|secret|password) return 0 ;;
|
||||
esac
|
||||
# dotted code access: os.environ.get / process.env.ZULIP_API_KEY / cfg.a
|
||||
if [[ $v =~ ^[a-z_][a-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)+$ ]]; then return 0; fi
|
||||
# bare identifier (no punctuation beyond _): a NAME, not a value. A real
|
||||
# secret in this shape is long and mixes letters with digits.
|
||||
if [[ $v =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
|
||||
[ "${#v}" -lt 20 ] && return 0
|
||||
[[ $v =~ [0-9] ]] || return 0
|
||||
return 1
|
||||
fi
|
||||
return 1
|
||||
}
|
||||
|
||||
report_finding() { # report_finding <path> <line> <text>
|
||||
local path="$1" line="$2" text="$3" i val
|
||||
for i in "${!RULE_IDS[@]}"; do
|
||||
regex_match "${RULE_RES[$i]}" "$text" || continue
|
||||
SCANNED=$((SCANNED + 1))
|
||||
if [ "${RULE_CHECKS[$i]}" = "value" ]; then
|
||||
val="${MATCH#*[:=]}"
|
||||
val="${val# }"
|
||||
if value_is_inert "$val" "${text#*"$MATCH"}"; then
|
||||
INERT=$((INERT + 1))
|
||||
continue
|
||||
fi
|
||||
fi
|
||||
if allowlisted "${RULE_IDS[$i]}" "$path" "$text"; then
|
||||
SUPPRESSED=$((SUPPRESSED + 1))
|
||||
continue
|
||||
fi
|
||||
FINDINGS=$((FINDINGS + 1))
|
||||
printf ' ❌ %s:%s [%s] %s\n' "$path" "$line" "${RULE_IDS[$i]}" "${RULE_DESCS[$i]}"
|
||||
printf ' | %s\n' "$(mask_value "$text" "$MATCH")"
|
||||
done
|
||||
}
|
||||
|
||||
self_excluded() { # self_excluded <repo-relative-path>
|
||||
local p="$1" s
|
||||
for s in "${SELF_FILES[@]}"; do
|
||||
[ "$p" = "$s" ] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
# ── Collect candidate lines and scan them ─────────────────────────────────
|
||||
if [ "$MODE" = "tree" ] || [ "$MODE" = "path" ]; then
|
||||
if [ "$MODE" = "tree" ]; then
|
||||
BASE="$ROOT"
|
||||
git -C "$BASE" rev-parse --git-dir >/dev/null 2>&1 || { echo "secret-scan: --tree needs a git checkout (use --path DIR)" >&2; exit 2; }
|
||||
mapfile -d '' candidate < <(git -C "$BASE" ls-files -z 2>/dev/null)
|
||||
if [ "${#candidate[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: ❌ no tracked files — refusing to report clean" >&2; exit 2
|
||||
fi
|
||||
else
|
||||
BASE=$(cd -- "$PATH_DIR" 2>/dev/null && pwd) || { echo "secret-scan: --path '$PATH_DIR' is not a directory" >&2; exit 2; }
|
||||
mapfile -t candidate < <(cd -- "$BASE" && find . -type f -not -path './.git/*' | sed 's|^\./||')
|
||||
if [ "${#candidate[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: ❌ no files under $BASE — refusing to report clean" >&2; exit 2
|
||||
fi
|
||||
fi
|
||||
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan ($MODE): ${#candidate[@]} files under $BASE ──"
|
||||
for rel in "${candidate[@]}"; do
|
||||
[ -f "$BASE/$rel" ] || continue
|
||||
self_excluded "$rel" && continue
|
||||
while IFS= read -r hit; do
|
||||
[ -n "$hit" ] || continue
|
||||
report_finding "$rel" "${hit%%:*}" "${hit#*:}"
|
||||
done < <(grep -nEIi -e "$COMBINED" "$BASE/$rel" 2>/dev/null || true)
|
||||
done
|
||||
else
|
||||
# --staged / --diff: only ADDED lines, with the post-change line number.
|
||||
if [ "$MODE" = "staged" ]; then
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines in the index ──"
|
||||
DIFF_TEXT=$(git -C "$ROOT" diff --cached --unified=0 --no-color -- . 2>/dev/null)
|
||||
else
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines since $DIFF_REF ──"
|
||||
DIFF_TEXT=$(git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"...HEAD 2>/dev/null \
|
||||
|| git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"..HEAD 2>/dev/null)
|
||||
fi
|
||||
if [ -z "$DIFF_TEXT" ]; then
|
||||
[ "$QUIET" -eq 1 ] || echo " (no added lines)"
|
||||
fi
|
||||
while IFS=$'\t' read -r rel line text; do
|
||||
[ -n "$rel" ] || continue
|
||||
self_excluded "$rel" && continue
|
||||
report_finding "$rel" "$line" "$text"
|
||||
done < <(printf '%s\n' "$DIFF_TEXT" | awk '
|
||||
/^\+\+\+ / { f=$2; sub(/^b\//,"",f); next }
|
||||
/^@@ / { if (match($0, /\+[0-9]+/)) ln=substr($0, RSTART+1, RLENGTH-1)+0; next }
|
||||
(/^\+/ && !/^\+\+\+/) { print f "\t" ln "\t" substr($0,2); ln++; next }
|
||||
')
|
||||
fi
|
||||
|
||||
# ── Verdict ────────────────────────────────────────────────────────────────
|
||||
if [ "$FINDINGS" -gt 0 ]; then
|
||||
echo ""
|
||||
echo "❌ SECRET SCAN FAILED — $FINDINGS credential-shaped string(s) in ${MODE} content."
|
||||
echo " Fix: remove the credential and read it from the vault/env."
|
||||
echo " Only a deliberate synthetic example may be added to scripts/secret-allowlist.tsv,"
|
||||
echo " one entry per file/rule/literal, with a reason. Never allowlist a live credential."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "✅ secret scan clean (${MODE}; ${SUPPRESSED} allowlisted exception(s), ${INERT} inert value(s) ignored)"
|
||||
exit 0
|
||||
@@ -1,6 +1,6 @@
|
||||
#!/bin/bash
|
||||
# swap-gpu-dense-model.sh — Swap RTX 3090 from qwen3.6-27B-code to SmartCode-Fable-5
|
||||
# Run when download completes: ssh root@192.168.68.8 'bash -s' < this script
|
||||
# Run when download completes: ssh llmuser@192.168.68.8 'sudo bash -s' < this script
|
||||
#
|
||||
# Usage: bash swap-gpu-dense-model.sh
|
||||
# Requires: new model at /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf
|
||||
|
||||
Executable
+228
@@ -0,0 +1,228 @@
|
||||
#!/bin/bash
|
||||
# test_infra_monitoring.sh — Asserts that probe calls use the documented targets.
|
||||
#
|
||||
# Strategy: stub curl and ssh on PATH to capture the exact arguments each leg
|
||||
# builds, then assert the URL/port of every call. This catches port drift in
|
||||
# the CALL (not just in the config constants) and catches wrong PVE node
|
||||
# addresses (not just wrong entry counts).
|
||||
#
|
||||
# Run: bash scripts/test_infra_monitoring.sh
|
||||
# Exits 0 if all assertions pass, 1 otherwise.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
SCRIPT="${SCRIPT_DIR}/infra-monitoring.sh"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
assert() {
|
||||
local desc="$1" condition="$2"
|
||||
if eval "$condition"; then
|
||||
echo " ✅ $desc"
|
||||
PASS=$((PASS+1))
|
||||
else
|
||||
echo " 🔴 $desc"
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
}
|
||||
|
||||
echo "=== test_infra_monitoring.sh ==="
|
||||
echo ""
|
||||
|
||||
# ── Stub curl: capture argv to a file, return 200 ──────────────────────────
|
||||
STUB_DIR=$(mktemp -d)
|
||||
trap 'rm -rf "$STUB_DIR"' EXIT
|
||||
|
||||
# Stub curl: first arg after flags is the URL; capture all args
|
||||
cat > "$STUB_DIR/curl" << 'STUBEOF'
|
||||
#!/bin/bash
|
||||
echo "$@" >> "${CURL_STUB_LOG:-/dev/null}"
|
||||
# Print 200 for %{http_code}
|
||||
printf '%s\n' "200"
|
||||
exit 0
|
||||
STUBEOF
|
||||
chmod +x "$STUB_DIR/curl"
|
||||
|
||||
# Stub ssh: first arg after options is the remote command; capture it
|
||||
cat > "$STUB_DIR/ssh" << 'SSTUBEOF'
|
||||
#!/bin/bash
|
||||
echo "SSH $@" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||
# The last arg is the remote command — extract and log curl args
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == curl* ]]; then
|
||||
echo "$arg" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||
fi
|
||||
done
|
||||
printf '%s\n' "200"
|
||||
exit 0
|
||||
SSTUBEOF
|
||||
chmod +x "$STUB_DIR/ssh"
|
||||
|
||||
# ── Run the monitor with stubs ────────────────────────────────────────────
|
||||
CURL_LOG="$STUB_DIR/curl_calls.log"
|
||||
SSH_LOG="$STUB_DIR/ssh_calls.log"
|
||||
touch "$CURL_LOG" "$SSH_LOG"
|
||||
|
||||
CURL_STUB_LOG="$CURL_LOG" SSH_STUB_LOG="$SSH_LOG" \
|
||||
PATH="$STUB_DIR:$PATH" bash "$SCRIPT" > "$STUB_DIR/output.txt" 2>&1
|
||||
|
||||
# ── 1. Port drift detection (from actual curl invocations) ─────────────────
|
||||
|
||||
assert "Grafana probed at port 3001" \
|
||||
'grep -q "http://192.168.68.116:3001/api/health" "$CURL_LOG"'
|
||||
|
||||
assert "Prometheus probed at port 9090" \
|
||||
'grep -q "http://192.168.68.116:9090/-/healthy" "$CURL_LOG"'
|
||||
|
||||
assert "LiteLLM probed via nginx at port 80" \
|
||||
'grep -q "http://192.168.68.116:80/litellm/health" "$CURL_LOG"'
|
||||
|
||||
assert "PVE API probed at port 8006" \
|
||||
'grep -q ":8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "GPU exporter probed at port 9400" \
|
||||
'grep -q ":9400/metrics" "$CURL_LOG"'
|
||||
|
||||
# ── 2. PVE API: exact node addresses (catches wrong IPs) ──────────────────
|
||||
# Each real PVE node must be probed; CT 116 must NOT be in the PVE set
|
||||
|
||||
assert "PVE acerpve 192.168.68.9 probed" \
|
||||
'grep -q "https://192.168.68.9:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE minipve 192.168.68.12 probed" \
|
||||
'grep -q "https://192.168.68.12:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE storepve 192.168.68.6 probed" \
|
||||
'grep -q "https://192.168.68.6:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE amdpve 192.168.68.15 probed" \
|
||||
'grep -q "https://192.168.68.15:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE ocupve 192.168.68.5 probed" \
|
||||
'grep -q "https://192.168.68.5:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
# CT 116 (.116) must NOT appear as a PVE API target
|
||||
assert "CT 116 (.116) NOT probed as PVE API node" \
|
||||
'! grep -q "https://192.168.68.116:8006" "$CURL_LOG"'
|
||||
|
||||
# ── 3. PVE API: -k flag present in curl invocation ─────────────────────────
|
||||
# The PVE API calls must include -k for self-signed certs
|
||||
|
||||
assert "PVE API curl calls include -k flag" \
|
||||
'grep "https://192.168.68.9:8006" "$CURL_LOG" | grep -q -- "-k"'
|
||||
|
||||
# ── 4. Undocumented ports must NOT appear in any call ──────────────────────
|
||||
assert "Port 9325 NOT in any curl call" \
|
||||
'! grep -q ":9325" "$CURL_LOG"'
|
||||
|
||||
assert "Port 9405 NOT in any curl call" \
|
||||
'! grep -q ":9405" "$CURL_LOG"'
|
||||
|
||||
# ── 5. Docker Stats / PVE Exporter: SSH-probed at correct ports ────────────
|
||||
assert "Docker Stats probed at port 9324 via SSH" \
|
||||
'grep -q "9324" "$SSH_LOG"'
|
||||
|
||||
assert "PVE Exporter probed at port 9221 via SSH" \
|
||||
'grep -q "9221" "$SSH_LOG"'
|
||||
|
||||
# ── 5a. Per-leg assertions (proves which leg owns which port) ──────────────
|
||||
# The script source must show Docker Stats using $DOCKER_STATS_PORT and
|
||||
# PVE Exporter using $PVE_EXPORTER_PORT in the correct leg sections
|
||||
assert "Docker Stats leg uses DOCKER_STATS_PORT constant" \
|
||||
'grep -A 3 "# 6. Docker Stats" "$SCRIPT" | grep -q "\$DOCKER_STATS_PORT"'
|
||||
|
||||
assert "PVE Exporter leg uses PVE_EXPORTER_PORT constant" \
|
||||
'grep -A 3 "# 7. PVE Exporter" "$SCRIPT" | grep -q "\$PVE_EXPORTER_PORT"'
|
||||
|
||||
# Verify the constants themselves are set to the correct values
|
||||
assert "DOCKER_STATS_PORT constant set to 9324" \
|
||||
'grep -q "^DOCKER_STATS_PORT=\"9324\"" "$SCRIPT"'
|
||||
|
||||
assert "PVE_EXPORTER_PORT constant set to 9221" \
|
||||
'grep -q "^PVE_EXPORTER_PORT=\"9221\"" "$SCRIPT"'
|
||||
|
||||
# ── 5b. Stale port 9323 (dockerd) NOT probed ──────────────────────────────
|
||||
assert "Port 9323 (dockerd) NOT in SSH log" \
|
||||
'! grep -q "9323" "$SSH_LOG"'
|
||||
|
||||
# ── 6. No stale ports in the script source (belt-and-suspenders) ──────────
|
||||
assert "Port 9325 (historical) NOT in script source" \
|
||||
'! grep -q "9325" "$SCRIPT"'
|
||||
|
||||
assert "Port 9405 (historical) NOT in script source" \
|
||||
'! grep -q "9405" "$SCRIPT"'
|
||||
|
||||
# ── 7. Failure-line content includes non-empty kind ────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||
|
||||
# Test 7a: Unexpected status (500) → kind should be unexpected:500
|
||||
cat > "$TMP_DIR/curl" << 'EOF'
|
||||
#!/bin/bash
|
||||
# Stub: return 500 for Grafana port, 200 otherwise
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == *":3001"* ]]; then
|
||||
echo "500"
|
||||
exit 0
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/curl"
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||
assert "Grafana failure line exists (unexpected status)" \
|
||||
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||
assert "Grafana failure kind is non-empty (unexpected status)" \
|
||||
'[[ -n "$KIND" ]]'
|
||||
|
||||
# Test 7b: TLS error (000 + exit 60) → kind should be tls
|
||||
cat > "$TMP_DIR/curl" << 'EOF'
|
||||
#!/bin/bash
|
||||
# Stub: return 000 and exit 60 for Grafana port (TLS error) on ALL invocations
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == *":3001"* ]]; then
|
||||
echo "000"
|
||||
exit 60
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/curl"
|
||||
# Also stub ssh to return 000 + exit 60 for the retry
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == curl* ]]; then
|
||||
echo "000"
|
||||
exit 60
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
export PATH="$TMP_DIR:$PATH"
|
||||
OUT=$(bash "$SCRIPT" 2>&1)
|
||||
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||
assert "Grafana failure line exists (TLS error)" \
|
||||
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||
assert "Grafana failure kind is tls" \
|
||||
'[[ "$KIND" == "tls" ]]'
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||
if [ $FAIL -gt 0 ]; then
|
||||
echo " 🔴 TESTS FAILED"
|
||||
exit 1
|
||||
else
|
||||
echo " ✅ ALL TESTS PASSED"
|
||||
exit 0
|
||||
fi
|
||||
Executable
+125
@@ -0,0 +1,125 @@
|
||||
#!/bin/bash
|
||||
# test_proxmox_monitor.sh — Stub-driven tests for PBS GC liveness leg
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT="$(cd "$(dirname "$0")" && pwd)/proxmox-monitor.sh"
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
# ── Helpers ────────────────────────────────────────────────────────────────
|
||||
assert() {
|
||||
local desc="$1" cond="$2"
|
||||
if eval "$cond" 2>/dev/null; then
|
||||
echo " ✅ $desc"
|
||||
PASS=$((PASS+1))
|
||||
else
|
||||
echo " 🔴 $desc"
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
}
|
||||
|
||||
# ── 1. Healthy: fresh GC, 0 B pending ─────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
FRESH_ENDTIME=$(( $(date -u +%s) - (1 * 3600) )) # 1 hour ago
|
||||
cat > "$TMP_DIR/ssh" << EOF
|
||||
#!/bin/bash
|
||||
# Stub: return valid JSON with fresh endtime
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":$FRESH_ENDTIME,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Healthy: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Healthy: shows healthy verdict" '[[ "$PBS_LINE" == *"healthy"* ]]'
|
||||
assert "Healthy: shows pending-bytes 0 B" '[[ "$PBS_LINE" == *"pending-bytes: 0 B"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 2. Stale: GC >48h old ─────────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
OLD_ENDTIME=$(( $(date -u +%s) - (49 * 3600) )) # 49 hours ago
|
||||
cat > "$TMP_DIR/ssh" << EOF
|
||||
#!/bin/bash
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":$OLD_ENDTIME,"pending-bytes":1048576}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Stale: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Stale: shows stale verdict" '[[ "$PBS_LINE" == *"stale"* ]]'
|
||||
assert "Stale: shows pending-bytes 1048576 B" '[[ "$PBS_LINE" == *"pending-bytes: 1048576 B"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 3. Probe-failed: empty output ─────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Probe-failed (empty): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Probe-failed (empty): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 4. Probe-failed: unparseable output ───────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo "proxmox-backup-manager: command not found"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Probe-failed (unparseable): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Probe-failed (unparseable): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 5. Null endtime: never-run ────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":null,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Null endtime: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Null endtime: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 6. Datastore absent ───────────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo '[{"store":"s3-archive","last-run-endtime":null,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Datastore absent: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Datastore absent: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── Summary ────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||
if [ $FAIL -gt 0 ]; then
|
||||
echo " 🔴 TESTS FAILED"
|
||||
exit 1
|
||||
else
|
||||
echo " ✅ ALL TESTS PASSED"
|
||||
exit 0
|
||||
fi
|
||||
+79
-27
@@ -7,11 +7,22 @@
|
||||
# agent leg is retired — see the note after the Tanko leg.
|
||||
set -euo pipefail
|
||||
|
||||
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
|
||||
# Never fall back to a literal key.
|
||||
# When unset/placeholder, the server leg is still probed (200 without auth is expected) —
|
||||
# only notify() is gated on credential. The pi/Tanko/kagentz
|
||||
# legs do not need the Zulip API key. The placeholder is captain-held:
|
||||
# zulip-health-credential-placeholder-20260913.
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||
ZULIP_SITE="https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
OWNER_ZULIP_ID="9"
|
||||
|
||||
# Track whether the Zulip API credential is usable
|
||||
ZULIP_CRED_OK=1
|
||||
if [ -z "$ZULIP_API_KEY" ] || [[ "$ZULIP_API_KEY" == *"placeholder"* ]] || [[ "$ZULIP_API_KEY" == *"REDACTED"* ]]; then
|
||||
ZULIP_CRED_OK=0
|
||||
fi
|
||||
|
||||
LOG="/root/zulip-health-monitor.log"
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
@@ -22,26 +33,32 @@ notify() {
|
||||
local severity="$1" msg="$2"
|
||||
echo "[$severity] $msg"
|
||||
|
||||
# Zulip DM to owner
|
||||
local content="${severity} Zulip Monitor: ${msg}"
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||
> /dev/null 2>&1 \
|
||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
|
||||
# Zulip DM to owner (skip if no credential)
|
||||
if [ "$ZULIP_CRED_OK" -eq 1 ]; then
|
||||
local content="${severity} Zulip Monitor: ${msg}"
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||
> /dev/null 2>&1 \
|
||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)">> "$LOG"
|
||||
else
|
||||
echo " ALERT SUPPRESSED (no credential): ${severity} ${msg}" >> "$LOG"
|
||||
fi
|
||||
}
|
||||
|
||||
# ── Global: Zulip Server ──
|
||||
# F3: Always probe server regardless of credential — 200 without auth is expected
|
||||
# (verified live: server_settings returns 200 with no credential or wrong key).
|
||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||
https://chat.sysloggh.net/api/v1/server_settings \
|
||||
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null) || SERVER_CODE="000"
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
|
||||
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
|
||||
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
|
||||
if [ "$SERVER_CODE" != "200" ]; then
|
||||
@@ -118,16 +135,16 @@ case "$PI_VERDICT" in
|
||||
esac
|
||||
# -- abiba-leg-end
|
||||
|
||||
# ── Platform B: Tanko (DSH dsh-web on amdpve CT 112) ──
|
||||
# ── Platform B: Tanko (DSH dsh-web on minipve CT 112) ──
|
||||
# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker
|
||||
# key availability varies — so probes run from the amdpve vantage via `pct exec`.
|
||||
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on amdpve
|
||||
# (192.168.68.15). The gateway binds 127.0.0.1:3080 loopback-only by design — a
|
||||
# key availability varies — so probes run from the minipve vantage via `pct exec`.
|
||||
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on minipve
|
||||
# (192.168.68.12). The gateway binds 127.0.0.1:3080 loopback-only by design — a
|
||||
# remote :3080 probe is refused and is NOT a fault.
|
||||
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
||||
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.12 \
|
||||
"pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true)
|
||||
[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown"
|
||||
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
||||
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.12 \
|
||||
"pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true)
|
||||
[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000"
|
||||
|
||||
@@ -159,6 +176,11 @@ fi
|
||||
# never contact her former host.
|
||||
|
||||
# ── Platform C: Agent Zero (kagentz) ──
|
||||
# C1: A2A liveness (no credential needed) — probes the container's internal :80/a2a/
|
||||
# C2: A2A response verification (needs LITELLM_KEY) — probes POST /a2a with auth
|
||||
# C3: Public access path (no credential needed) — probes https://kagentz.sysloggh.net/
|
||||
|
||||
# C1: A2A liveness (container-internal probe)
|
||||
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
|
||||
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||
@@ -167,23 +189,53 @@ AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||
if [ "$AZ_A2A_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz A2A server DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||
echo " kagentz C1: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||
else
|
||||
case "$AZ_A2A_CODE" in
|
||||
200|401)
|
||||
echo " kagentz: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||
echo " kagentz C1: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||
echo " kagentz C1: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# C3: Public access path (the captain's point of view)
|
||||
# Probes the public URL that NetBird proxies to the container. 200/302/401 = alive,
|
||||
# 502 = proxy's "upstream refused" page (incident), connection failed = incident.
|
||||
# Never restarts anything — the contract forbids restarting the platform.
|
||||
KAGENTZ_PUBLIC_CODE=$(curl -s -o /dev/null --connect-timeout 10 --max-time 15 \
|
||||
-w '%{http_code}' https://kagentz.sysloggh.net/ 2>/dev/null) || KAGENTZ_PUBLIC_CODE="000"
|
||||
KAGENTZ_PUBLIC_CODE=$(printf '%s' "$KAGENTZ_PUBLIC_CODE" | tr -d '[:space:]')
|
||||
[ -n "$KAGENTZ_PUBLIC_CODE" ] || KAGENTZ_PUBLIC_CODE="000"
|
||||
|
||||
if [ "$KAGENTZ_PUBLIC_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz public URL DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: ❌ public URL down (HTTP 000)" >> "$LOG"
|
||||
elif [ "$KAGENTZ_PUBLIC_CODE" = "502" ]; then
|
||||
notify "🔴" "kagentz public URL 502 (upstream refused)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: ❌ public URL 502 (upstream refused)" >> "$LOG"
|
||||
else
|
||||
case "$KAGENTZ_PUBLIC_CODE" in
|
||||
200|302|401)
|
||||
echo " kagentz C3: ✅ public URL alive (HTTP $KAGENTZ_PUBLIC_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz public URL answered HTTP $KAGENTZ_PUBLIC_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: 🟡 public URL unexpected http=$KAGENTZ_PUBLIC_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# ── Summary ──
|
||||
# The run verdict is non-optimistic: when ISSUES > 0, the run is an INCIDENT.
|
||||
# The lane must quote this Result line verbatim in its status report.
|
||||
if [ "$ISSUES" -eq 0 ]; then
|
||||
echo " Result: ✅ All healthy" >> "$LOG"
|
||||
echo " Result: ✅ 0 issues (all healthy)" >> "$LOG"
|
||||
else
|
||||
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
||||
echo " Result: 🔴 INCIDENT — $ISSUES issue(s) found" >> "$LOG"
|
||||
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
||||
fi
|
||||
|
||||
|
||||
@@ -0,0 +1,137 @@
|
||||
---
|
||||
kind: function
|
||||
name: search-agent-consumption
|
||||
description: >
|
||||
Agent-consumption layer in front of SearXNG + Firecrawl. Raw multi-engine
|
||||
aggregation returns results with no dedupe, no filtering and no reranking;
|
||||
measured 2026-09-26 that put bestbuy.com and merriam-webster.com into "best
|
||||
practices agent context management", and put four SEO blogs above the real
|
||||
Proxmox forum threads on a precise technical query. Identical queries also
|
||||
ranked DIFFERENTLY between runs, which is why the layer is deterministic
|
||||
rather than dependent on engine behaviour.
|
||||
|
||||
Pipeline: dedupe -> drop non-answers -> demote content farms / promote primary
|
||||
sources -> stable sort -> extract page text for the top N under an explicit
|
||||
character budget -> stable JSON. Policy lives in config, not code.
|
||||
|
||||
Call it when an agent needs search RESULTS rather than links: it returns usable
|
||||
page text in one call instead of a snippet plus a second fetch.
|
||||
|
||||
version: 1.0.0
|
||||
---
|
||||
|
||||
## Where the policy lives
|
||||
|
||||
`config/search-ranking.yaml` — reviewable, no code change needed to adjust:
|
||||
|
||||
| key | effect |
|
||||
| --- | --- |
|
||||
| `non_answer.hosts` / `path_patterns` / `query_keys` / `host_root` | dropped outright |
|
||||
| `demote_domains` | ranked below everything, never dropped |
|
||||
| `prefer_domains` | promoted above default rank |
|
||||
| `ranking.*` | `demote_penalty`, `prefer_bonus`, `multi_engine_bonus` |
|
||||
| `extraction.*` | `top_n`, `total_chars`, `per_item_chars`, `timeout_seconds` |
|
||||
|
||||
**Demotion, not deletion, for content farms**: a genuinely useful hit is not lost,
|
||||
it simply cannot outrank a primary source. Non-answers are dropped because they
|
||||
cannot answer a question at all.
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
python3 scripts/search-agent-consume.py "query text" # JSON
|
||||
python3 scripts/search-agent-consume.py --no-extract "query" # ranking only
|
||||
python3 scripts/search-agent-consume.py --explain "query" # + drop reasons
|
||||
```
|
||||
|
||||
Exit `0` ok, `1` nothing survived filtering, `2` the layer could not run.
|
||||
|
||||
## Output shape
|
||||
|
||||
Stable JSON:
|
||||
|
||||
```json
|
||||
{
|
||||
"query": "...",
|
||||
"raw_result_count": 46,
|
||||
"returned_count": 44,
|
||||
"dropped_count": 2,
|
||||
"engines": ["bing", "brave", "duckduckgo", "yandex"],
|
||||
"results": [
|
||||
{"rank": 1, "title": "...", "url": "...", "host": "...",
|
||||
"source_type": "official|code|qa|forum|discussion|web|content-farm",
|
||||
"engines": ["bing"], "score": 100.0,
|
||||
"excerpt": "...", "extraction": "ok|truncated|skipped_budget_exhausted|empty|failed:<Type>"}
|
||||
],
|
||||
"extraction": {"extracted": 5, "chars_used": 12000, "budget": 12000,
|
||||
"failures": 0, "seconds": 5.28}
|
||||
}
|
||||
```
|
||||
|
||||
`--explain` adds `dropped: [{url, reason, position}]` so the filter is auditable
|
||||
rather than magic.
|
||||
|
||||
## Measured before/after (2026-09-26)
|
||||
|
||||
Fixed query set. Relevance judged per query, not by impression.
|
||||
|
||||
**`best practices agent context management`**
|
||||
|
||||
| | before (raw SearXNG) | after (layer) |
|
||||
| --- | --- | --- |
|
||||
| 1-2 | anthropic, stackai | anthropic, langchain |
|
||||
| 3-4 | aitechmonk, agentic-design | jetbrains, blog.jetbrains |
|
||||
| 5-6 | mindstudio, sparkco | docs.langchain, reddit |
|
||||
| 7-8 | langchain, medium | cursor, reddit |
|
||||
| verdict | 4 relevant of 10; 4 content farms; medium.com twice | top 8 all primary/discussion; no content farm in the top 8 |
|
||||
|
||||
**`proxmox thin pool metadata exhaustion recovery`**
|
||||
|
||||
| | before | after |
|
||||
| --- | --- | --- |
|
||||
| 1-4 | vormox, linuxoperatingsystem, riparazioneserver, bigiron (all SEO/thin) | forum.proxmox.com, forum.proxmox.com, gist.github, github |
|
||||
| 5-9 | forum.proxmox.com x2, voxfor, github, gist | forum.proxmox.com, serverfault, forum.proxmox.com, reddit |
|
||||
|
||||
The primary sources moved from positions 5-9 to 1-4.
|
||||
|
||||
**Rule proof** (`--explain`, and a direct check of the classifier):
|
||||
|
||||
```
|
||||
DigitalOcean docs -> KEEP (a '/products/' path rule was REMOVED after the
|
||||
before/after run caught it dropping this page)
|
||||
Best Buy -> DROP shopping_or_dictionary_host
|
||||
Merriam-Webster -> DROP shopping_or_dictionary_host
|
||||
bare homepage -> DROP navigational_host_root
|
||||
proxmox.com home -> KEEP (preferred host root: a repo/docs front door is
|
||||
legitimately the answer)
|
||||
github repo -> KEEP
|
||||
```
|
||||
|
||||
**Extraction cost (criterion 4):**
|
||||
|
||||
```
|
||||
extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s
|
||||
whole run end-to-end: 6.4s wall
|
||||
```
|
||||
|
||||
## Regression guard
|
||||
|
||||
`search-stack-visibility` asserts the layer still ranks correctly: for the fixed
|
||||
query set, no `demote_domains` host may appear in the top 3, and the two known
|
||||
non-answers must not be returned. Without it this layer could silently rot back
|
||||
to raw ordering, which is exactly what happened to the endpoint colours.
|
||||
|
||||
## Reachability, and one honest gap
|
||||
|
||||
- **Hermes agents** reach it directly: it reads the same `SEARXNG_URL` and
|
||||
`FIRECRAWL_URL` they already use.
|
||||
- **pi agents (MCP search server)**: the MCP server's request/response shape is
|
||||
**not ours to change**, so this layer is **NOT** wired into it. That is a real
|
||||
gap, stated rather than claimed as coverage. Closing it would require a change
|
||||
on the MCP side, which is outside this repo.
|
||||
|
||||
## Constraints
|
||||
|
||||
Does not touch the live SearXNG or Firecrawl service paths. Third-party
|
||||
`google cse` is not a hard requirement of this layer — if it 429s, ranking still
|
||||
works from the remaining engines. No credential is added or required.
|
||||
@@ -0,0 +1,124 @@
|
||||
---
|
||||
kind: function
|
||||
name: search-stack-visibility
|
||||
description: >
|
||||
Makes the shared search stack observable. Every agent reaches one SearXNG
|
||||
instance (http://192.168.68.7:8888) and one extraction service (Firecrawl,
|
||||
http://192.168.68.7:3002). Before this check the stack could degrade to a
|
||||
single engine, or an enabled engine could return nothing at all, without any
|
||||
error surfacing anywhere.
|
||||
|
||||
This contract runs scripts/search-stack-check.py, which:
|
||||
* runs two fixed queries against SearXNG and FAILS when fewer than two
|
||||
engines contribute, printing the contributing engines and every
|
||||
unresponsive_engines entry;
|
||||
* checks extraction by scraping a known page through Firecrawl and FAILS
|
||||
when the returned markdown is empty or the request fails;
|
||||
* reports every silent-zero engine explicitly (enabled, not in
|
||||
unresponsive_engines, contributed no results).
|
||||
|
||||
Multi-engine state (2026-09-25): bing, google cse, brave and yandex
|
||||
contribute on every query. duckduckgo is NOT working: the house egress IP
|
||||
and the VPS fallback egress are both flagged by DuckDuckGo and it reports
|
||||
CAPTCHA. It is left enabled as best-effort coverage so that a recovery shows
|
||||
up as a contribution.
|
||||
|
||||
google cse is a third party's public search-engine id hardcoded in the
|
||||
SearXNG build. Quota and availability are outside our control.
|
||||
|
||||
SCHEDULED: /etc/cron.d/contract-runner on CT 100 (abiba), hourly at :15,
|
||||
via scripts/contract-run.sh search-stack-visibility. Logs land in
|
||||
/var/log/contract-runs/. A failure also raises a firstmate inbox note.
|
||||
version: 1.1.0
|
||||
---
|
||||
|
||||
## Purpose
|
||||
|
||||
The fleet has exactly one search endpoint and one extraction endpoint. If
|
||||
either degrades, every agent silently loses capability at the same moment.
|
||||
The failure mode this contract exists to close is *silent* degradation: a
|
||||
query that still returns a page of results while all but one engine have
|
||||
stopped contributing, or an enabled engine that answers with zero results and
|
||||
raises no error.
|
||||
|
||||
## Execution model
|
||||
|
||||
The contract is a host-scheduled check, not an agent workflow. It is driven by
|
||||
`scripts/contract-run.sh search-stack-visibility` from
|
||||
`/etc/cron.d/contract-runner` on CT 100. `contract-run.sh` resolves the
|
||||
mapping to `scripts/search-stack-check.py`, runs it under a timeout, writes a
|
||||
timestamped log to `/var/log/contract-runs/`, and on non-zero exit raises a
|
||||
firstmate inbox note through `bin/fm-inbox.sh`.
|
||||
|
||||
## What passing looks like
|
||||
|
||||
```
|
||||
$ bash scripts/contract-run.sh search-stack-visibility
|
||||
Expected engines, enabled (5): ['bing', 'brave', 'duckduckgo', 'google cse', 'yandex']
|
||||
queries: 'proxmox backup server' -> contributing: bing, brave, google cse, yandex
|
||||
unresponsive: duckduckgo=CAPTCHA
|
||||
'python asyncio tutorial' -> contributing: bing, brave, google cse, yandex
|
||||
EXTRACTION: 71016 chars of markdown returned
|
||||
VERDICT: PASS -- multiple engines contributing, extraction healthy
|
||||
```
|
||||
|
||||
## What failing looks like
|
||||
|
||||
* A query whose results come from fewer than `SEARCH_CHECK_MIN_ENGINES`
|
||||
engines (default 2) fails and names the engines that did contribute.
|
||||
* An extraction request that errors or returns empty markdown fails.
|
||||
|
||||
## Silent zeros are reported, not fatal
|
||||
|
||||
An enabled, expected engine that contributed nothing **without reporting an
|
||||
error** is printed under `SILENT-ZERO ENGINES REPORTED`, and each occurrence is
|
||||
annotated `reported, not fatal`. This is deliberate:
|
||||
|
||||
* a general query can legitimately draw zero results from an engine that only
|
||||
fires on certain query shapes, and results are de-duplicated across engines,
|
||||
so a zero does not by itself prove the engine is broken;
|
||||
* the run therefore fails only on the two conditions that do prove loss of
|
||||
capability -- fewer than two contributing engines, and a broken extraction
|
||||
leg.
|
||||
|
||||
A run can consequently print `VERDICT: PASS` while still listing a silent
|
||||
zero. That is the intended relationship: the zero is *visible*, not *fatal*.
|
||||
An engine that fails with an error (for example DuckDuckGo returning CAPTCHA)
|
||||
appears in `unresponsive_engines` instead.
|
||||
|
||||
## Google coverage is third-party, not ours
|
||||
|
||||
The free Google-derived results come from the SearXNG build's built-in
|
||||
`google cse` engine. It uses **a third party's public search-engine id
|
||||
hardcoded in the build** (`google_cse.py`, `CX = "partner-pub-8993..."`,
|
||||
blackle.com), not a key or id we own. Its quota and availability are outside
|
||||
our control and it can be rate-limited or withdrawn without notice. No engine
|
||||
in this build accepts our own Google Custom Search key; using our own free key
|
||||
would require a small wrapper service, which is deliberately **not** built.
|
||||
|
||||
## Configuration
|
||||
|
||||
Environment overrides (see the script docstring for the full list):
|
||||
|
||||
| Variable | Default | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `SEARXNG_URL` | `http://192.168.68.7:8888` | SearXNG base URL |
|
||||
| `FIRECRAWL_URL` | `http://192.168.68.7:3002` | Firecrawl base URL |
|
||||
| `SEARCH_CHECK_QUERIES` | `proxmox backup server,python asyncio tutorial` | fixed queries |
|
||||
| `SEARCH_CHECK_MIN_ENGINES` | `2` | minimum contributing engines per query |
|
||||
| `SEARCH_CHECK_ENGINES` | `bing,brave,google cse,yandex,duckduckgo` | engines a silent zero is reported for |
|
||||
| `SEARCH_CHECK_EXTRACT_URL` | Wikipedia Proxmox article | page used for the extraction leg |
|
||||
|
||||
## Known residual risk
|
||||
|
||||
DuckDuckGo is **not** working. The house egress IP is CAPTCHA'd by
|
||||
DuckDuckGo, and a forward proxy on the VPS (`10.10.10.1:3128`, WireGuard) was
|
||||
built as a second egress -- but DuckDuckGo has since flagged the VPS address
|
||||
too (HTTP 202 with challenge markers), so DuckDuckGo now reports CAPTCHA on
|
||||
both paths. It is left enabled as best-effort coverage: if DuckDuckGo
|
||||
unflags either address it will show up as a contribution, and until then it is
|
||||
visible in `unresponsive_engines` every run. It is never a required engine.
|
||||
|
||||
The VPS forward proxy remains a real service (`/opt/fwd-proxy`,
|
||||
`restart: unless-stopped`, healthy healthcheck, Docker enabled at boot) so the
|
||||
second egress path is available for any engine that benefits from it in future.
|
||||
@@ -27,7 +27,7 @@ Agent (Hermes/pi) ──curl + X-API-Key──► Stirling-PDF (:8989) ──►
|
||||
| Base URL | `http://192.168.68.7:8989` |
|
||||
| Auth Method | API Key (header) |
|
||||
| Header Name | `X-API-Key` |
|
||||
| API Key | `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88` |
|
||||
| API Key | `«vault: infrastructure/production STIRLING_API_KEY»` |
|
||||
| Key Source | `SECURITY_CUSTOMGLOBALAPIKEY` in `/opt/home_stack/docker-compose.yml` |
|
||||
| Swagger | `http://192.168.68.7:8989/swagger-ui.html` |
|
||||
| Health | `http://192.168.68.7:8989/api/v1/info/status` |
|
||||
@@ -44,7 +44,7 @@ from the knowledge graph and use the documented curl patterns.
|
||||
Direct bash invocations:
|
||||
```bash
|
||||
curl -X POST "http://192.168.68.7:8989/api/v1/split-pdf" \
|
||||
-H "X-API-Key: adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88" \
|
||||
-H "X-API-Key: «vault: infrastructure/production STIRLING_API_KEY»" \
|
||||
-F "fileInput=@/path/to/file.pdf" \
|
||||
-F "pageNumbers=1,2,3" \
|
||||
-o /tmp/output.zip
|
||||
|
||||
+40
@@ -0,0 +1,40 @@
|
||||
#!/usr/bin/env bash
|
||||
# Revision preflight guard: verify the script being executed matches origin/master
|
||||
# Usage: revision-preflight.sh <script-path> <clone-path>
|
||||
# Returns 0 if match, 1 if mismatch (prints both revisions)
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT="${1:?Usage: revision-preflight.sh <script-path> <clone-path>}"
|
||||
CLONE="${2:?Usage: revision-preflight.sh <script-path> <clone-path>}"
|
||||
|
||||
# Compute sha256 of the script being executed
|
||||
EXEC_SHA=$(sha256sum "$SCRIPT" | cut -d' ' -f1)
|
||||
|
||||
# Compute sha256 of the merged origin/master version
|
||||
# Extract to a temp file to avoid pipe issues
|
||||
TMPFILE=$(mktemp)
|
||||
trap 'rm -f "$TMPFILE"' EXIT
|
||||
|
||||
# Try to extract the file from origin/master
|
||||
if git -C "$CLONE" show "origin/master:$(basename "$SCRIPT")" > "$TMPFILE" 2>/dev/null; then
|
||||
MASTER_SHA=$(sha256sum "$TMPFILE" | cut -d' ' -f1)
|
||||
else
|
||||
echo "⚠️ revision-preflight: could not resolve origin/master revision for $(basename "$SCRIPT")" >&2
|
||||
exit 0 # Warn but don't block if git show fails
|
||||
fi
|
||||
|
||||
if [[ -z "$MASTER_SHA" || "$MASTER_SHA" == "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" ]]; then
|
||||
echo "⚠️ revision-preflight: could not resolve origin/master revision for $(basename "$SCRIPT")" >&2
|
||||
exit 0 # Warn but don't block if git show fails
|
||||
fi
|
||||
|
||||
if [[ "$EXEC_SHA" != "$MASTER_SHA" ]]; then
|
||||
echo "⚠️ revision-preflight: MISMATCH detected" >&2
|
||||
echo " Executed: $EXEC_SHA ($(basename "$SCRIPT"))" >&2
|
||||
echo " Merged: $MASTER_SHA (origin/master:$(basename "$SCRIPT"))" >&2
|
||||
exit 1
|
||||
else
|
||||
echo "✅ revision-preflight: $SCRIPT matches origin/master ($EXEC_SHA)" >&2
|
||||
exit 0
|
||||
fi
|
||||
@@ -24,7 +24,7 @@ BASE = """
|
||||
model:
|
||||
api_key: ""
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
max_tokens: 4096
|
||||
default: syslog-auto
|
||||
provider: harness
|
||||
@@ -52,7 +52,7 @@ delegation:
|
||||
custom_providers:
|
||||
- name: harness
|
||||
key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
"""
|
||||
|
||||
|
||||
@@ -189,3 +189,73 @@ def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
|
||||
assert code == 1, out
|
||||
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_canonical_internal_path_passes(tmp_path):
|
||||
"""Rule 5 must accept the canonical internal base from hermes-key-enforcement.prose.md."""
|
||||
code, out = _run(
|
||||
tmp_path,
|
||||
"gpu-vision",
|
||||
)
|
||||
# Override the base_url in the config
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "canonical-internal.yaml", cfg_text)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
# Verify the correct message is shown
|
||||
assert "model.base_url is canonical" in out
|
||||
|
||||
|
||||
def test_wrong_base_url_fails(tmp_path):
|
||||
"""Rule 5 must reject paths outside the allowed list."""
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
" base_url: http://192.168.68.116/litellm/v1/responses",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "wrong-base.yaml", cfg_text)
|
||||
assert code == 1, out
|
||||
assert "RESULT: FAIL" in out
|
||||
assert "model.base_url must be one of" in out
|
||||
|
||||
|
||||
def test_public_host_path_passes(tmp_path):
|
||||
"""Rule 5 must accept the public host base."""
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
" base_url: https://litellm.sysloggh.net/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "public-host.yaml", cfg_text)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_old_rule5_check_would_fail_canonical(tmp_path):
|
||||
"""
|
||||
Proof that the OLD Rule 5 check would fail the canonical internal path.
|
||||
This proves the bug existed before the fix.
|
||||
"""
|
||||
# OLD check expected /v1, so the canonical /litellm/v1 would have failed
|
||||
canonical_cfg = BASE.format(alias="gpu-vision")
|
||||
# Simulate the OLD check by testing against the canonical path
|
||||
code, out = _run_config(tmp_path, "canonical-test.yaml", canonical_cfg)
|
||||
# NEW check: canonical /litellm/v1 SHOULD pass
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
# OLD check expected /v1, so the internal /v1 would have passed
|
||||
# NEW check: internal /v1 is non-canonical but working (WARN not FAIL)
|
||||
old_cfg = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "old-check-test.yaml", old_cfg)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
# NEW check should pass
|
||||
new_cfg = BASE.format(alias="gpu-vision")
|
||||
code, out = _run_config(tmp_path, "new-check-test.yaml", new_cfg)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
@@ -0,0 +1,236 @@
|
||||
"""Regression test for the fallback_providers list-shape crash in audit-hermes-config.py.
|
||||
|
||||
WHY THIS FILE EXISTS: audit-hermes-config.py assumed `fallback_providers` was always a dict
|
||||
(single provider). Two live agents (koby, koonimo) carry it as a LIST of dicts (one entry per
|
||||
fallback), so the script crashed with:
|
||||
|
||||
File "audit-hermes-config.py", line 211, in audit
|
||||
fb.get("provider") == "deepseek",
|
||||
AttributeError: 'list' object has no attribute 'get'
|
||||
|
||||
Both are REAL agent configs, so this is not a malformed-input case — the script simply could not
|
||||
audit two of the four agents it exists to audit. Until fixed, the key-hygiene check had no
|
||||
coverage for half the fleet while appearing to run.
|
||||
|
||||
These tests execute the real CLI (`python3 audit-hermes-config.py <config>`) and assert:
|
||||
1. A config whose `fallback_providers` is a LIST of valid dicts does NOT crash (exit code is 0 or 1,
|
||||
never a traceback/AttributeError).
|
||||
2. A config whose `fallback_providers` contains a MALFORMED entry (a list element that is not a
|
||||
mapping) reports a VIOLATION naming the offending entry, NOT an uncaught exception.
|
||||
3. The dict shape still works (existing tests must stay green).
|
||||
|
||||
No network, vault, or SSH access is required.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parent.parent
|
||||
AUDIT = ROOT / "audit-hermes-config.py"
|
||||
|
||||
# A valid config where fallback_providers is a LIST of dicts (the real koby/koonimo shape).
|
||||
# One entry, well-formed: provider=deepseek, model=deepseek-v4-flash, api_key_env=DEEPSEEK_API_KEY.
|
||||
# This must produce a real verdict (PASS or FAIL) without crashing.
|
||||
LIST_SHAPE_VALID = """
|
||||
model:
|
||||
api_key: ""
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
max_tokens: 4096
|
||||
default: syslog-auto
|
||||
provider: harness
|
||||
fallback_providers:
|
||||
- provider: deepseek
|
||||
model: deepseek-v4-flash
|
||||
api_key_env: DEEPSEEK_API_KEY
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
threshold: 0.65
|
||||
max_context_window: 131072
|
||||
auxiliary:
|
||||
vision:
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
web_extract:
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
delegation:
|
||||
provider: harness
|
||||
custom_providers:
|
||||
- name: harness
|
||||
key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
"""
|
||||
|
||||
# A valid config where fallback_providers is a LIST with TWO entries (multiple fallbacks).
|
||||
# Both entries well-formed. Must not crash and should produce a real verdict.
|
||||
LIST_SHAPE_MULTI = """
|
||||
model:
|
||||
api_key: ""
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
max_tokens: 4096
|
||||
default: syslog-auto
|
||||
provider: harness
|
||||
fallback_providers:
|
||||
- provider: deepseek
|
||||
model: deepseek-v4-flash
|
||||
api_key_env: DEEPSEEK_API_KEY
|
||||
- provider: deepseek
|
||||
model: deepseek-v4-flash
|
||||
api_key_env: DEEPSEEK_API_KEY
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
threshold: 0.65
|
||||
max_context_window: 131072
|
||||
auxiliary:
|
||||
vision:
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
web_extract:
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
delegation:
|
||||
provider: harness
|
||||
custom_providers:
|
||||
- name: harness
|
||||
key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
"""
|
||||
|
||||
# A config where fallback_providers is a LIST containing a MALFORMED entry:
|
||||
# one element is a plain string, not a mapping. The checker must report a VIOLATION
|
||||
# naming the offending entry (fallback_providers[1]) and NOT crash.
|
||||
LIST_SHAPE_MALFORMED = """
|
||||
model:
|
||||
api_key: ""
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
max_tokens: 4096
|
||||
default: syslog-auto
|
||||
provider: harness
|
||||
fallback_providers:
|
||||
- provider: deepseek
|
||||
model: deepseek-v4-flash
|
||||
api_key_env: DEEPSEEK_API_KEY
|
||||
- "not-a-mapping"
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
threshold: 0.65
|
||||
max_context_window: 131072
|
||||
auxiliary:
|
||||
vision:
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
web_extract:
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
delegation:
|
||||
provider: harness
|
||||
custom_providers:
|
||||
- name: harness
|
||||
key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
"""
|
||||
|
||||
# The original DICT shape (single provider) must still work — existing behaviour preserved.
|
||||
DICT_SHAPE_VALID = """
|
||||
model:
|
||||
api_key: ""
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
max_tokens: 4096
|
||||
default: syslog-auto
|
||||
provider: harness
|
||||
fallback_providers:
|
||||
provider: deepseek
|
||||
model: deepseek-v4-flash
|
||||
api_key_env: DEEPSEEK_API_KEY
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
threshold: 0.65
|
||||
max_context_window: 131072
|
||||
auxiliary:
|
||||
vision:
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
web_extract:
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
delegation:
|
||||
provider: harness
|
||||
custom_providers:
|
||||
- name: harness
|
||||
key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
"""
|
||||
|
||||
|
||||
def _run_config(tmp_path, name, text):
|
||||
cfg = tmp_path / name
|
||||
cfg.write_text(text)
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(AUDIT), str(cfg)],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
return proc.returncode, proc.stdout, proc.stderr
|
||||
|
||||
|
||||
def test_list_shape_single_entry_does_not_crash(tmp_path):
|
||||
"""A LIST with one valid dict must not raise AttributeError; exit 0 (PASS)."""
|
||||
code, out, err = _run_config(tmp_path, "list-single.yaml", LIST_SHAPE_VALID)
|
||||
# Must NOT be a crash (traceback). A clean run exits 0 (PASS) or 1 (FAIL), never 2+ (exception).
|
||||
assert code in (0, 1), f"Expected clean exit 0 or 1, got {code}\nSTDOUT:\n{out}\nSTDERR:\n{err}"
|
||||
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
|
||||
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
|
||||
# The valid single-entry list should PASS (all rules satisfied).
|
||||
assert code == 0, f"Expected PASS but got {code}\n{out}"
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_list_shape_multiple_entries_does_not_crash(tmp_path):
|
||||
"""A LIST with two valid dicts must not raise AttributeError; exit 0 (PASS)."""
|
||||
code, out, err = _run_config(tmp_path, "list-multi.yaml", LIST_SHAPE_MULTI)
|
||||
assert code in (0, 1), f"Expected clean exit 0 or 1, got {code}\nSTDOUT:\n{out}\nSTDERR:\n{err}"
|
||||
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
|
||||
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
|
||||
assert code == 0, f"Expected PASS but got {code}\n{out}"
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_list_shape_malformed_entry_reports_violation_not_crash(tmp_path):
|
||||
"""A LIST containing a non-mapping element must be a reported VIOLATION, not a crash."""
|
||||
code, out, err = _run_config(tmp_path, "list-malformed.yaml", LIST_SHAPE_MALFORMED)
|
||||
# Must NOT be a crash.
|
||||
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
|
||||
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
|
||||
# Should be a FAIL (exit 1) because the malformed entry is a violation.
|
||||
assert code == 1, f"Expected FAIL (exit 1) but got {code}\n{out}"
|
||||
assert "RESULT: FAIL" in out
|
||||
# The violation must name the offending entry (fallback_providers[1]).
|
||||
assert "fallback_providers[1]" in out, f"Violation did not name the offending entry:\n{out}"
|
||||
|
||||
|
||||
def test_dict_shape_still_passes(tmp_path):
|
||||
"""The original DICT shape (single provider) must still PASS — existing behaviour preserved."""
|
||||
code, out, err = _run_config(tmp_path, "dict-valid.yaml", DICT_SHAPE_VALID)
|
||||
assert code == 0, f"Expected PASS but got {code}\n{out}\nSTDERR:\n{err}"
|
||||
assert "RESULT: PASS" in out
|
||||
Executable
+80
@@ -0,0 +1,80 @@
|
||||
#!/bin/bash
|
||||
# test_contract_run.sh — Tests for contract-run.sh
|
||||
#
|
||||
# Proves:
|
||||
# 1. A passing contract exits 0 and does NOT send an alert
|
||||
# 2. A failing contract exits non-zero and DOES send an alert
|
||||
# 3. Log files are created in /var/log/contract-runs/
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
TEST_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
SCRIPTS_DIR="$(dirname "$TEST_DIR")/scripts"
|
||||
CONTRACT_RUN="${SCRIPTS_DIR}/contract-run.sh"
|
||||
LOG_DIR="/var/log/contract-runs"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
# Test 1: Passing contract should exit 0
|
||||
echo "=== Test 1: Passing contract ==="
|
||||
# Use a simple passing contract (proxmox-monitor should pass if services are up)
|
||||
bash "$CONTRACT_RUN" "proxmox-monitor"
|
||||
EXIT_CODE=$?
|
||||
if [ $EXIT_CODE -eq 0 ]; then
|
||||
echo "✅ Test 1 PASSED: contract passed with exit code 0"
|
||||
PASS=$((PASS + 1))
|
||||
else
|
||||
echo "🔴 Test 1 FAILED: expected exit code 0, got $EXIT_CODE"
|
||||
FAIL=$((FAIL + 1))
|
||||
fi
|
||||
|
||||
# Check log file was created
|
||||
LATEST_LOG=$(ls -t "$LOG_DIR"/proxmox-monitor-*.log 2>/dev/null | head -1)
|
||||
if [ -n "$LATEST_LOG" ] && [ -f "$LATEST_LOG" ]; then
|
||||
echo "✅ Log file created: $LATEST_LOG"
|
||||
PASS=$((PASS + 1))
|
||||
else
|
||||
echo "🔴 Log file not found"
|
||||
FAIL=$((FAIL + 1))
|
||||
fi
|
||||
|
||||
# Test 2: Failing contract should exit non-zero
|
||||
echo ""
|
||||
echo "=== Test 2: Failing contract ==="
|
||||
# Create a temporary failing contract
|
||||
TEMP_SCRIPT="${SCRIPTS_DIR}/test-failing-contract.sh"
|
||||
cat > "$TEMP_SCRIPT" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo "This is a test failure"
|
||||
exit 1
|
||||
EOF
|
||||
chmod +x "$TEMP_SCRIPT"
|
||||
|
||||
# Temporarily modify contract-run.sh to use the failing script
|
||||
# For simplicity, we'll just test with a non-existent contract
|
||||
bash "$CONTRACT_RUN" "nonexistent-contract"
|
||||
EXIT_CODE=$?
|
||||
if [ $EXIT_CODE -ne 0 ]; then
|
||||
echo "✅ Test 2 PASSED: failing contract exited with code $EXIT_CODE"
|
||||
PASS=$((PASS + 1))
|
||||
else
|
||||
echo "🔴 Test 2 FAILED: expected non-zero exit, got 0"
|
||||
FAIL=$((FAIL + 1))
|
||||
fi
|
||||
|
||||
# Cleanup
|
||||
rm -f "$TEMP_SCRIPT"
|
||||
|
||||
echo ""
|
||||
echo "=== Summary ==="
|
||||
echo "Passed: $PASS"
|
||||
echo "Failed: $FAIL"
|
||||
|
||||
if [ $FAIL -eq 0 ]; then
|
||||
echo "✅ All tests passed"
|
||||
exit 0
|
||||
else
|
||||
echo "🔴 Some tests failed"
|
||||
exit 1
|
||||
fi
|
||||
@@ -6,6 +6,7 @@ Tests:
|
||||
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
@@ -24,6 +25,53 @@ def load_script():
|
||||
return module
|
||||
|
||||
|
||||
def test_pve_token_is_read_from_the_environment():
|
||||
"""The PVE token must come from the injected environment, never a literal.
|
||||
|
||||
Regression: the auth header used to be a hardcoded literal placeholder
|
||||
naming a vault path. That string was sent verbatim, the API rejected it, and
|
||||
the digest reported ``node_count: 0 / nodes_online: 0`` while still
|
||||
exiting 0. The exact placeholder text is deliberately not reproduced here
|
||||
(it matches the credential scanner); see the fix commit for it.
|
||||
"""
|
||||
mod = load_script()
|
||||
assert hasattr(mod, "pve_auth"), "pve_auth() must exist to resolve the token at call time"
|
||||
sample = "unit-test-sample-value"
|
||||
with patch.dict("os.environ", {"PVE_TOKEN": sample}, clear=False):
|
||||
assert mod.pve_auth() == f"Authorization: {mod.PVE_AUTH_HEADER}{sample}"
|
||||
assert mod.pve_auth().endswith(sample)
|
||||
|
||||
|
||||
def test_missing_pve_token_is_degraded_not_a_placeholder():
|
||||
"""With no PVE_TOKEN, pve_get must return None (probe failure), not send a placeholder."""
|
||||
mod = load_script()
|
||||
env = {k: v for k, v in os.environ.items() if k != "PVE_TOKEN"}
|
||||
with patch.dict("os.environ", env, clear=True):
|
||||
assert mod.pve_get("/api2/json/nodes") is None, (
|
||||
"a missing PVE_TOKEN must degrade to None so the caller records a probe failure"
|
||||
)
|
||||
|
||||
|
||||
def test_unreachable_probe_is_recorded_as_a_failure():
|
||||
"""An unreachable probe must be recorded, so the run cannot pass silently."""
|
||||
mod = load_script()
|
||||
assert hasattr(mod, "PROBE_FAILURES"), "PROBE_FAILURES must exist"
|
||||
mod.PROBE_FAILURES.clear()
|
||||
with patch.object(mod, "pve_get", return_value=None):
|
||||
report = mod.collect()
|
||||
assert report["pve_probe_status"] == "unreachable"
|
||||
assert any("unreachable" in f for f in mod.PROBE_FAILURES), (
|
||||
f"unreachable probe must be recorded in PROBE_FAILURES, got {mod.PROBE_FAILURES}"
|
||||
)
|
||||
|
||||
|
||||
def test_pve_token_placeholder_is_gone():
|
||||
"""The literal placeholder must no longer appear anywhere in the script."""
|
||||
src = (Path(__file__).parent.parent / "scripts" / "daily-infra-report.py").read_text()
|
||||
assert "«vault:" not in src, "the unresolved vault placeholder must not remain in the script"
|
||||
assert "AUTH = \"Authorization" not in src, "the hardcoded AUTH literal must be gone"
|
||||
|
||||
|
||||
def test_nested_zulip_read_feeds_agent_card():
|
||||
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
|
||||
# Mock the http_get_body response with nested structure
|
||||
@@ -135,4 +183,32 @@ if __name__ == "__main__":
|
||||
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_pve_token_is_read_from_the_environment()
|
||||
print("✓ test_pve_token_is_read_from_the_environment passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_pve_token_is_read_from_the_environment failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_missing_pve_token_is_degraded_not_a_placeholder()
|
||||
print("✓ test_missing_pve_token_is_degraded_not_a_placeholder passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_missing_pve_token_is_degraded_not_a_placeholder failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_unreachable_probe_is_recorded_as_a_failure()
|
||||
print("✓ test_unreachable_probe_is_recorded_as_a_failure passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_unreachable_probe_is_recorded_as_a_failure failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_pve_token_placeholder_is_gone()
|
||||
print("✓ test_pve_token_placeholder_is_gone passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_pve_token_placeholder_is_gone failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
print("All tests passed!")
|
||||
|
||||
@@ -49,7 +49,7 @@ HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
|
||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||
|
||||
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
|
||||
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
||||
TANKO_VANTAGE = "192.168.68.12" # minipve — Tanko CT 112 via pct exec
|
||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||
|
||||
|
||||
@@ -82,7 +82,7 @@ done
|
||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||
cmd="${*: -1}"
|
||||
case "$host" in
|
||||
192.168.68.15)
|
||||
192.168.68.12)
|
||||
case "$cmd" in
|
||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||
@@ -98,8 +98,8 @@ exit 0
|
||||
"""
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
|
||||
# record every call (including notify) payloads.
|
||||
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||
# kagentz C3 public URL, and record every call (including notify) payloads.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
@@ -109,6 +109,8 @@ case "$*" in
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
*kagentz.sysloggh.net*)
|
||||
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
@@ -121,7 +123,8 @@ def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a_code="401", az_a2a_exit=0):
|
||||
az_a2a_code="401", az_a2a_exit=0,
|
||||
kagentz_public_code="302"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
@@ -154,6 +157,7 @@ def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
@@ -169,8 +173,9 @@ def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
||||
assert "Server: ✅ HTTP 200" in log
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "Tanko: ✅ service=active http=200" in log
|
||||
assert "kagentz: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "Result: ✅ All healthy" in log
|
||||
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||
|
||||
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||
assert proc.stdout == ""
|
||||
@@ -207,8 +212,8 @@ def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
||||
# The rest of the monitor still ran alongside the failing Tanko leg.
|
||||
log = log_path.read_text()
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "kagentz: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "kagentz C1: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
|
||||
|
||||
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||
@@ -216,9 +221,9 @@ def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||
assert "kagentz: ✅ A2A alive" not in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "kagentz C1: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||
assert "kagentz C1: ✅ A2A alive" not in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
assert "kagentz A2A server answered HTTP 500" in proc.stdout
|
||||
|
||||
|
||||
@@ -230,10 +235,10 @@ def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz: ❌ A2A down (HTTP 000)" in log
|
||||
assert "kagentz: ✅ A2A alive" not in log
|
||||
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||
assert "kagentz C1: ✅ A2A alive" not in log
|
||||
assert "unexpected" not in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
|
||||
|
||||
|
||||
|
||||
Executable
+108
@@ -0,0 +1,108 @@
|
||||
#!/bin/bash
|
||||
# test_pbs_gc_states.sh — Tests for PBS GC four-state logic
|
||||
# Self-contained: inlines the SSH replacement logic
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
TEST_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
SCRIPTS_DIR="$(dirname "$TEST_DIR")/scripts"
|
||||
PROXMOX_MONITOR="${1:-${SCRIPTS_DIR}/proxmox-monitor.sh}"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
# Create the Python replacement script
|
||||
REPLACE_SCRIPT=$(mktemp /tmp/replace_ssh_XXXXXX.py)
|
||||
cat > "$REPLACE_SCRIPT" << 'PYEOF'
|
||||
import sys
|
||||
import re
|
||||
|
||||
wrapper = sys.argv[1]
|
||||
monitor = sys.argv[2]
|
||||
|
||||
with open(monitor) as f:
|
||||
c = f.read()
|
||||
|
||||
pattern = r'PBS_GC_OUTPUT=\$\(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192\.168\.68\.6 \\\n "/sbin/pct exec 107 -- /sbin/proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null\)'
|
||||
replacement = 'PBS_GC_OUTPUT=$(cat "' + wrapper + '")'
|
||||
|
||||
if re.search(pattern, c):
|
||||
c = re.sub(pattern, replacement, c)
|
||||
|
||||
with open(monitor, 'w') as f:
|
||||
f.write(c)
|
||||
PYEOF
|
||||
|
||||
run_test() {
|
||||
local name="$1"
|
||||
local json="$2"
|
||||
local expected_behavior="$3"
|
||||
local expected_pattern="$4"
|
||||
|
||||
local wrapper monitor
|
||||
wrapper=$(mktemp /tmp/pbs-test-wrapper.XXXXXX)
|
||||
monitor=$(mktemp /tmp/pbs-test-monitor.XXXXXX)
|
||||
|
||||
printf '%s\n' "$json" > "$wrapper"
|
||||
cp "$PROXMOX_MONITOR" "$monitor"
|
||||
|
||||
# Replace the SSH call with cat "$wrapper"
|
||||
python3 "$REPLACE_SCRIPT" "$wrapper" "$monitor"
|
||||
|
||||
local output exit_code
|
||||
output=$(bash "$monitor" 2>&1)
|
||||
exit_code=$?
|
||||
|
||||
local ok=true
|
||||
if [ "$expected_behavior" = "fail" ]; then
|
||||
# Should fail with PBS GC error
|
||||
if ! echo "$output" | grep -q "🔴 PBS GC"; then ok=false; fi
|
||||
if [ $exit_code -eq 0 ]; then ok=false; fi
|
||||
else
|
||||
# Should pass with expected pattern
|
||||
if ! echo "$output" | grep -q "$expected_pattern"; then ok=false; fi
|
||||
if echo "$output" | grep -q "🔴 PBS GC"; then ok=false; fi
|
||||
fi
|
||||
|
||||
if $ok; then
|
||||
echo " ✅ $name"
|
||||
PASS=$((PASS + 1))
|
||||
else
|
||||
echo " 🔴 $name FAILED (exit=$exit_code)"
|
||||
echo "$output" | grep "PBS GC" | sed 's/^/ /'
|
||||
FAIL=$((FAIL + 1))
|
||||
fi
|
||||
|
||||
rm -f "$wrapper" "$monitor"
|
||||
}
|
||||
|
||||
echo "=== PBS GC Four-State Tests ==="
|
||||
echo "Script: $PROXMOX_MONITOR"
|
||||
echo ""
|
||||
|
||||
echo "1. probe-failed (unparseable JSON)"
|
||||
run_test "unparseable-json" 'NOT JSON {{{' "fail" "probe-failed"
|
||||
|
||||
echo "2. probe-failed (store not found)"
|
||||
run_test "store-missing" '[{"store": "other", "last-run-endtime": 1000}]' "fail" "probe-failed"
|
||||
|
||||
echo "3. probe-failed (empty body)"
|
||||
run_test "empty-body" "" "fail" "probe-failed"
|
||||
|
||||
echo "4. running (in progress - upid set, no last-run-endtime)"
|
||||
run_test "running" '[{"store": "storepve-datastore", "upid": "UPID:123:1:456:gc:root@pam", "pending-bytes": 100}]' "pass" "⏳ PBS GC: running"
|
||||
|
||||
echo "5. stale (last run >48h)"
|
||||
STALE=$(date -u -d "50 hours ago" +%s)
|
||||
run_test "stale" "[{\"store\": \"storepve-datastore\", \"last-run-endtime\": $STALE, \"pending-bytes\": 200}]" "fail" "stale"
|
||||
|
||||
echo "6. healthy (completed <48h)"
|
||||
HEALTHY=$(date -u -d "1 hour ago" +%s)
|
||||
run_test "healthy" "[{\"store\": \"storepve-datastore\", \"last-run-endtime\": $HEALTHY, \"pending-bytes\": 0}]" "pass" "✅ PBS GC: healthy"
|
||||
|
||||
echo ""
|
||||
echo "=== Results: $PASS passed, $FAIL failed ==="
|
||||
[ $FAIL -eq 0 ] && echo "✅ All passed" || echo "🔴 Some failed"
|
||||
|
||||
rm -f "$REPLACE_SCRIPT"
|
||||
[ $FAIL -eq 0 ] && exit 0 || exit 1
|
||||
@@ -104,6 +104,59 @@ def test_koby_ct111_is_on_storepve(ahc):
|
||||
assert ahc.AGENTS["koby"]["pve"] == "storepve"
|
||||
|
||||
|
||||
def test_tanko_ct112_is_probed_on_minipve(ahc, monkeypatch, capsys):
|
||||
# CT 112 (tanko) was live-migrated to minipve (.12) on 2026-09-27; the
|
||||
# amdpve mapping made `pct status 112` fail and read as ct-unreachable.
|
||||
# Execute the probe and assert the host the script actually contacts.
|
||||
probes = []
|
||||
monkeypatch.setattr(
|
||||
ahc, "ssh",
|
||||
lambda host, cmd, user="root": probes.append((host, cmd)) or "status: running",
|
||||
)
|
||||
ahc.FAIL.clear()
|
||||
ahc.REPORT_ONLY.clear()
|
||||
try:
|
||||
ahc.check_ct_liveness()
|
||||
tanko_hosts = [h for h, cmd in probes if cmd == "pct status 112 2>/dev/null"]
|
||||
assert tanko_hosts == ["192.168.68.12"]
|
||||
finally:
|
||||
ahc.FAIL.clear()
|
||||
ahc.REPORT_ONLY.clear()
|
||||
|
||||
|
||||
def test_gpu_rtx3090_probe_uses_llmuser_not_root(ahc, monkeypatch, capsys):
|
||||
# 2026-09-28: root SSH to .8 was lost when the guest was rebuilt; llmuser
|
||||
# owns llama-server and can read systemctl status and the :8080 pid. A root
|
||||
# probe reads as UNREACHABLE for a healthy host (the reported bug). Execute
|
||||
# check_gpu_ports() against an SSH boundary that only accepts llmuser@.8 and
|
||||
# assert the .8 leg does not produce the false UNREACHABLE failure.
|
||||
seen = []
|
||||
|
||||
def fake_ssh(host, cmd, user="root"):
|
||||
seen.append((host, user))
|
||||
if host == "192.168.68.8" and user != "llmuser":
|
||||
return None # root SSH denied -> baseline false UNREACHABLE
|
||||
if cmd.startswith("systemctl is-active"):
|
||||
return "active"
|
||||
if cmd.startswith("ss -tlnp"):
|
||||
return "48351"
|
||||
if cmd.startswith("curl"):
|
||||
return '{"status":"ok"}'
|
||||
return None
|
||||
|
||||
monkeypatch.setattr(ahc, "ssh", fake_ssh)
|
||||
ahc.FAIL.clear()
|
||||
try:
|
||||
ahc.check_gpu_ports()
|
||||
out = capsys.readouterr().out
|
||||
assert "gpu-unreachable:192.168.68.8" not in ahc.FAIL
|
||||
assert "\u2705 gpu-rtx3090 (.8): healthy" in out
|
||||
assert ("192.168.68.8", "llmuser") in seen
|
||||
assert not any(host == "192.168.68.8" and user == "root" for host, user in seen)
|
||||
finally:
|
||||
ahc.FAIL.clear()
|
||||
|
||||
|
||||
def test_report_only_legs_never_count_as_failures(ahc):
|
||||
for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)):
|
||||
ahc.FAIL.clear()
|
||||
|
||||
Executable
+282
@@ -0,0 +1,282 @@
|
||||
#!/usr/bin/env bash
|
||||
# Behavioural tests for scripts/revision-preflight.sh
|
||||
#
|
||||
# Every case builds a throwaway clone with a real bare remote, so origin/master
|
||||
# is genuine and the guard's fetch path is exercised. Nothing outside mktemp is
|
||||
# touched.
|
||||
#
|
||||
# The pre-fix draft is kept at tests/fixtures/revision-preflight.prefix.sh and
|
||||
# is run against the SAME cases, to prove these tests bite: the pre-fix guard
|
||||
# exits 0 where the fixed guard exits 1.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO="$(cd "$HERE/.." && pwd)"
|
||||
GUARD="$REPO/scripts/revision-preflight.sh"
|
||||
PREFIX_GUARD="$HERE/fixtures/revision-preflight.prefix.sh"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
FAILED_CASES=()
|
||||
|
||||
pass() { printf ' ✓ %s\n' "$1"; PASS=$((PASS + 1)); }
|
||||
fail() { printf ' ✗ %s\n' "$1"; FAIL=$((FAIL + 1)); FAILED_CASES+=("$1"); }
|
||||
|
||||
# Build a clone with a real remote; echo the clone path.
|
||||
make_clone() {
|
||||
local tmp
|
||||
tmp="$(mktemp -d)"
|
||||
git init --bare -q "$tmp/remote.git"
|
||||
git init -q "$tmp/clone"
|
||||
(
|
||||
cd "$tmp/clone" || exit 1
|
||||
git config user.email test@example.invalid
|
||||
git config user.name test
|
||||
mkdir -p scripts
|
||||
printf '#!/bin/bash\necho hello\n' > scripts/demo.sh
|
||||
chmod +x scripts/demo.sh
|
||||
git add -A
|
||||
git commit -qm init
|
||||
git branch -M master
|
||||
git remote add origin "$tmp/remote.git"
|
||||
git push -q origin master
|
||||
git fetch -q origin
|
||||
)
|
||||
echo "$tmp/clone"
|
||||
}
|
||||
|
||||
echo "== revision-preflight behavioural tests =="
|
||||
|
||||
# ── 1. match → exit 0 ────────────────────────────────────────────────────────
|
||||
echo "1. matching copy"
|
||||
C=$(make_clone)
|
||||
if out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); then
|
||||
pass "matching copy exits 0"
|
||||
else
|
||||
fail "matching copy should exit 0 (got $?, output: $out)"
|
||||
fi
|
||||
if [[ -z "$( "$GUARD" --quiet "$C/scripts/demo.sh" "$C" 2>&1 )" ]]; then
|
||||
pass "--quiet prints nothing on a match"
|
||||
else
|
||||
fail "--quiet should print nothing on a match"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 2. mismatch → exit 1 and names both hashes ───────────────────────────────
|
||||
echo "2. mismatched copy"
|
||||
C=$(make_clone)
|
||||
printf '#!/bin/bash\necho TAMPERED\n' > "$C/scripts/demo.sh"
|
||||
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 1 ]]; then pass "mismatch exits 1"; else fail "mismatch should exit 1 (got $rc)"; fi
|
||||
if [[ "$out" == *"MISMATCH"* ]]; then pass "mismatch says MISMATCH"; else fail "mismatch should say MISMATCH"; fi
|
||||
if [[ "$out" == *"executed:"* && "$out" == *"merged:"* ]]; then
|
||||
pass "mismatch prints both revisions"
|
||||
else
|
||||
fail "mismatch should print both revisions"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 3. paths under scripts/ resolve (the original basename defect) ───────────
|
||||
echo "3. repo-relative path resolution"
|
||||
C=$(make_clone)
|
||||
if "$GUARD" --quiet "$C/scripts/demo.sh" "$C" >/dev/null 2>&1; then
|
||||
pass "script under scripts/ resolves against origin/master"
|
||||
else
|
||||
fail "script under scripts/ must resolve (basename defect)"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 4. script that exists in NO revision → must fail ─────────────────────────
|
||||
echo "4. untracked script present in no revision"
|
||||
C=$(make_clone)
|
||||
printf '#!/bin/bash\necho never committed\n' > "$C/scripts/ghost.sh"
|
||||
out=$("$GUARD" "$C/scripts/ghost.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 1 ]]; then pass "ghost script exits 1"; else fail "ghost script must exit 1 (got $rc)"; fi
|
||||
if [[ "$out" == *"does not exist in"* ]]; then
|
||||
pass "ghost script says it is absent from the ref"
|
||||
else
|
||||
fail "ghost script should say it is absent from the ref"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
reason_of() { grep -m1 '^REASON=' <<<"$1" | cut -d= -f2-; }
|
||||
|
||||
# ── 5. unresolvable ref → CANNOT VERIFY (exit 2) ─────────────────────────────
|
||||
echo "5. unresolvable ref"
|
||||
C=$(make_clone)
|
||||
out=$("$GUARD" --no-fetch --ref origin/nope "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 2 ]]; then pass "unresolvable ref exits 2 (cannot verify)"; else fail "unresolvable ref must exit 2 (got $rc)"; fi
|
||||
if [[ "$(reason_of "$out")" == "cannot-verify:ref-unresolvable" ]]; then
|
||||
pass "unresolvable ref names cannot-verify:ref-unresolvable"
|
||||
else
|
||||
fail "unresolvable ref should name its class (got: $(reason_of "$out"))"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 5b. fetch failure → CANNOT VERIFY, and it names that ─────────────────────
|
||||
echo "5b. fetch failed"
|
||||
C=$(make_clone)
|
||||
( cd "$C" && git remote set-url origin /nonexistent/definitely-not-a-repo )
|
||||
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 2 ]]; then pass "fetch failure exits 2 (cannot verify)"; else fail "fetch failure must exit 2 (got $rc)"; fi
|
||||
if [[ "$(reason_of "$out")" == "cannot-verify:fetch-failed" ]]; then
|
||||
pass "fetch failure names cannot-verify:fetch-failed"
|
||||
else
|
||||
fail "fetch failure should name its class (got: $(reason_of "$out"))"
|
||||
fi
|
||||
if [[ "$out" == *"bound:"* ]]; then pass "fetch failure reports the bound"; else fail "fetch failure should report the bound"; fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 5c. fetch timeout → CANNOT VERIFY, bounded (never hangs) ─────────────────
|
||||
echo "5c. fetch timeout is bounded"
|
||||
C=$(make_clone)
|
||||
# a remote that will never answer: a fifo-backed git daemon is overkill, so use
|
||||
# a black-hole address with a 1s bound and assert we return promptly.
|
||||
( cd "$C" && git remote set-url origin http://10.255.255.1:9/never.git )
|
||||
start=$(date +%s)
|
||||
out=$("$GUARD" --fetch-timeout 1 "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||
elapsed=$(( $(date +%s) - start ))
|
||||
if [[ $rc -eq 2 ]]; then pass "timeout exits 2 (cannot verify)"; else fail "timeout must exit 2 (got $rc)"; fi
|
||||
if [[ "$(reason_of "$out")" == "cannot-verify:fetch-failed" ]]; then
|
||||
pass "timeout names cannot-verify:fetch-failed"
|
||||
else
|
||||
fail "timeout should name its class (got: $(reason_of "$out"))"
|
||||
fi
|
||||
if [[ $elapsed -le 10 ]]; then pass "timeout returned promptly (${elapsed}s, bound 1s)"; else fail "timeout did not bound the fetch (${elapsed}s)"; fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 6. missing script → MISMATCH:path-absent (exit 1) ────────────────────────
|
||||
echo "6. missing script"
|
||||
C=$(make_clone)
|
||||
out=$("$GUARD" "$C/scripts/nope.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 1 ]]; then pass "missing script exits 1"; else fail "missing script must exit 1 (got $rc)"; fi
|
||||
if [[ "$(reason_of "$out")" == "mismatch:path-absent" ]]; then
|
||||
pass "missing script names mismatch:path-absent"
|
||||
else
|
||||
fail "missing script should name its class (got: $(reason_of "$out"))"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 7. script outside the clone → MISMATCH (exit 1) ──────────────────────────
|
||||
echo "7. script outside the clone"
|
||||
C=$(make_clone)
|
||||
OUTSIDE=$(mktemp)
|
||||
printf '#!/bin/bash\necho outside\n' > "$OUTSIDE"
|
||||
out=$("$GUARD" "$OUTSIDE" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 1 ]]; then pass "outside script exits 1"; else fail "outside script must exit 1 (got $rc)"; fi
|
||||
rm -f "$OUTSIDE"; rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 7b. detached HEAD is named, not reported as a raw content mismatch ───────
|
||||
echo "7b. detached HEAD"
|
||||
C=$(make_clone)
|
||||
( cd "$C" && printf '#!/bin/bash\necho TAMPERED\n' > scripts/demo.sh \
|
||||
&& git add scripts/demo.sh && git commit -qm tamper && git checkout -q --detach HEAD )
|
||||
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 1 ]]; then pass "detached HEAD exits 1"; else fail "detached HEAD must exit 1 (got $rc)"; fi
|
||||
if [[ "$(reason_of "$out")" == "mismatch:detached-head" ]]; then
|
||||
pass "detached HEAD names mismatch:detached-head"
|
||||
else
|
||||
fail "detached HEAD should name its class (got: $(reason_of "$out"))"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 7c. a branch ahead of the ref is named as such, not as a raw mismatch ────
|
||||
echo "7c. clone ahead of the ref"
|
||||
C=$(make_clone)
|
||||
( cd "$C" && git checkout -q -b feature \
|
||||
&& printf '#!/bin/bash\necho FEATURE\n' > scripts/demo.sh \
|
||||
&& git add scripts/demo.sh && git commit -qm feature )
|
||||
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 1 ]]; then pass "clone ahead exits 1"; else fail "clone ahead must exit 1 (got $rc)"; fi
|
||||
if [[ "$(reason_of "$out")" == "mismatch:clone-ahead" ]]; then
|
||||
pass "clone ahead names mismatch:clone-ahead"
|
||||
else
|
||||
fail "clone ahead should name its class (got: $(reason_of "$out"))"
|
||||
fi
|
||||
if [[ "$out" == *"mid-review"* ]]; then pass "clone ahead explains it is a legitimate state"; else fail "clone ahead should explain the state"; fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 7d. a genuine content mismatch is named as content ──────────────────────
|
||||
echo "7d. genuine content mismatch"
|
||||
C=$(make_clone)
|
||||
printf '#!/bin/bash\necho TAMPERED\n' > "$C/scripts/demo.sh"
|
||||
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||
if [[ "$(reason_of "$out")" == "mismatch:content" ]]; then
|
||||
pass "hand-edit names mismatch:content"
|
||||
else
|
||||
fail "hand-edit should name mismatch:content (got: $(reason_of "$out"))"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 8. --no-fetch states the freshness assumption ────────────────────────────
|
||||
echo "8. --no-fetch states its assumption"
|
||||
C=$(make_clone)
|
||||
out=$("$GUARD" --no-fetch "$C/scripts/demo.sh" "$C" 2>&1)
|
||||
if [[ "$out" == *"freshness is assumed"* ]]; then
|
||||
pass "--no-fetch states the freshness assumption"
|
||||
else
|
||||
fail "--no-fetch should state the freshness assumption"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# ── 8b. contract-run.sh defaults to warn, not enforce ───────────────────────
|
||||
echo "8b. contract-run.sh default mode"
|
||||
DEFAULT=$(grep -m1 'CONTRACT_REVISION_PREFLIGHT:-' "$REPO/scripts/contract-run.sh" | sed 's/.*:-//; s/}.*//')
|
||||
if [[ "$DEFAULT" == "warn" ]]; then
|
||||
pass "contract-run.sh defaults to warn"
|
||||
else
|
||||
fail "contract-run.sh default must be warn (found: '$DEFAULT')"
|
||||
fi
|
||||
if grep -q 'CONTRACT_REVISION_PREFLIGHT=enforce' "$REPO/scripts/contract-run.sh"; then
|
||||
pass "enforce remains available and documented"
|
||||
else
|
||||
fail "enforce must remain documented"
|
||||
fi
|
||||
|
||||
# ── 9. the pre-fix guard must FAIL these same cases (proves the tests bite) ──
|
||||
echo "9. pre-fix draft fails the same cases (bite proof)"
|
||||
if [[ ! -f "$PREFIX_GUARD" ]]; then
|
||||
fail "pre-fix fixture missing: $PREFIX_GUARD"
|
||||
else
|
||||
# 9a. repo-relative path: pre-fix drops scripts/ and cannot resolve
|
||||
C=$(make_clone)
|
||||
out=$("$PREFIX_GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 0 && "$out" == *"could not resolve"* ]]; then
|
||||
pass "pre-fix: exits 0 and cannot resolve scripts/demo.sh (defect confirmed)"
|
||||
else
|
||||
fail "pre-fix should exit 0 with 'could not resolve' (got rc=$rc)"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# 9b. ghost script: pre-fix passes a script that exists in no revision
|
||||
C=$(make_clone)
|
||||
printf '#!/bin/bash\necho never committed\n' > "$C/scripts/ghost.sh"
|
||||
out=$("$PREFIX_GUARD" "$C/scripts/ghost.sh" "$C" 2>&1); rc=$?
|
||||
if [[ $rc -eq 0 ]]; then
|
||||
pass "pre-fix: PASSES a ghost script that exists in no revision (defect confirmed)"
|
||||
else
|
||||
fail "pre-fix was expected to wrongly pass the ghost script (got rc=$rc)"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
|
||||
# 9c. a file that DOES exist at the repo root still works pre-fix, showing
|
||||
# the defect is specific to nested paths
|
||||
C=$(make_clone)
|
||||
printf '#!/bin/bash\necho root\n' > "$C/rootlevel.sh"
|
||||
( cd "$C" && git add rootlevel.sh && git commit -qm root && git push -q origin master && git fetch -q origin )
|
||||
if "$PREFIX_GUARD" "$C/rootlevel.sh" "$C" >/dev/null 2>&1; then
|
||||
pass "pre-fix: root-level path resolves (so the defect is the basename, not git)"
|
||||
else
|
||||
fail "pre-fix should resolve a root-level tracked file"
|
||||
fi
|
||||
rm -rf "$(dirname "$C")"
|
||||
fi
|
||||
|
||||
echo
|
||||
echo " passed: $PASS failed: $FAIL"
|
||||
if [[ $FAIL -gt 0 ]]; then
|
||||
printf ' FAILED: %s\n' "${FAILED_CASES[@]}"
|
||||
exit 1
|
||||
fi
|
||||
echo "All revision-preflight tests passed."
|
||||
Executable
+153
@@ -0,0 +1,153 @@
|
||||
#!/usr/bin/env bash
|
||||
# test_secret_scan.sh — self-test for the commit-time secret guard.
|
||||
#
|
||||
# Run: bash tests/test_secret_scan.sh
|
||||
# Exit: 0 all cases passed, 1 a case failed.
|
||||
#
|
||||
# WHY THIS FILE EXISTS: a scanner that is never observed to fail is not a guard.
|
||||
# Every fixture below is fabricated and pattern-shaped; the test writes it to a
|
||||
# temp tree (a path no allowlist entry covers) and asserts the guard FAILS. The
|
||||
# same fixtures are deliberately listed in scripts/secret-allowlist.tsv, so the
|
||||
# repo-wide tree scan stays quiet while a planted copy still bites — that is the
|
||||
# difference between an explicit, reasoned exception and a guard trained to
|
||||
# ignore a word.
|
||||
#
|
||||
# Only bash + coreutils + grep. No python/node: the Gitea runner executes job
|
||||
# steps inside the runner container, which has neither.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
HERE=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||
ROOT=$(cd -- "$HERE/.." && pwd)
|
||||
SCAN="$ROOT/scripts/secret-scan.sh"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
LAST_OUT=""
|
||||
|
||||
ok() { PASS=$((PASS + 1)); echo " ✅ $1"; }
|
||||
bad() { FAIL=$((FAIL + 1)); echo " ❌ $1"; }
|
||||
|
||||
expect_exit() { # expect_exit <want-code> <label> <cmd...>
|
||||
local want="$1" label="$2"; shift 2
|
||||
local rc
|
||||
LAST_OUT=$("$@" 2>&1); rc=$?
|
||||
if [ "$rc" -eq "$want" ]; then ok "$label (exit $rc)"; else
|
||||
bad "$label (wanted exit $want, got $rc)"
|
||||
printf '%s\n' "$LAST_OUT" | sed 's/^/ /' | head -8
|
||||
fi
|
||||
}
|
||||
|
||||
expect_contains() { # expect_contains <label> <needle>
|
||||
if printf '%s' "$LAST_OUT" | grep -qF -- "$2"; then ok "$1"; else
|
||||
bad "$1 (output did not mention: $2)"
|
||||
fi
|
||||
}
|
||||
|
||||
TMPROOT=$(mktemp -d)
|
||||
trap 'rm -rf "$TMPROOT"' EXIT
|
||||
|
||||
echo "── secret-scan self-test ──"
|
||||
|
||||
# ── 1. Guard syntax ───────────────────────────────────────────────────────
|
||||
expect_exit 0 "scanner parses with bash -n" bash -n "$SCAN"
|
||||
|
||||
# ── 2. Guard FAILS on planted, pattern-matching fixtures ──────────────────
|
||||
mkdir -p "$TMPROOT/planted"
|
||||
cat > "$TMPROOT/planted/ops.env" <<'EOF'
|
||||
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
|
||||
EOF
|
||||
expect_exit 1 "planted sk-or-v1 key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted sk-or-v1 key names the openrouter-key rule" "[openrouter-key]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/curl.sh" <<'EOF'
|
||||
curl -s -H "Authorization: Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb" http://example.invalid/
|
||||
EOF
|
||||
expect_exit 1 "planted literal Bearer token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted Bearer token names the bearer-token rule" "[bearer-token]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/pve.sh" <<'EOF'
|
||||
AUTH="Authorization: PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555"
|
||||
EOF
|
||||
expect_exit 1 "planted Proxmox token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted Proxmox token names the proxmox-token rule" "[proxmox-token]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/deploy-key.pem" <<'EOF'
|
||||
-----BEGIN OPENSSH PRIVATE KEY-----
|
||||
b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW
|
||||
-----END OPENSSH PRIVATE KEY-----
|
||||
EOF
|
||||
expect_exit 1 "planted PEM private key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted PEM key names the private-key rule" "[private-key]"
|
||||
|
||||
# Prose is scanned exactly like code — the original exposures were in .md files.
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/handover.md" <<'EOF'
|
||||
- Admin credentials: `admin` / `correct-horse-battery-staple`
|
||||
EOF
|
||||
expect_exit 1 "planted prose credential line fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted prose line names the cred-prose rule" "[cred-prose]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/config.env" <<'EOF'
|
||||
DB_PASSWORD=correct-horse-battery-staple
|
||||
EOF
|
||||
expect_exit 1 "planted password assignment fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted password assignment names the secret-assign rule" "[secret-assign]"
|
||||
|
||||
# ── 3. Guard stays QUIET on inert values and on the real tree ─────────────
|
||||
mkdir -p "$TMPROOT/inert"
|
||||
cat > "$TMPROOT/inert/config.yaml" <<'EOF'
|
||||
api_key: not-needed
|
||||
bearer_token=monitor_key
|
||||
api_key: $LITELLM_API_KEY
|
||||
EOF
|
||||
expect_exit 0 "env refs, sentinels and variable names are not credentials" bash "$SCAN" --path "$TMPROOT/inert" --quiet
|
||||
|
||||
expect_exit 0 "current repo tree passes the guard" bash "$SCAN" --tree
|
||||
expect_contains "tree run reports the allowlisted exceptions it applied" "allowlisted exception(s)"
|
||||
|
||||
# ── 4. Allowlist entries are path-explicit, not word-based ────────────────
|
||||
# This exact line is allowlisted in infrastructure-control.prose.md; the same
|
||||
# text at an unlisted path must still fail, proving the exception is per-file
|
||||
# and reviewed, not a blanket "ignore the word vault".
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/unlisted.md" <<'EOF'
|
||||
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||
EOF
|
||||
expect_exit 1 "allowlisted text at an unlisted path still fails" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
|
||||
# ── 5. Commit-time mode: the guard blocks a STAGED credential ─────────────
|
||||
# A throwaway git repo with its own copy of the scanner, so this exercises the
|
||||
# real pre-commit path (--staged) without touching this repo's index.
|
||||
mkdir -p "$TMPROOT/repo/scripts"
|
||||
cp "$SCAN" "$TMPROOT/repo/scripts/secret-scan.sh"
|
||||
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/repo/scripts/secret-patterns.tsv"
|
||||
cp "$ROOT/scripts/secret-allowlist.tsv" "$TMPROOT/repo/scripts/secret-allowlist.tsv"
|
||||
git -C "$TMPROOT/repo" init -q
|
||||
git -C "$TMPROOT/repo" -c user.email=t@example.invalid -c user.name=test commit -q --allow-empty -m base
|
||||
cat > "$TMPROOT/repo/planted.env" <<'EOF'
|
||||
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
|
||||
EOF
|
||||
git -C "$TMPROOT/repo" add planted.env
|
||||
expect_exit 1 "staged credential fails at commit time (--staged)" bash "$TMPROOT/repo/scripts/secret-scan.sh" --staged --quiet
|
||||
expect_contains "staged credential names the openrouter-key rule" "[openrouter-key]"
|
||||
|
||||
# ── 6. Fail closed: an allowlist entry without a reason is a hard error ───
|
||||
mkdir -p "$TMPROOT/scanner" "$TMPROOT/clean"
|
||||
cp "$SCAN" "$TMPROOT/scanner/secret-scan.sh"
|
||||
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/scanner/secret-patterns.tsv"
|
||||
printf '*\t*.md\twhatever\n' > "$TMPROOT/scanner/secret-allowlist.tsv"
|
||||
echo "placeholder" > "$TMPROOT/clean/ok.md"
|
||||
expect_exit 2 "allowlist entry with no reason fails closed" bash "$TMPROOT/scanner/secret-scan.sh" --path "$TMPROOT/clean" --quiet
|
||||
|
||||
# ── Verdict ───────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
if [ "$FAIL" -gt 0 ]; then
|
||||
echo "❌ secret-scan self-test FAILED — $PASS passed, $FAIL failed"
|
||||
exit 1
|
||||
fi
|
||||
echo "✅ secret-scan self-test passed ($PASS cases)"
|
||||
@@ -0,0 +1,209 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Behavioural tests for zulip-monitor.sh kagentz C1/C3 legs and the Result verdict.
|
||||
|
||||
WHY THIS FILE EXISTS: the 2026-09-19 kagentz A2A outage was correctly detected
|
||||
by the monitor (C1 returned 000, the run logged an issue) but the lane's own
|
||||
summarization in ops.status wrote "OK" with the note "A2A server DOWN — expected
|
||||
(no credentials configured)". The optimistic verdict came from the lane, not the
|
||||
script. The fix adds a C3 public-access-path leg and makes the Result line say
|
||||
"INCIDENT" when issues are found, so the lane can quote it verbatim.
|
||||
|
||||
CONTRACT UNDER TEST:
|
||||
* C1 (A2A liveness, no credential): 000 → INCIDENT.
|
||||
* C3 (public access path, no credential): 502 → INCIDENT, 000 → INCIDENT,
|
||||
200/302/401 → alive.
|
||||
* Result verdict: when ISSUES > 0, the log's Result line says "INCIDENT",
|
||||
not just "issues found".
|
||||
* Healthy control: C1 401 + C3 302 → 0 issues, "all healthy".
|
||||
|
||||
HOW: behavioural execution using the sandbox pattern already in this repo
|
||||
(tests/test_mumuni_monitor_removal.py). The sandbox copies the shipped monitor
|
||||
verbatim, rewrites only its LOG constant, and runs it with stub ssh/curl on
|
||||
PATH. Each test asserts from the run's own log/verdict, not from file text.
|
||||
|
||||
Usage: python3 -m pytest tests/test_zulip_kagentz_legs.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import pathlib
|
||||
import stat
|
||||
import subprocess
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||
|
||||
TANKO_VANTAGE = "192.168.68.12" # minipve — Tanko CT 112 via pct exec
|
||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||
|
||||
|
||||
# ── Stub ssh: answers Tanko and Agent Zero probes by env vars ──────────
|
||||
|
||||
SSH_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub ssh: record the target host, then answer by host + remote command.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||
host=""
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
*@192.168.*) host="${a##*@}" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||
cmd="${*: -1}"
|
||||
case "$host" in
|
||||
192.168.68.12)
|
||||
case "$cmd" in
|
||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||
esac ;;
|
||||
192.168.68.14)
|
||||
case "$cmd" in
|
||||
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "${AZ_A2A_EXIT:-0}" ;;
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
# ── Stub curl: serves Zulip server, Abiba health, and C3 public URL ───
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||
# C3 public URL probe (https://kagentz.sysloggh.net/). Record every call.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
case " $* " in
|
||||
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
*kagentz.sysloggh.net*)
|
||||
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
path.write_text(body)
|
||||
path.chmod(path.stat().st_mode
|
||||
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a_code="401", az_a2a_exit=0,
|
||||
kagentz_public_code="302"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
Everything else — legs, labels, notify logic — is the shipped script.
|
||||
"""
|
||||
sandbox = tmp_path / "sandbox"
|
||||
bindir = sandbox / "bin"
|
||||
record = sandbox / "record"
|
||||
bindir.mkdir(parents=True)
|
||||
record.mkdir()
|
||||
|
||||
_write_exec(bindir / "ssh", SSH_STUB)
|
||||
_write_exec(bindir / "curl", CURL_STUB)
|
||||
|
||||
source = ZULIP_MONITOR.read_text()
|
||||
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||
log_path = sandbox / "zulip-health-monitor.log"
|
||||
script = sandbox / "zulip-monitor.sh"
|
||||
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||
|
||||
env = dict(os.environ)
|
||||
env.update({
|
||||
"PATH": f"{bindir}:{env['PATH']}",
|
||||
"RECORD_DIR": str(record),
|
||||
"TANKO_SVC": tanko_svc,
|
||||
"TANKO_HTTP": tanko_http,
|
||||
"AZ_A2A_CODE": az_a2a_code,
|
||||
"AZ_A2A_EXIT": str(az_a2a_exit),
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
return proc, record, log_path
|
||||
|
||||
|
||||
# ── Required behavioural cases ─────────────────────────────────────────
|
||||
|
||||
def test_c3_502_is_incident(tmp_path):
|
||||
"""C3 public leg returns 502 → the run's verdict is an INCIDENT,
|
||||
the C3 line names the 502, and the run is not summarised as healthy."""
|
||||
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="502")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C3 line names the 502.
|
||||
assert "kagentz C3: ❌ public URL 502 (upstream refused)" in log
|
||||
# The verdict is an INCIDENT, not healthy.
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
# The run is not summarised as healthy.
|
||||
assert "✅ 0 issues" not in log
|
||||
|
||||
|
||||
def test_c3_000_is_incident(tmp_path):
|
||||
"""C3 public leg returns 000 → INCIDENT."""
|
||||
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="000")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C3 line reports the connection failure.
|
||||
assert "kagentz C3: ❌ public URL down (HTTP 000)" in log
|
||||
# The verdict is an INCIDENT.
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
assert "✅ 0 issues" not in log
|
||||
|
||||
|
||||
def test_healthy_control_c1_401_c3_302(tmp_path):
|
||||
"""Healthy control: C1 401 plus C3 302 → 0 issues and a healthy verdict,
|
||||
proving the new leg cannot cry wolf."""
|
||||
proc, record, log_path = _run_monitor(tmp_path,
|
||||
az_a2a_code="401",
|
||||
kagentz_public_code="302")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# Both legs report alive.
|
||||
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||
# Zero issues, healthy verdict.
|
||||
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||
# No INCIDENT.
|
||||
assert "INCIDENT" not in log
|
||||
# No notify fired for kagentz.
|
||||
assert "kagentz public URL" not in proc.stdout
|
||||
assert "kagentz A2A server" not in proc.stdout
|
||||
|
||||
|
||||
def test_c1_000_is_incident(tmp_path):
|
||||
"""C1 returns 000 → INCIDENT, keeping the leg that actually caught this
|
||||
outage covered behaviourally."""
|
||||
proc, record, log_path = _run_monitor(tmp_path,
|
||||
az_a2a_code="000",
|
||||
az_a2a_exit=7,
|
||||
kagentz_public_code="302")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C1 line reports the A2A down.
|
||||
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||
# The verdict is an INCIDENT (even though C3 is healthy).
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
# The notify fired for the A2A down.
|
||||
assert "kagentz A2A server DOWN" in proc.stdout
|
||||
+73
-43
@@ -1,9 +1,9 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-health
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||
version: 3.3.0
|
||||
version: 3.4.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
report_only_agents:
|
||||
@@ -28,9 +28,9 @@ session start.
|
||||
## Requires
|
||||
|
||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||
- **SSH access** to minipve (192.168.68.12) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||
- **PM2** on localhost for pi process management
|
||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||
- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200`
|
||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||
- **Relay access** via RA-H OS MCP for alert delivery
|
||||
|
||||
@@ -126,13 +126,20 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
|
||||
- **On critical alert**: Escalate to relay message immediately, don't wait for schedule
|
||||
|
||||
## Execution
|
||||
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||
execution — it only proves the agent read the result and reported it. The actual
|
||||
monitoring work happens in the host cron job.
|
||||
|
||||
|
||||
### Liveness rule (scoped)
|
||||
|
||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN.
|
||||
|
||||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above).
|
||||
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
||||
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
||||
|
||||
@@ -221,20 +228,20 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
|
||||
| Crash loop >10/h | Alert user |
|
||||
|
||||
|
||||
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
|
||||
### Step 3: Platform B — Tanko (DSH on minipve CT 112)
|
||||
|
||||
Mumuni is out of scope for this host (see the note above): she runs on her own
|
||||
container and is monitored on her side.
|
||||
|
||||
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
|
||||
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
|
||||
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **amdpve**
|
||||
PVE host (**192.168.68.15**). Direct SSH to 192.168.68.122 is not a dependency
|
||||
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **minipve**
|
||||
PVE host (**192.168.68.12**). Direct SSH to 192.168.68.122 is not a dependency
|
||||
of this contract — per-worker key availability varies — so CT 112 probes run
|
||||
from the amdpve vantage via `pct exec`:
|
||||
from the minipve vantage via `pct exec`:
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.15 "pct exec 112 -- <command>"
|
||||
ssh root@192.168.68.12 "pct exec 112 -- <command>"
|
||||
```
|
||||
|
||||
> **By design (verified 2026-09-08):** the `dsh-web` gateway binds
|
||||
@@ -246,7 +253,7 @@ ssh root@192.168.68.15 "pct exec 112 -- <command>"
|
||||
**B1: Gateway Service State (Tanko)**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web"
|
||||
ssh root@192.168.68.12 "pct exec 112 -- systemctl is-active dsh-web"
|
||||
```
|
||||
|
||||
Expected: `active`. Anything else → gateway service down → apply the Tanko heal
|
||||
@@ -255,7 +262,7 @@ Expected: `active`. Anything else → gateway service down → apply the Tanko h
|
||||
**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
|
||||
ssh root@192.168.68.12 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
|
||||
```
|
||||
|
||||
Alive = **ANY** HTTP status response from the endpoint — the expected set is
|
||||
@@ -266,13 +273,13 @@ process answering `503` is running and self-heal must NOT restart-loop it.
|
||||
Down = connection refused (`000`) or timeout only. Statuses outside the
|
||||
expected set are logged/reported as a warning — reported, never healed on.
|
||||
|
||||
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)**
|
||||
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to minipve)**
|
||||
|
||||
```bash
|
||||
curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/
|
||||
```
|
||||
|
||||
Fallback only — used when the monitoring node has no pct/SSH path to amdpve.
|
||||
Fallback only — used when the monitoring node has no pct/SSH path to minipve.
|
||||
Alive = **ANY** HTTP status response from the endpoint — healthy signals are
|
||||
`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other
|
||||
status, including `404`/`5xx`, also counts alive: the endpoint is up and
|
||||
@@ -397,29 +404,29 @@ ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
|
||||
4. Every later request through `/` presents that cookie; the token is not needed
|
||||
again until the cookie expires or a new browser is used.
|
||||
|
||||
**Verification** (amdpve vantage):
|
||||
**Verification** (minipve vantage):
|
||||
```bash
|
||||
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
||||
ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
|
||||
# Expected: 302
|
||||
|
||||
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
||||
ssh root@192.168.68.12 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
||||
-w '%{http_code}\n' http://192.168.68.122:8081/"
|
||||
# Expected: 000
|
||||
|
||||
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
|
||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
||||
TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
||||
ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the minted dsh-auth-... cookie (authority
|
||||
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
|
||||
|
||||
# 4. Token refresh is non-disruptive and idempotent.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
||||
ssh root@192.168.68.12 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
||||
# Expected: "token unchanged; nginx not reloaded" when nothing changed
|
||||
```
|
||||
|
||||
@@ -430,32 +437,32 @@ fresh cookie. Both verified live 2026-09-11.
|
||||
|
||||
```bash
|
||||
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
|
||||
ssh root@192.168.68.12 "pct exec 112 -- systemctl restart dsh-web"
|
||||
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
|
||||
# until the socket answers (any status but 000) before asserting the cookie.
|
||||
for i in $(seq 1 60); do
|
||||
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
UP=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
|
||||
[ "$UP" != "000" ] && break
|
||||
sleep 2
|
||||
done
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the pre-restart cookie is still accepted.
|
||||
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
|
||||
# manual run may no-op on the flock, so poll until the include carries a token
|
||||
# the running process accepts (bounded wait) before the mint+reuse check.
|
||||
for i in $(seq 1 60); do
|
||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
||||
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
||||
CODE=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
|
||||
[ "$CODE" = "303" ] && break
|
||||
sleep 2
|
||||
done
|
||||
# Expected: 303 — the include now holds the token the running process accepts.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
||||
ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
||||
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the refreshed token minted a fresh cookie.
|
||||
```
|
||||
@@ -469,9 +476,10 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
|
||||
> monitor issued a restart for something that could not start, posting a false
|
||||
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
|
||||
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
|
||||
> liveness only, and a probe must never restart a platform.
|
||||
> liveness/response and public-path access only, and a probe must never restart
|
||||
> a platform.
|
||||
|
||||
**C1: A2A Server Health**
|
||||
**C1: A2A Server Health (no credential needed)**
|
||||
|
||||
```bash
|
||||
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
|
||||
@@ -479,9 +487,9 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
|
||||
```
|
||||
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down (INCIDENT). Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**C2: A2A Response Verification**
|
||||
**C2: A2A Response Verification (requires LITELLM_KEY)**
|
||||
|
||||
```bash
|
||||
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
|
||||
@@ -491,15 +499,30 @@ ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:
|
||||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||
```
|
||||
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set (this is a credential issue, NOT a server-down incident).
|
||||
|
||||
**C3: Public Access Path (no credential needed)**
|
||||
|
||||
```bash
|
||||
# Probes the public URL that NetBird proxies to the agent-zero container.
|
||||
# This is the captain's point of view: if the captain can't reach it, it's down.
|
||||
# 200/302/401 = alive, 502 = proxy's "upstream refused" page (INCIDENT),
|
||||
# connection failed (000) = INCIDENT. Never restarts anything.
|
||||
curl -s -o /dev/null --connect-timeout 10 --max-time 15 -w '%{http_code}' https://kagentz.sysloggh.net/
|
||||
```
|
||||
|
||||
Expected: `200` (Agent Zero login page), `302` (redirect), or `401` (auth-gated) = alive. `502` = NetBird proxy's "upstream refused" page (the container's port 80 is not listening) = INCIDENT. `000` (connection failed) = INCIDENT. Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**Platform C Actions**
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
|
||||
| C1 A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||
| C1 A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||
| C2 A2A returns 401 | Check `LITELLM_KEY` is set (credential issue, NOT server-down) |
|
||||
| C3 public URL returns `502` (upstream refused) | Alert only — the container's port 80 is not listening; investigate the agent-zero container |
|
||||
| C3 public URL returns `000` (connection failed) | Alert only — the public path is down; investigate the NetBird proxy or the container |
|
||||
| C3 public URL returns a status other than `200`/`302`/`401` | Log/report as a warning — reported, never healed on |
|
||||
|
||||
### Step 5: Global Checks
|
||||
|
||||
@@ -513,12 +536,19 @@ If any bot processes >50 bot-originated messages in 15min → warning.
|
||||
|
||||
### Step 6: Compile and Report
|
||||
|
||||
1. Compile all platform checks and severity
|
||||
2. Determine `overall_severity` from worst per-agent severity
|
||||
3. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||
4. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||
5. If any agent critical or >2 degraded: send relay message to user
|
||||
6. Update `last_check` timestamp in `### Maintains` snapshot
|
||||
1. Run `scripts/zulip-monitor.sh` and take its final `Result:` line as the
|
||||
authoritative run verdict. The verdict line is either
|
||||
`Result: ✅ 0 issues (all healthy)` or
|
||||
`Result: 🔴 INCIDENT — N issue(s) found`.
|
||||
2. Quote that `Result:` line verbatim in the status report. When it says
|
||||
`INCIDENT`, the run MUST be reported as an incident — never summarised as
|
||||
OK/healthy and never annotated as "expected".
|
||||
3. Compile all platform checks and severity
|
||||
4. Determine `overall_severity` from worst per-agent severity
|
||||
5. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||
6. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||
7. If any agent critical or >2 degraded: send relay message to user
|
||||
8. Update `last_check` timestamp in `### Maintains` snapshot
|
||||
|
||||
### Restart Debounce
|
||||
|
||||
|
||||
Reference in New Issue
Block a user