Compare commits
139
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4320369bb9 | ||
|
|
3b74ca28d1 | ||
|
|
af9397d672 | ||
|
|
97dc2d772f | ||
|
|
2238777a2f | ||
|
|
6b2ba1bba5 | ||
|
|
a48b947242 | ||
|
|
ba9d29b4b9 | ||
|
|
d6376e5142 | ||
|
|
2ea6b4fc17 | ||
|
|
c6fd8eece3 | ||
|
|
4c715526ef | ||
|
|
308265e7ce | ||
|
|
88e9243ce4 | ||
|
|
13ac189365 | ||
|
|
2851a0cfc8 | ||
|
|
e0c9852de8 | ||
|
|
bf3a1ba523 | ||
|
|
f230812e3a | ||
|
|
96769a103f | ||
|
|
d697baa7b6 | ||
|
|
fb7f351a2b | ||
|
|
e42b970dec | ||
|
|
eadb927ec1 | ||
|
|
d4e238047d | ||
|
|
73d5097555 | ||
|
|
c460ef905c | ||
|
|
d99b552448 | ||
|
|
fc0cd7a032 | ||
|
|
ba38efcd75 | ||
|
|
0d30091f62 | ||
|
|
0a41a2d584 | ||
|
|
de32f54337 | ||
|
|
042fd3ccc6 | ||
|
|
226f2ad55e | ||
|
|
552c776c0c | ||
|
|
778424acd4 | ||
|
|
c5a6dcd42a | ||
|
|
9faffe4f6b | ||
|
|
cd00cd0475 | ||
|
|
819cd53bce | ||
|
|
574cb99d76 | ||
|
|
9d64b0bd66 | ||
|
|
cdc7ad2c79 | ||
|
|
5409dfd73a | ||
|
|
8b2eba4f7a | ||
|
|
040fecef3e | ||
|
|
483a66b7b4 | ||
|
|
c0d04a2c02 | ||
|
|
fda6c844ff | ||
|
|
748ea389be | ||
|
|
f16a890d0e | ||
|
|
c666d3e15c | ||
|
|
c07aa5e382 | ||
|
|
87205e6fdb | ||
|
|
8f1e5eebc4 | ||
|
|
bb1b65340e | ||
|
|
9e87927444 | ||
|
|
b0683e9566 | ||
|
|
dd11c8f14f | ||
|
|
0732eed329 | ||
|
|
59ed7cdbf7 | ||
|
|
5d70bbf25b | ||
|
|
ccc916d1ec | ||
|
|
48fb263d4b | ||
|
|
64790ebd19 | ||
|
|
6fa0a255df | ||
|
|
c72436b406 | ||
|
|
a17379676d | ||
|
|
63990b84f7 | ||
|
|
7a5ddb46a9 | ||
|
|
568fec2efa | ||
|
|
641a52c6da | ||
|
|
58033f39c5 | ||
|
|
6c59988e7f | ||
|
|
9a2ee6faec | ||
|
|
e71ded3c8c | ||
|
|
a12abbeb14 | ||
|
|
0b9aebca37 | ||
|
|
fa458afa26 | ||
|
|
aa3da83af5 | ||
|
|
aee2de25ac | ||
|
|
76653381ec | ||
|
|
3b32cc9658 | ||
|
|
d07c4494b5 | ||
|
|
66ad5ac89d | ||
|
|
b10fd6fc98 | ||
|
|
8ff13d38f3 | ||
|
|
a13457bcd6 | ||
|
|
f59d1a2159 | ||
|
|
da8f5f43c9 | ||
|
|
8ae4b59150 | ||
|
|
ef7f90ef5a | ||
|
|
43e891e679 | ||
|
|
933cfd223b | ||
|
|
f77d6ca1d1 | ||
|
|
f4f8a4cab8 | ||
|
|
7e257ce512 | ||
|
|
c0454811bb | ||
|
|
3c7f5d7d65 | ||
|
|
03be9b13d0 | ||
|
|
c295322c85 | ||
|
|
385f7e0623 | ||
|
|
7efbfffe44 | ||
|
|
315fcbae23 | ||
|
|
93f15709d1 | ||
|
|
1137dd4582 | ||
|
|
7400dfd833 | ||
|
|
c65f5219e1 | ||
|
|
077972fa2b | ||
|
|
d2bca5405a | ||
|
|
8a5cba8515 | ||
|
|
20f882412f | ||
|
|
30b2fe3fdc | ||
|
|
5112c566c8 | ||
|
|
83307eb9b2 | ||
|
|
8245716286 | ||
|
|
cfb6c03572 | ||
|
|
85f70f65bc | ||
|
|
a820b3f7dd | ||
|
|
0b92ab17b1 | ||
|
|
9edefe036e | ||
|
|
dd6e1e8b22 | ||
|
|
c712d4faf0 | ||
|
|
7f62f19c24 | ||
|
|
57bfe7e06a | ||
|
|
9100ea3326 | ||
|
|
0f26119859 | ||
|
|
5c1c8d7c19 | ||
|
|
8a2ea2d0d7 | ||
|
|
dae8d14880 | ||
|
|
39209c7ac9 | ||
|
|
8ed3b9c606 | ||
|
|
6c616a9e58 | ||
|
|
b9b1712ac6 | ||
|
|
6fb411613e | ||
|
|
4bdd88613b | ||
|
|
bc7a55122f | ||
|
|
713b9ce80c |
@@ -71,6 +71,18 @@ jobs:
|
|||||||
git fetch origin "${{ gitea.ref }}" --depth=50
|
git fetch origin "${{ gitea.ref }}" --depth=50
|
||||||
git checkout "${{ gitea.sha }}"
|
git checkout "${{ gitea.sha }}"
|
||||||
|
|
||||||
|
- name: Committed-credential scan (secret guard)
|
||||||
|
run: |
|
||||||
|
# Fails the build on a credential-shaped string in the tree. Patterns
|
||||||
|
# live in scripts/secret-patterns.tsv; the only tolerated literal
|
||||||
|
# examples are in scripts/secret-allowlist.tsv, each with a reason.
|
||||||
|
# Do not turn this into a warning: a warning in a stream nobody reads
|
||||||
|
# is how six live credentials sat in this repo for weeks.
|
||||||
|
bash scripts/secret-scan.sh
|
||||||
|
|
||||||
|
- name: Secret guard self-test
|
||||||
|
run: bash tests/test_secret_scan.sh
|
||||||
|
|
||||||
- name: Structure + regression + consistency lint
|
- name: Structure + regression + consistency lint
|
||||||
run: bash scripts/prose-lint.sh
|
run: bash scripts/prose-lint.sh
|
||||||
|
|
||||||
|
|||||||
@@ -1 +1,2 @@
|
|||||||
__pycache__/
|
__pycache__/
|
||||||
|
state/host-disk-bands.json
|
||||||
|
|||||||
@@ -51,6 +51,12 @@ Two incidents taught us this:
|
|||||||
- `/grafana/` nginx route — was reverted Jul 2, must not reappear
|
- `/grafana/` nginx route — was reverted Jul 2, must not reappear
|
||||||
- `CT 122` or `CT 123` as CT ID labels — don't exist in the cluster
|
- `CT 122` or `CT 123` as CT ID labels — don't exist in the cluster
|
||||||
- These rules are hardcoded in `scripts/prose-lint.sh`
|
- These rules are hardcoded in `scripts/prose-lint.sh`
|
||||||
|
- **Committed-credential guard:** `scripts/secret-scan.sh` FAILS the build on
|
||||||
|
credential-shaped strings (patterns in `scripts/secret-patterns.tsv`, prose
|
||||||
|
included). Tolerated literals are listed one-per-example with a reason in
|
||||||
|
`scripts/secret-allowlist.tsv`; never allowlist a live credential. It runs in
|
||||||
|
the CI lint job, in `scripts/prose-lint.sh`, and via
|
||||||
|
`bash scripts/secret-scan.sh --staged` before committing.
|
||||||
|
|
||||||
### Stage 3 — AI Review
|
### Stage 3 — AI Review
|
||||||
- Diff is sent to `syslog-auto` model via LiteLLM
|
- Diff is sent to `syslog-auto` model via LiteLLM
|
||||||
|
|||||||
@@ -0,0 +1,142 @@
|
|||||||
|
---
|
||||||
|
kind: responsibility
|
||||||
|
name: agent-health-check
|
||||||
|
description: >
|
||||||
|
Consolidated agent health verification for the LiteLLM + GPU + Zulip +
|
||||||
|
gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours
|
||||||
|
via cron and on-demand via "run contract: agent-health-check". Verifies:
|
||||||
|
LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway
|
||||||
|
liveness, CT liveness, gateway log health, config YAML integrity,
|
||||||
|
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
|
||||||
|
title: Agent Health Check — Consolidated
|
||||||
|
version: 1.0.0
|
||||||
|
runtime_contract: 2
|
||||||
|
agent: abiba
|
||||||
|
---
|
||||||
|
|
||||||
|
# Agent Health Check
|
||||||
|
|
||||||
|
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
|
||||||
|
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only.
|
||||||
|
|
||||||
|
**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
|
||||||
|
|
||||||
|
## Requires
|
||||||
|
|
||||||
|
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
|
||||||
|
- **SSH access** to GPU hosts — `llmuser` on .8 (owns `llama-server`), `root` on .110 and .15 — and agent CTs (.122, .129, .114, .24)
|
||||||
|
- **Python 3** for script execution
|
||||||
|
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- last_check: timestamp — When the last full diagnostic ran
|
||||||
|
- overall_severity: "healthy" | "degraded" | "critical"
|
||||||
|
- liteLLM_keys: map of agent → key validity
|
||||||
|
- gpu_ports: map of host → port conflict status
|
||||||
|
- agents: map of agent → streaming health + gateway liveness
|
||||||
|
- ct_liveness: map of CT → active status
|
||||||
|
- config_integrity: map of config file → valid/invalid
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||||
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||||
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||||
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||||
|
execution — it only proves the agent read the result and reported it. The actual
|
||||||
|
monitoring work happens in the host cron job.
|
||||||
|
|
||||||
|
|
||||||
|
### check-health
|
||||||
|
|
||||||
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool
|
||||||
|
calls; never repeat a prior report unless a live probe fails.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Run the consolidated health check script
|
||||||
|
python3 /root/scripts/agent-health-check.py --json
|
||||||
|
```
|
||||||
|
|
||||||
|
**Report format**: Begin every report with the **absolute path the script
|
||||||
|
executed from** so a stale-consumer report is distinguishable from a real fault
|
||||||
|
at read time. Summarize actual results from each check. Apply the standing probe
|
||||||
|
rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
|
||||||
|
|
||||||
|
**Expected output**: JSON with `overall_severity` field. If `healthy`, report
|
||||||
|
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
||||||
|
failures and their severity.
|
||||||
|
|
||||||
|
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
|
||||||
|
MUST include one clause per check leg, in every state: healthy, degraded/warn,
|
||||||
|
skipped, or failed. A missing leg must never look the same as a healthy leg.
|
||||||
|
Required legs and their templates in every state:
|
||||||
|
|
||||||
|
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
|
||||||
|
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
|
||||||
|
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
|
||||||
|
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
|
||||||
|
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
|
||||||
|
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
|
||||||
|
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
|
||||||
|
- The GPU leg has six non-healthy states the code can produce:
|
||||||
|
(i) `gpu-unreachable:{host}` — SSH probe failed;
|
||||||
|
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
|
||||||
|
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
|
||||||
|
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
|
||||||
|
(v) unit active, /health body contains "error" — error response;
|
||||||
|
(vi) unit active, /health body unrecognised — unknown health.
|
||||||
|
In every case the failing host and reason must be named.
|
||||||
|
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
|
||||||
|
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
|
||||||
|
- skipped: `CTs: SKIPPED (SSH access unavailable)`
|
||||||
|
- `Vault secrets: 3/3 present`
|
||||||
|
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
|
||||||
|
- skipped: `Vault secrets: SKIPPED (vault not configured)`
|
||||||
|
|
||||||
|
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
|
||||||
|
long as the host is identifiable from context; the full `probe-failed: <target>
|
||||||
|
<kind>` form is required when a leg reports a failure in the detail section.
|
||||||
|
|
||||||
|
### Probe Shape (per standing rules from 1150.msg)
|
||||||
|
|
||||||
|
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||||
|
service answered — report the code, never "down". A redirect is not a failure.
|
||||||
|
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||||
|
2. **A failed probe is never a service verdict.** Print
|
||||||
|
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
||||||
|
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
||||||
|
then report.
|
||||||
|
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
||||||
|
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
||||||
|
(retried at 25s: also timeout)" is actionable.
|
||||||
|
|
||||||
|
## Strategies
|
||||||
|
|
||||||
|
### When LiteLLM keys are invalid
|
||||||
|
Report the specific agent + key name. Do not attempt to fix — credential
|
||||||
|
rotation is a separate operation.
|
||||||
|
|
||||||
|
### When GPU port conflicts are detected
|
||||||
|
Report the conflicting ports and processes. Do not kill processes — that's a
|
||||||
|
destructive action requiring captain approval.
|
||||||
|
|
||||||
|
### When gateway liveness is degraded
|
||||||
|
Report the specific CT + gateway status. Do not restart unless the restart
|
||||||
|
debounce window has passed.
|
||||||
|
|
||||||
|
### When CT liveness is down
|
||||||
|
Report the specific CT. Do not restart — that's a destructive action.
|
||||||
|
|
||||||
|
### When config YAML is invalid
|
||||||
|
Report the specific file + parse error. Do not fix — that's a config change.
|
||||||
|
|
||||||
|
### When gateway log health is degraded
|
||||||
|
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
|
||||||
|
|
||||||
|
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
|
||||||
|
|
||||||
|
## Continuity
|
||||||
|
|
||||||
|
- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
|
||||||
|
- **On `agent-health` command**: Run on-demand and report to user
|
||||||
|
- **On critical alert**: Escalate to relay message immediately
|
||||||
@@ -16,8 +16,8 @@ litellm.exceptions.AuthenticationError: OpenrouterException -
|
|||||||
```
|
```
|
||||||
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
||||||
|
|
||||||
**Old Key**: `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`
|
**Old Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||||
**New Key**: `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`
|
**New Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||||
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||||
|
|
||||||
### 2. Telegram Bot Conflict (CRITICAL)
|
### 2. Telegram Bot Conflict (CRITICAL)
|
||||||
@@ -48,14 +48,14 @@ McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 sec
|
|||||||
```bash
|
```bash
|
||||||
# Container .env update
|
# Container .env update
|
||||||
sudo docker exec agent-zero bash -c '
|
sudo docker exec agent-zero bash -c '
|
||||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
|
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»|" /a0/usr/.env
|
||||||
'
|
'
|
||||||
```
|
```
|
||||||
|
|
||||||
**Verification**:
|
**Verification**:
|
||||||
```bash
|
```bash
|
||||||
curl -s https://openrouter.ai/api/v1/auth/key \
|
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||||
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
|
-H "Authorization: Bearer «vault: agents/production OPENROUTER_API_KEY»" | python3 -m json.tool
|
||||||
```
|
```
|
||||||
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
||||||
|
|
||||||
@@ -140,7 +140,7 @@ Added section:
|
|||||||
|
|
||||||
| Component | Status | Details |
|
| Component | Status | Details |
|
||||||
|-----------|--------|---------|
|
|-----------|--------|---------|
|
||||||
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
|
| **OpenRouter Key** | ✅ Valid | `«vault: agents/production OPENROUTER_API_KEY»` user verified, |
|
||||||
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
||||||
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
||||||
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
||||||
|
|||||||
@@ -54,7 +54,7 @@ description: >
|
|||||||
```
|
```
|
||||||
|
|
||||||
4. **Return status**
|
4. **Return status**
|
||||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-synthetic...", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||||
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
||||||
- If vault secret is missing: `{ vault_synced: false }`
|
- If vault secret is missing: `{ vault_synced: false }`
|
||||||
|
|
||||||
@@ -89,8 +89,8 @@ description: >
|
|||||||
|
|
||||||
| Field | Value |
|
| Field | Value |
|
||||||
|-------|-------|
|
|-------|-------|
|
||||||
| **Key Prefix** | `sk-or-v1-0af3f3` |
|
| **Key Prefix** | `«vault: agents/production OPENROUTER_API_KEY»` |
|
||||||
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
|
| **Full Key** | `«vault: agents/production OPENROUTER_API_KEY»` (in vault + /a0/usr/.env) |
|
||||||
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
||||||
| **Free Tier** | No |
|
| **Free Tier** | No |
|
||||||
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
||||||
@@ -101,7 +101,7 @@ description: >
|
|||||||
|
|
||||||
| Date | Action | Notes |
|
| Date | Action | Notes |
|
||||||
|------|--------|-------|
|
|------|--------|-------|
|
||||||
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
| 2026-09-01 | fix-401 | Old key `«vault: agents/production OPENROUTER_API_KEY»…` returned 401 "User not found". Replaced with new key `«vault: agents/production OPENROUTER_API_KEY»…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||||
|
|
||||||
## Infrastructure References
|
## Infrastructure References
|
||||||
|
|
||||||
|
|||||||
+90
-23
@@ -94,7 +94,16 @@ def audit(path):
|
|||||||
cfg = yaml.safe_load(f)
|
cfg = yaml.safe_load(f)
|
||||||
|
|
||||||
model = cfg.get("model", {})
|
model = cfg.get("model", {})
|
||||||
fb = cfg.get("fallback_providers", {})
|
fb_raw = cfg.get("fallback_providers", {})
|
||||||
|
# Normalize: fallback_providers may be a dict (single provider) or a list of dicts
|
||||||
|
# (one entry per fallback). Both shapes are valid; we must handle both without crashing.
|
||||||
|
if isinstance(fb_raw, dict):
|
||||||
|
fb_entries = [fb_raw]
|
||||||
|
elif isinstance(fb_raw, list):
|
||||||
|
fb_entries = fb_raw
|
||||||
|
else:
|
||||||
|
fb_entries = [fb_raw] # Let it fail the check below as malformed
|
||||||
|
fb = fb_entries[0] if fb_entries else {}
|
||||||
comp = cfg.get("compression", {})
|
comp = cfg.get("compression", {})
|
||||||
aux = cfg.get("auxiliary", {})
|
aux = cfg.get("auxiliary", {})
|
||||||
deleg = cfg.get("delegation", {})
|
deleg = cfg.get("delegation", {})
|
||||||
@@ -114,12 +123,22 @@ def audit(path):
|
|||||||
)
|
)
|
||||||
|
|
||||||
# --- Rule 5: Main Config Base URL ---
|
# --- Rule 5: Main Config Base URL ---
|
||||||
expected_base = "http://192.168.68.116/v1"
|
# Canonical internal base (hermes-key-enforcement.prose.md:19/38/57/84/91/96)
|
||||||
check(
|
# and public base (serves /v1 only, per 2026-09-19 probe from CT 116).
|
||||||
model.get("base_url") == expected_base,
|
# The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only.
|
||||||
"Rule 5",
|
# FAIL anything else (do not widen to accept any path ending in /v1).
|
||||||
f"model.base_url must be {expected_base} (got {model.get('base_url')!r}) — /v1 not /litellm/v1",
|
# Internal /v1 is non-canonical but working (authenticated via nginx), so WARN not FAIL.
|
||||||
)
|
canonical_internal = "http://192.168.68.116/litellm/v1"
|
||||||
|
public_host = "https://litellm.sysloggh.net/v1"
|
||||||
|
non_canonical_internal = "http://192.168.68.116/v1"
|
||||||
|
allowed_bases = (canonical_internal, public_host)
|
||||||
|
actual_base = model.get("base_url")
|
||||||
|
if actual_base in allowed_bases:
|
||||||
|
check(True, "Rule 5", f"model.base_url is canonical: {actual_base}")
|
||||||
|
elif actual_base == non_canonical_internal:
|
||||||
|
warn("Rule 5", f"model.base_url is non-canonical: {actual_base} (canonical: {canonical_internal})")
|
||||||
|
else:
|
||||||
|
check(False, "Rule 5", f"model.base_url must be one of {allowed_bases} (got {actual_base!r})")
|
||||||
|
|
||||||
# --- Rule 6: max_tokens Is Required ---
|
# --- Rule 6: max_tokens Is Required ---
|
||||||
check(
|
check(
|
||||||
@@ -197,22 +216,32 @@ def audit(path):
|
|||||||
"Rule 14",
|
"Rule 14",
|
||||||
f"delegation.provider must be 'harness' (got {deleg.get('provider')!r})",
|
f"delegation.provider must be 'harness' (got {deleg.get('provider')!r})",
|
||||||
)
|
)
|
||||||
check(
|
# Check each fallback entry. A malformed entry (not a mapping) is a VIOLATION, not a crash.
|
||||||
fb.get("provider") == "deepseek",
|
for idx, entry in enumerate(fb_entries):
|
||||||
"Rule 14",
|
prefix = f"fallback_providers[{idx}]"
|
||||||
f"fallback_providers.provider must be 'deepseek' (got {fb.get('provider')!r}) — "
|
if not isinstance(entry, dict):
|
||||||
f"true fallback diversity, not same endpoint as primary",
|
check(
|
||||||
)
|
False,
|
||||||
check(
|
"Rule 14",
|
||||||
fb.get("model") == "deepseek-v4-flash",
|
f"{prefix} must be a mapping (got {type(entry).__name__})",
|
||||||
"Rule 14",
|
)
|
||||||
f"fallback_providers.model must be 'deepseek-v4-flash' (got {fb.get('model')!r})",
|
continue
|
||||||
)
|
check(
|
||||||
check(
|
entry.get("provider") == "deepseek",
|
||||||
fb.get("api_key_env") == "DEEPSEEK_API_KEY",
|
"Rule 14",
|
||||||
"Rule 14",
|
f"{prefix}.provider must be 'deepseek' (got {entry.get('provider')!r}) — "
|
||||||
f"fallback_providers.api_key_env must be DEEPSEEK_API_KEY (got {fb.get('api_key_env')!r})",
|
f"true fallback diversity, not same endpoint as primary",
|
||||||
)
|
)
|
||||||
|
check(
|
||||||
|
entry.get("model") == "deepseek-v4-flash",
|
||||||
|
"Rule 14",
|
||||||
|
f"{prefix}.model must be 'deepseek-v4-flash' (got {entry.get('model')!r})",
|
||||||
|
)
|
||||||
|
check(
|
||||||
|
entry.get("api_key_env") == "DEEPSEEK_API_KEY",
|
||||||
|
"Rule 14",
|
||||||
|
f"{prefix}.api_key_env must be DEEPSEEK_API_KEY (got {entry.get('api_key_env')!r})",
|
||||||
|
)
|
||||||
|
|
||||||
# --- custom_providers sanity ---
|
# --- custom_providers sanity ---
|
||||||
check(
|
check(
|
||||||
@@ -262,6 +291,44 @@ def audit(path):
|
|||||||
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# --- MCP Server Checks (Rule 15) ---
|
||||||
|
# Valid MCP server endpoints
|
||||||
|
VALID_MCP_ENDPOINTS = {
|
||||||
|
'ra-h-os': 'http://192.168.68.65:3100/mcp',
|
||||||
|
'litellm': 'https://litellm.sysloggh.net/mcp',
|
||||||
|
}
|
||||||
|
|
||||||
|
# Check MCP servers if they exist
|
||||||
|
mcp_servers = cfg.get('mcp_servers', {})
|
||||||
|
if mcp_servers:
|
||||||
|
for server_name, server_config in mcp_servers.items():
|
||||||
|
url = server_config.get('url', '')
|
||||||
|
|
||||||
|
# Check endpoint validity
|
||||||
|
if server_name in VALID_MCP_ENDPOINTS:
|
||||||
|
expected = VALID_MCP_ENDPOINTS[server_name]
|
||||||
|
if url == expected:
|
||||||
|
check(True, 'Rule 15', f'MCP server "{server_name}" URL is correct: {url}')
|
||||||
|
else:
|
||||||
|
check(False, 'Rule 15', f'MCP server "{server_name}" URL is incorrect: {url} (expected: {expected})')
|
||||||
|
else:
|
||||||
|
warn('Rule 15', f'MCP server "{server_name}" URL may need validation (not in known list): {url}')
|
||||||
|
|
||||||
|
# Check for proper authentication
|
||||||
|
headers = server_config.get('headers', {})
|
||||||
|
has_auth = False
|
||||||
|
for key, value in headers.items():
|
||||||
|
if 'key' in key.lower() or 'auth' in key.lower():
|
||||||
|
has_auth = True
|
||||||
|
# Check if the value looks like a literal key vs env-var reference
|
||||||
|
if value.startswith('Bearer ') and value[7:].startswith('sk-'):
|
||||||
|
check(True, 'Rule 15', f'MCP server "{server_name}" has valid auth header: {key}')
|
||||||
|
else:
|
||||||
|
warn('Rule 15', f'MCP server "{server_name}" header may use env-var instead of literal key: {key} = {value}')
|
||||||
|
break
|
||||||
|
if not has_auth:
|
||||||
|
warn('Rule 15', f'MCP server "{server_name}" has no authentication header')
|
||||||
|
|
||||||
# --- Report ---
|
# --- Report ---
|
||||||
print(f"{'=' * 60}")
|
print(f"{'=' * 60}")
|
||||||
print(f"Hermes Config Audit: {path}")
|
print(f"Hermes Config Audit: {path}")
|
||||||
|
|||||||
@@ -0,0 +1,173 @@
|
|||||||
|
# Search ranking policy for the agent-consumption layer.
|
||||||
|
#
|
||||||
|
# Everything here is CONFIG, not code, so it is reviewable and changeable without
|
||||||
|
# touching the module. Read by scripts/search-agent-consume.py.
|
||||||
|
#
|
||||||
|
# Why this file exists: multi-engine aggregation returns results with no
|
||||||
|
# filtering, no dedupe and no reranking. On 2026-09-26 that put a shopping page
|
||||||
|
# and a dictionary definition into "best practices agent context management",
|
||||||
|
# and put four SEO blogs ABOVE the actual Proxmox forum threads on a precise
|
||||||
|
# technical query. Identical queries also ranked differently between runs, which
|
||||||
|
# is the strongest argument for a deterministic layer rather than hoping the
|
||||||
|
# engines behave.
|
||||||
|
|
||||||
|
version: 1
|
||||||
|
|
||||||
|
# ── Non-answers: dropped outright, never returned ────────────────────────────
|
||||||
|
# These are pages that cannot answer a question: navigational homepages,
|
||||||
|
# shopping/product pages, dictionary definitions, and login walls.
|
||||||
|
non_answer:
|
||||||
|
# URL path is empty -> it is a site's front door, not an answer. Still allowed
|
||||||
|
# when the host is explicitly preferred (see prefer_domains), because some
|
||||||
|
# docs/repo front doors ARE the answer.
|
||||||
|
host_root: true
|
||||||
|
path_patterns:
|
||||||
|
- '/dictionary/'
|
||||||
|
- '/dictionary?'
|
||||||
|
- '/wiki/Wiktionary:'
|
||||||
|
- '/search?'
|
||||||
|
- '/cart'
|
||||||
|
- '/checkout'
|
||||||
|
- '/login'
|
||||||
|
- '/signin'
|
||||||
|
- '/sign-in'
|
||||||
|
- '/account/login'
|
||||||
|
- '/shop/'
|
||||||
|
- '/store/'
|
||||||
|
- '/dp/' # Amazon-style product URL
|
||||||
|
- '/gp/product/'
|
||||||
|
- '/add-to-cart'
|
||||||
|
- '/checkout'
|
||||||
|
# NOTE: '/products/' and '/product/' were REMOVED as path patterns. They fired
|
||||||
|
# on docs.digitalocean.com/products/inference/... — a legitimate documentation
|
||||||
|
# page — which the 2026-09-26 before/after run caught. Shopping is caught by
|
||||||
|
# the shopping HOST list instead, which does not have that false positive.
|
||||||
|
# Query strings that betray a search/shopping surface rather than an article.
|
||||||
|
query_keys:
|
||||||
|
- 'q'
|
||||||
|
- 'query'
|
||||||
|
- 's'
|
||||||
|
- 'search'
|
||||||
|
- 'add-to-cart'
|
||||||
|
# Hosts that are shopping/retail and never answer a technical question.
|
||||||
|
hosts:
|
||||||
|
- bestbuy.com
|
||||||
|
- amazon.com
|
||||||
|
- ebay.com
|
||||||
|
- walmart.com
|
||||||
|
- etsy.com
|
||||||
|
- aliexpress.com
|
||||||
|
- merriam-webster.com
|
||||||
|
- dictionary.com
|
||||||
|
- thesaurus.com
|
||||||
|
- vocabulary.com
|
||||||
|
- collinsdictionary.com
|
||||||
|
|
||||||
|
# ── Demotion: ranked below everything else, never dropped ────────────────────
|
||||||
|
# Low-authority content farms / SEO aggregators. Demoted rather than dropped so
|
||||||
|
# a genuinely useful hit is not lost, but it can never outrank a primary source.
|
||||||
|
# Reviewable: add or remove hosts here, no code change required.
|
||||||
|
demote_domains:
|
||||||
|
- medium.com
|
||||||
|
- sparkco.ai
|
||||||
|
- mindstudio.ai
|
||||||
|
- aitechmonk.com
|
||||||
|
- stackai.com
|
||||||
|
- agentic-design.ai
|
||||||
|
- voxfor.com
|
||||||
|
- bigiron.cc
|
||||||
|
- linuxoperatingsystem.net
|
||||||
|
- riparazioneserver.com
|
||||||
|
- rossmanngroup.com
|
||||||
|
- dev.to
|
||||||
|
- hashnode.dev
|
||||||
|
- substack.com
|
||||||
|
- towardsdatascience.com
|
||||||
|
- analyticsvidhya.com
|
||||||
|
- geeksforgeeks.org
|
||||||
|
- tutorialspoint.com
|
||||||
|
- javatpoint.com
|
||||||
|
- w3schools.com
|
||||||
|
- scaler.com
|
||||||
|
- simplilearn.com
|
||||||
|
- udemy.com
|
||||||
|
- coursera.org
|
||||||
|
|
||||||
|
# ── Preference: promoted above the default rank ──────────────────────────────
|
||||||
|
# Primary sources: upstream repositories, official docs, Q&A, vendor
|
||||||
|
# engineering blogs. These are what an agent should be reading.
|
||||||
|
prefer_domains:
|
||||||
|
# upstream repositories and code hosting
|
||||||
|
- github.com
|
||||||
|
- gitlab.com
|
||||||
|
- codeberg.org
|
||||||
|
- sourceforge.net
|
||||||
|
- kernel.org
|
||||||
|
- git.kernel.org
|
||||||
|
# Q&A
|
||||||
|
- stackoverflow.com
|
||||||
|
- stackexchange.com
|
||||||
|
- superuser.com
|
||||||
|
- serverfault.com
|
||||||
|
- askubuntu.com
|
||||||
|
- discourse.org
|
||||||
|
# vendor / project documentation and forums
|
||||||
|
- proxmox.com
|
||||||
|
- forum.proxmox.com
|
||||||
|
- pve.proxmox.com
|
||||||
|
- docs.python.org
|
||||||
|
- developer.mozilla.org
|
||||||
|
- kernelnewbies.org
|
||||||
|
- man7.org
|
||||||
|
- gnu.org
|
||||||
|
- debian.org
|
||||||
|
- ubuntu.com
|
||||||
|
- redhat.com
|
||||||
|
- kernel.dk # io_uring / Jens Axboe
|
||||||
|
- github.io # project pages (docs, papers) — promoted, not authoritative by itself
|
||||||
|
# vendor engineering blogs
|
||||||
|
- anthropic.com
|
||||||
|
- openai.com
|
||||||
|
- googleblog.com
|
||||||
|
- developers.googleblog.com
|
||||||
|
- engineering.fb.com
|
||||||
|
- netflixtechblog.com
|
||||||
|
- aws.amazon.com
|
||||||
|
- cloud.google.com
|
||||||
|
- microsoft.com
|
||||||
|
- learn.microsoft.com
|
||||||
|
- apple.com
|
||||||
|
- nvidia.com
|
||||||
|
- intel.com
|
||||||
|
- amd.com
|
||||||
|
- redislabs.com
|
||||||
|
- cloudflare.com
|
||||||
|
- langchain.com
|
||||||
|
- jetbrains.com
|
||||||
|
- cursor.com
|
||||||
|
# community discussion with high signal
|
||||||
|
- news.ycombinator.com
|
||||||
|
- lobste.rs
|
||||||
|
- reddit.com
|
||||||
|
|
||||||
|
# ── Ranking weights ──────────────────────────────────────────────────────────
|
||||||
|
# Final score = engine_score - demote_penalty + prefer_bonus, then a stable
|
||||||
|
# tiebreak on original position so ordering is reproducible run to run.
|
||||||
|
ranking:
|
||||||
|
demote_penalty: 1000
|
||||||
|
prefer_bonus: 100
|
||||||
|
# Results that several engines independently returned are more likely real.
|
||||||
|
multi_engine_bonus: 25
|
||||||
|
# Shallow paths (e.g. /blog/x) are slightly less likely to be primary docs.
|
||||||
|
host_root_allowed_when_preferred: true
|
||||||
|
|
||||||
|
# ── Extraction budget (criterion 4) ──────────────────────────────────────────
|
||||||
|
# Return CONTENT, not just links, so an agent gets usable material in ONE call.
|
||||||
|
extraction:
|
||||||
|
top_n: 5 # how many results get page text extracted
|
||||||
|
total_chars: 12000 # global budget across all extracted items
|
||||||
|
per_item_chars: 4000 # cap for any single item, so one page cannot eat the budget
|
||||||
|
timeout_seconds: 45 # per scrape
|
||||||
|
# If extraction fails, the result is still returned with an empty excerpt —
|
||||||
|
# a link is better than nothing, but the failure is recorded in the output.
|
||||||
|
on_failure: keep_with_empty_excerpt
|
||||||
+118
-2
@@ -38,6 +38,7 @@ index:
|
|||||||
by_category:
|
by_category:
|
||||||
compliance:
|
compliance:
|
||||||
- hermes-key-enforcement
|
- hermes-key-enforcement
|
||||||
|
- litellm-api-keys
|
||||||
- hermes-config-template
|
- hermes-config-template
|
||||||
- hermes-agent-baseline
|
- hermes-agent-baseline
|
||||||
monitoring:
|
monitoring:
|
||||||
@@ -46,6 +47,7 @@ index:
|
|||||||
- infrastructure-monitoring
|
- infrastructure-monitoring
|
||||||
- zulip-health
|
- zulip-health
|
||||||
- litellm-health
|
- litellm-health
|
||||||
|
- daily-health-digest
|
||||||
remediation:
|
remediation:
|
||||||
- litellm-self-heal
|
- litellm-self-heal
|
||||||
- pm2-self-heal
|
- pm2-self-heal
|
||||||
@@ -95,12 +97,14 @@ index:
|
|||||||
- infrastructure-maintenance
|
- infrastructure-maintenance
|
||||||
- pm2-self-heal
|
- pm2-self-heal
|
||||||
- disk-gc-threat-response
|
- disk-gc-threat-response
|
||||||
|
- daily-health-digest
|
||||||
gpu:
|
gpu:
|
||||||
- gpu-monitor
|
- gpu-monitor
|
||||||
- gpu-fleet
|
- gpu-fleet
|
||||||
proxmox:
|
proxmox:
|
||||||
- proxmox-monitor
|
- proxmox-monitor
|
||||||
litellm:
|
litellm:
|
||||||
|
- litellm-api-keys
|
||||||
- litellm-health
|
- litellm-health
|
||||||
- litellm-self-heal
|
- litellm-self-heal
|
||||||
memory:
|
memory:
|
||||||
@@ -628,7 +632,7 @@ contracts:
|
|||||||
sensitivity: high
|
sensitivity: high
|
||||||
status: active
|
status: active
|
||||||
owner: abiba
|
owner: abiba
|
||||||
version: 3.3.0
|
version: 3.4.0
|
||||||
trigger:
|
trigger:
|
||||||
type: scheduled
|
type: scheduled
|
||||||
cadence: '*/15 * * * *'
|
cadence: '*/15 * * * *'
|
||||||
@@ -639,7 +643,7 @@ contracts:
|
|||||||
timeout: 120
|
timeout: 120
|
||||||
requires:
|
requires:
|
||||||
- Zulip API key for abiba-bot@chat.sysloggh.net
|
- Zulip API key for abiba-bot@chat.sysloggh.net
|
||||||
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
|
- SSH access to minipve (192.168.68.12) for Tanko (CT 112) and the Agent Zero Docker host (.14)
|
||||||
verification:
|
verification:
|
||||||
postconditions:
|
postconditions:
|
||||||
- check: bot registration active
|
- check: bot registration active
|
||||||
@@ -1869,6 +1873,118 @@ contracts:
|
|||||||
drift_alerts: []
|
drift_alerts: []
|
||||||
# Koby Report-Only Registry (2026-08-17 — Captain)
|
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||||
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||||
|
- name: litellm-api-keys
|
||||||
|
file: litellm-api-keys.prose.md
|
||||||
|
kind: function
|
||||||
|
category: compliance
|
||||||
|
sensitivity: critical
|
||||||
|
status: active
|
||||||
|
owner: abiba
|
||||||
|
version: 1.1.0
|
||||||
|
trigger:
|
||||||
|
type: on_demand
|
||||||
|
cadence: null
|
||||||
|
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
|
||||||
|
cron_job_id: null
|
||||||
|
execution:
|
||||||
|
agent: abiba
|
||||||
|
timeout: 120
|
||||||
|
requires: []
|
||||||
|
protocol:
|
||||||
|
- Load contract from prose-contracts/main
|
||||||
|
- Retrieve master key from Infisical (project=infrastructure env=production)
|
||||||
|
- Read live key-scoped model roster from CT 116 /v1/models
|
||||||
|
- Create/rotate/verify the requested agent key with an EXPLICIT models list
|
||||||
|
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
|
||||||
|
verification:
|
||||||
|
postconditions:
|
||||||
|
- check: standard agent key is local-only
|
||||||
|
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
|
||||||
|
expect: 0 cloud models
|
||||||
|
- check: key exists with correct alias
|
||||||
|
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
|
||||||
|
expect: 200 with matching alias
|
||||||
|
artifact: key creation/rotation report
|
||||||
|
receipt:
|
||||||
|
format: json
|
||||||
|
storage: ~/.hermes/runs/litellm-api-keys/
|
||||||
|
graph_node: true
|
||||||
|
escalation:
|
||||||
|
info:
|
||||||
|
action: log_to_receipt
|
||||||
|
notify: []
|
||||||
|
warning:
|
||||||
|
action: relay_alert
|
||||||
|
notify:
|
||||||
|
- abiba
|
||||||
|
- mumuni
|
||||||
|
critical:
|
||||||
|
action: relay_alert
|
||||||
|
notify:
|
||||||
|
- abiba
|
||||||
|
- mumuni
|
||||||
|
- ops
|
||||||
|
- name: daily-health-digest
|
||||||
|
file: daily-health-digest.prose.md
|
||||||
|
kind: function
|
||||||
|
category: monitoring
|
||||||
|
sensitivity: normal
|
||||||
|
status: active
|
||||||
|
owner: abiba
|
||||||
|
version: 1.0.0
|
||||||
|
trigger:
|
||||||
|
type: scheduled
|
||||||
|
cadence: 30 10 * * *
|
||||||
|
description: Daily at 10:30 UTC, dispatched on CT 100 as a firstmate message
|
||||||
|
to the ops lane, which executes the pinned producer
|
||||||
|
cron_job_id: null
|
||||||
|
execution:
|
||||||
|
agent: abiba
|
||||||
|
timeout: 300
|
||||||
|
requires:
|
||||||
|
- infisical (vault credentials injected at run time)
|
||||||
|
protocol:
|
||||||
|
- 'Execute from the PINNED clone only: /root/abiba-workspace/projects/prose-contracts'
|
||||||
|
- cd /root/abiba-workspace/projects/prose-contracts
|
||||||
|
- infisical run --env=prod -- python3 scripts/daily-infra-report.py
|
||||||
|
- 'Never execute from a per-agent working copy (treehouse) - it drifts onto feature branches'
|
||||||
|
- 'On failure: do not treat a 0/0 Proxmox section as evidence about the estate - it means could not look'
|
||||||
|
verification:
|
||||||
|
postconditions:
|
||||||
|
- check: every Proxmox probe is reachable
|
||||||
|
verify: >-
|
||||||
|
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
|
||||||
|
| grep -c 'pve_probe_status'
|
||||||
|
expect: '1'
|
||||||
|
- check: all nodes reported online
|
||||||
|
verify: >-
|
||||||
|
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
|
||||||
|
| grep 'nodes_online'
|
||||||
|
expect: nodes_online == node_count
|
||||||
|
artifact: timestamped HTML dashboard emailed to jerome@sysloggh.com
|
||||||
|
verify_commands:
|
||||||
|
- infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
|
||||||
|
- python3 -m pytest tests/test_daily_infra_report.py -q
|
||||||
|
delivery:
|
||||||
|
transport: zulip-dm-attachment
|
||||||
|
recipient_user_id: 9
|
||||||
|
sender: abiba-bot@chat.sysloggh.net
|
||||||
|
key_source: abiba-bot Zulip key already on the execution host, read from the
|
||||||
|
600-mode env file /root/.pi/agent/extensions/zulip/.env
|
||||||
|
key_policy: do NOT add a vault entry - that is a captain decision under the auth-keys charter
|
||||||
|
body: short Markdown pointer; the HTML attachment IS the report
|
||||||
|
artifact: /var/log/daily-infra-report/infra-report-<UTCstamp>.html
|
||||||
|
note: Replaced SMTP/mail on 2026-09-26 by captain decision. Removes the Google
|
||||||
|
dependency entirely; closes daily-digest-mail-transport-20260921.
|
||||||
|
exit_semantics:
|
||||||
|
'1': missing PVE_TOKEN, unreachable Proxmox probe, missing/rejected Zulip
|
||||||
|
credential, or a failed upload/post - raises an alert
|
||||||
|
'0': healthy delivery only - there is no degraded delivery leg any more
|
||||||
|
depends_on: []
|
||||||
|
last_run: null
|
||||||
|
last_status: null
|
||||||
|
drift_alerts: []
|
||||||
|
|
||||||
koby_report_only: true
|
koby_report_only: true
|
||||||
koby_host: "CT 111 (tdunna)"
|
koby_host: "CT 111 (tdunna)"
|
||||||
koby_ip: ".129"
|
koby_ip: ".129"
|
||||||
|
|||||||
@@ -0,0 +1,203 @@
|
|||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: daily-health-digest
|
||||||
|
description: >
|
||||||
|
Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox
|
||||||
|
nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails
|
||||||
|
it as an HTML report.
|
||||||
|
|
||||||
|
Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to
|
||||||
|
the ops lane, which executes the producer below. Until 2026-09-25 this ran
|
||||||
|
with NO contract file at all, which is why the choice of execution copy was
|
||||||
|
silently the operator's rather than the contract's.
|
||||||
|
|
||||||
|
EXECUTION IS PINNED. The producer must be run from the clone named under
|
||||||
|
"Execution pinning" — not from an agent working copy.
|
||||||
|
|
||||||
|
Exit-code semantics (as they actually behave, verified 2026-09-25):
|
||||||
|
* missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert
|
||||||
|
* missing or rejected Zulip credential -> exit 1 (delivery is the only
|
||||||
|
output path, so it is a real failure, not a degraded leg)
|
||||||
|
* delivery failure -> exit 1, and the report body is printed AND persisted
|
||||||
|
so the content is never swallowed
|
||||||
|
|
||||||
|
version: 2.0.0
|
||||||
|
---
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Give one daily, machine-collected picture of the estate so drift and outages
|
||||||
|
are seen the day they happen rather than when something breaks. It is a
|
||||||
|
*report*, not a repair: it changes nothing.
|
||||||
|
|
||||||
|
## Execution pinning
|
||||||
|
|
||||||
|
**Pinned execution path:**
|
||||||
|
|
||||||
|
```
|
||||||
|
/root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py
|
||||||
|
```
|
||||||
|
|
||||||
|
**Pinned clone:** `/root/abiba-workspace/projects/prose-contracts`
|
||||||
|
|
||||||
|
That is the cron's `FM_HOME` clone and the only stable, non-ephemeral copy.
|
||||||
|
The treehouse clone (`/root/.treehouse/agent-workspace-*/…/projects/prose-contracts`)
|
||||||
|
is a **per-agent working copy and must NOT be pinned or executed from** — it
|
||||||
|
drifts onto feature branches, which is exactly how a stale producer reported a
|
||||||
|
stale picture and nobody noticed.
|
||||||
|
|
||||||
|
See `docs/contract-execution-pinning.md`. Schedule and alerting live in
|
||||||
|
`/etc/cron.d/contract-runner` on CT 100:
|
||||||
|
|
||||||
|
```
|
||||||
|
30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1
|
||||||
|
```
|
||||||
|
|
||||||
|
Invocation (credentials come from the vault; never inline them):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd /root/abiba-workspace/projects/prose-contracts
|
||||||
|
infisical run --env=prod -- python3 scripts/daily-infra-report.py
|
||||||
|
```
|
||||||
|
|
||||||
|
## Output shape
|
||||||
|
|
||||||
|
Modes:
|
||||||
|
|
||||||
|
| invocation | effect |
|
||||||
|
| --- | --- |
|
||||||
|
| *(none)* | collect, build the HTML dashboard, email it |
|
||||||
|
| `--test-email` | same but with a `🧪 TEST —` subject prefix |
|
||||||
|
| `--json` | print the collected data as JSON to stdout and **send no email** |
|
||||||
|
|
||||||
|
`--json` emits a single object with these top-level keys (observed on a live
|
||||||
|
run 2026-09-25):
|
||||||
|
|
||||||
|
| key | type | meaning |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `nodes` | object (5) | per-node cpu/ram/disk/uptime/status |
|
||||||
|
| `node_count`, `nodes_online` | int | Proxmox node totals |
|
||||||
|
| `pve_probe_status`, `resources_probe_status` | `ok`\|`unreachable` | probe outcome |
|
||||||
|
| `total_vms`, `running_vms`, `stopped_vms`, `vms_by_node` | — | guest inventory |
|
||||||
|
| `storage`, `nfs` | array | datastore and mount usage |
|
||||||
|
| `litellm` | object | inference checks |
|
||||||
|
| `zulip_ext` | object | Zulip queue/serving state |
|
||||||
|
| `agents` | object | per-agent health |
|
||||||
|
| `docker_vm`, `docker_syslog`, `docker_netbird`, `endpoints` | object/array | Docker hosts and probed endpoints |
|
||||||
|
|
||||||
|
## What a healthy run looks like
|
||||||
|
|
||||||
|
```
|
||||||
|
$ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
|
||||||
|
"nodes_online": 5, "node_count": 5, "pve_probe_status": "ok",
|
||||||
|
"resources_probe_status": "ok", "running_vms": 22, "total_vms": 22
|
||||||
|
EXIT=0
|
||||||
|
```
|
||||||
|
|
||||||
|
and in delivery mode:
|
||||||
|
|
||||||
|
```
|
||||||
|
report ready: 16208 chars of HTML (delivered as a file attachment)
|
||||||
|
Sending to the captain's Zulip DM...
|
||||||
|
✅ Delivered to Zulip DM (user 9), message id 86221, attachment 16208 bytes
|
||||||
|
at /user_uploads/2/45/m1cQesBFV78BGeNY2lN8xkN5/infra-report-20260926-153406.html
|
||||||
|
```
|
||||||
|
|
||||||
|
Healthy means: every probe reports `ok`, `nodes_online == node_count`, and the
|
||||||
|
email leg reports a successful send.
|
||||||
|
|
||||||
|
## Exit-code semantics — as they actually behave
|
||||||
|
|
||||||
|
Verified on 2026-09-25 by running each case deliberately.
|
||||||
|
|
||||||
|
| condition | exit | alert | notes |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| all probes reachable, email sent | 0 | — | healthy |
|
||||||
|
| **missing `PVE_TOKEN`** | **1** | yes | `PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down)`, and `cluster resources unreachable` |
|
||||||
|
| **Proxmox probe unreachable** | **1** | yes | same path as above; `pve_probe_status: unreachable` |
|
||||||
|
| **missing/rejected Zulip credential** | **1** | yes | delivery is the only output path; report printed and persisted |
|
||||||
|
| **upload or message post fails** | **1** | yes | report printed and persisted; message names which step failed |
|
||||||
|
|
||||||
|
The distinction is deliberate and must not be flattened:
|
||||||
|
|
||||||
|
* A **missing PVE token or an unreachable probe is a real failure** — the report
|
||||||
|
would otherwise claim zero nodes and still look successful. That was the
|
||||||
|
2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1.
|
||||||
|
* A **missing email credential is survivable** — the report is still produced
|
||||||
|
and is still useful. It is a `DEGRADED` leg and exits 0 by design.
|
||||||
|
|
||||||
|
`PROBE_FAILURES` and `DEGRADED_LEGS` are separate lists for exactly this
|
||||||
|
reason. Do not merge them.
|
||||||
|
|
||||||
|
## Delivery: Zulip DM carrying the report as an HTML ATTACHMENT
|
||||||
|
|
||||||
|
Captain's decision 2026-09-26, clarified the same day: the digest is delivered to
|
||||||
|
his **Zulip DM (user id 9)** from `abiba-bot@chat.sysloggh.net`, as an **HTML
|
||||||
|
FILE** — an attachment, not HTML rendered in the message body and not a Markdown
|
||||||
|
translation of it.
|
||||||
|
|
||||||
|
* the styled dashboard is built exactly as before and written to
|
||||||
|
`/var/log/daily-infra-report/infra-report-<UTCstamp>.html`;
|
||||||
|
* it is uploaded through `POST /api/v1/user_uploads`;
|
||||||
|
* the **message body stays short Markdown** — subject line, top-line status
|
||||||
|
(nodes online, guests running, any degraded legs), and a link to the
|
||||||
|
attachment. The attachment IS the report; the body does not reproduce it.
|
||||||
|
|
||||||
|
This removes the Google dependency entirely: **no SMTP, no `EMAIL_PASSWORD`, no
|
||||||
|
app password, nothing to rotate.** `daily-digest-mail-transport-20260921` is
|
||||||
|
closed under this option.
|
||||||
|
|
||||||
|
The **10,000-character message cap does not apply** — it bounds message TEXT
|
||||||
|
only, and the report travels as a file. Do not shrink the report to fit it.
|
||||||
|
|
||||||
|
The credential is abiba-bot's Zulip key already on the execution host at
|
||||||
|
`/root/.pi/agent/extensions/zulip/.env` (`ABIBA_ZULIP_API_KEY`, mode 600,
|
||||||
|
root-readable). **Do not place a new credential in the vault** — under the
|
||||||
|
auth-keys charter that is a captain decision.
|
||||||
|
|
||||||
|
## What counts as a failure
|
||||||
|
|
||||||
|
A run FAILS (exit 1) when the report cannot be trusted or delivered:
|
||||||
|
|
||||||
|
* any probe is unreachable, so a section would silently be empty;
|
||||||
|
* `PVE_TOKEN` is missing;
|
||||||
|
* the Zulip credential is missing or rejected, or the upload/post fails.
|
||||||
|
|
||||||
|
There is **no degraded delivery leg any more**. Delivery is the only output
|
||||||
|
path, so a missing credential is a failure rather than a survivable degradation —
|
||||||
|
the previous "missing `EMAIL_PASSWORD` still exits 0" rule is retired with the
|
||||||
|
mail transport.
|
||||||
|
|
||||||
|
**A delivery failure must never swallow the report.** On failure the script
|
||||||
|
prints the report body to stdout *and* leaves the HTML artifact on disk, so the
|
||||||
|
content is always recoverable from the run log. That closes the queued defect
|
||||||
|
where a failed send printed only the transport error and the report never
|
||||||
|
surfaced.
|
||||||
|
|
||||||
|
## Failure behaviour
|
||||||
|
|
||||||
|
* Non-zero exit with the alert text above; on the scheduled path the dispatch is
|
||||||
|
a firstmate message, so the ops lane sees it and reports it.
|
||||||
|
* On a probe failure the report must **not** be treated as evidence about the
|
||||||
|
estate — a `0/0` Proxmox section means "could not look", not "nothing there".
|
||||||
|
That reading is why the 2026-09-25 defect went unnoticed.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# data path, no email
|
||||||
|
cd /root/abiba-workspace/projects/prose-contracts
|
||||||
|
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \
|
||||||
|
| grep -E 'pve_probe_status|node_count|nodes_online'
|
||||||
|
|
||||||
|
# delivery path
|
||||||
|
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-zulip
|
||||||
|
```
|
||||||
|
|
||||||
|
Regression tests: `tests/test_daily_infra_report.py` (7 tests). Four of them
|
||||||
|
fail against the pre-fix script, which is what makes them bite.
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- daily-infra-dashboard: { status: "ok|undelivered", transport: zulip-dm-attachment, last_check: timestamp }
|
||||||
|
- pve-probe: { status: "ok|unreachable", last_check: timestamp }
|
||||||
@@ -43,7 +43,7 @@ Docker hosts get special attention:
|
|||||||
|
|
||||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
|
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
|
||||||
|
|
||||||
## Threat Levels
|
## Threat Levels (GUEST filesystems)
|
||||||
|
|
||||||
| Level | Threshold | Response | Escalation |
|
| Level | Threshold | Response | Escalation |
|
||||||
|-------|-----------|----------|------------|
|
|-------|-----------|----------|------------|
|
||||||
@@ -53,6 +53,39 @@ Docker hosts get special attention:
|
|||||||
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
|
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
|
||||||
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
|
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
|
||||||
|
|
||||||
|
## Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)
|
||||||
|
|
||||||
|
Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are **report-only** — no automatic deletion of media or datastore content ever.
|
||||||
|
|
||||||
|
| Level | Threshold | Response | Escalation |
|
||||||
|
|-------|-----------|----------|------------|
|
||||||
|
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
||||||
|
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
||||||
|
| **HOST-RED** | 95% | Name the volume + % + absolute free space; **media volumes: "capacity decision — owner to decide"; pbs-datastore/host-root: "immediate owner attention"** | Zulip DM + channel alert (state-change only) |
|
||||||
|
|
||||||
|
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
||||||
|
|
||||||
|
**State lives in a small JSON state file:** the scanner resolves it to an absolute path from the script's location — `$(dirname "$0")/state/host-disk-bands.json` (i.e., the `state/` directory next to the `scripts/` directory in the prose-contracts repo). The state file is gitignored runtime state — the scanner creates it on first run. Keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
|
||||||
|
|
||||||
|
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
|
||||||
|
|
||||||
|
```
|
||||||
|
storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
|
||||||
|
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
|
||||||
|
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
|
||||||
|
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
|
||||||
|
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN
|
||||||
|
```
|
||||||
|
|
||||||
|
**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.
|
||||||
|
|
||||||
|
**Action classes by volume type:**
|
||||||
|
- **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
|
||||||
|
- **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
|
||||||
|
- **pbs-datastore** (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.
|
||||||
|
|
||||||
|
**Justification (measured 2026-09-15, firstmate):** storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.
|
||||||
|
|
||||||
## Requires
|
## Requires
|
||||||
|
|
||||||
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
|
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
|
||||||
@@ -127,6 +160,17 @@ from the `report_only_guests` YAML block above.
|
|||||||
- `retry`: 2 attempts for SSH failures before marking a CT unreachable
|
- `retry`: 2 attempts for SSH failures before marking a CT unreachable
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||||
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||||
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||||
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||||
|
execution — it only proves the agent read the result and reported it. The actual
|
||||||
|
monitoring work happens in the host cron job.
|
||||||
|
|
||||||
|
|
||||||
|
### Host filesystems: report-only, NEVER auto-delete
|
||||||
|
|
||||||
|
Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.
|
||||||
|
|
||||||
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
|
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
|
||||||
|
|
||||||
@@ -191,6 +235,21 @@ call summary-reporter
|
|||||||
plan: plan
|
plan: plan
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## GC SCHEDULE (PBS datastore only)
|
||||||
|
|
||||||
|
The PBS GC schedule is defined in ONE authoritative place: `/etc/cron.d/pbs-gc` on storepve.
|
||||||
|
The schedule is `0 20 * * *` (20:00 LOCAL = 00:00 UTC, since host timezone is America/New_York).
|
||||||
|
This applies **only** to the PBS datastore (`/tank/pbs-backup`), NOT to media volumes.
|
||||||
|
Media volumes (/media/*) are report-only at all threat levels.
|
||||||
|
|
||||||
|
The cron runs `/usr/local/bin/pbs-gc.sh` which executes:
|
||||||
|
```bash
|
||||||
|
proxmox-backup-manager garbage-collection start storepve-datastore
|
||||||
|
```
|
||||||
|
This is NOT a `prune` operation; it is a GC pass that reclaims unreferenced chunks.
|
||||||
|
There is no `--keep-daily` flag; retention is governed by jobs.cfg (keep-daily=35).
|
||||||
|
The GC does not touch media volumes or any other filesystem.
|
||||||
|
|
||||||
## GC Strategies by Host Type
|
## GC Strategies by Host Type
|
||||||
|
|
||||||
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
||||||
@@ -243,6 +302,12 @@ done
|
|||||||
|
|
||||||
## Alert Templates
|
## Alert Templates
|
||||||
|
|
||||||
|
### HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)
|
||||||
|
```
|
||||||
|
⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
|
||||||
|
Action: {volume_type-specific action}
|
||||||
|
```
|
||||||
|
|
||||||
### AMBER (75-84%)
|
### AMBER (75-84%)
|
||||||
```
|
```
|
||||||
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
||||||
@@ -349,7 +414,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
|||||||
| 108 | media | storepve | lxc | ✅ reachable |
|
| 108 | media | storepve | lxc | ✅ reachable |
|
||||||
| 110 | gitea | minipve | lxc | ✅ reachable |
|
| 110 | gitea | minipve | lxc | ✅ reachable |
|
||||||
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
|
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
|
||||||
| 112 | tanko | amdpve | lxc | ✅ reachable |
|
| 112 | tanko | minipve | lxc | ✅ reachable |
|
||||||
| 113 | baggy | amdpve | lxc | ✅ reachable |
|
| 113 | baggy | amdpve | lxc | ✅ reachable |
|
||||||
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
|
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
|
||||||
| 116 | syslog-api | minipve | lxc | ✅ reachable |
|
| 116 | syslog-api | minipve | lxc | ✅ reachable |
|
||||||
|
|||||||
@@ -0,0 +1,169 @@
|
|||||||
|
# Contract execution pinning
|
||||||
|
|
||||||
|
Which copy of a contract script actually ran, and how that is proven.
|
||||||
|
|
||||||
|
## Why this exists
|
||||||
|
|
||||||
|
Three times on 2026-09-25 a contract reported a verdict from a copy that was
|
||||||
|
not the merged one:
|
||||||
|
|
||||||
|
1. The ops lane's own clone sat on the merged feature branch
|
||||||
|
`fix/search-stack-multi-engine-20260925` at `8b2eba4` with no `pve_auth`
|
||||||
|
fix, while it executed the daily digest from a different clone. Nothing in
|
||||||
|
the workflow noticed.
|
||||||
|
2. `scripts/search-stack-check.py` was deployed into the pinned runner clone
|
||||||
|
by hand rather than through git.
|
||||||
|
3. A stale local `origin/master` ref made an ancestry check report
|
||||||
|
"unlanded work" for a branch that had in fact merged — the same staleness
|
||||||
|
would have passed a stale script as current.
|
||||||
|
|
||||||
|
A contract verdict is only meaningful if it came from the merged copy. The
|
||||||
|
control is `scripts/revision-preflight.sh`.
|
||||||
|
|
||||||
|
## The rule
|
||||||
|
|
||||||
|
**Every contract pins exactly one clone for execution: the clone that
|
||||||
|
`scripts/contract-run.sh` itself lives in.**
|
||||||
|
|
||||||
|
`contract-run.sh` derives that from its own location (`SCRIPTS_DIR`) and checks
|
||||||
|
the script it is about to run against `origin/master` in the same clone. There
|
||||||
|
is no second path to configure, and no contract may be executed from a
|
||||||
|
hand-copied location.
|
||||||
|
|
||||||
|
| Contract | Script | Pinned clone |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `infrastructure-monitoring` | `scripts/infra-monitoring.sh` | the clone containing `contract-run.sh` |
|
||||||
|
| `proxmox-monitor` | `scripts/proxmox-monitor.sh` | same |
|
||||||
|
| `zulip-health` | `scripts/zulip-monitor.sh` | same |
|
||||||
|
| `agent-health-check` | `scripts/agent-health-check.py` | same |
|
||||||
|
| `litellm-health` | `scripts/litellm-health-check.py` | same |
|
||||||
|
| `disk-gc-threat-response` | `scripts/disk-gc-scan.py` | same |
|
||||||
|
| `pm2-self-heal` | `scripts/pm2-self-heal.sh` | same |
|
||||||
|
| `search-stack-visibility` | `scripts/search-stack-check.py` | same |
|
||||||
|
|
||||||
|
### The deployed runner
|
||||||
|
|
||||||
|
The scheduler on **CT 100 (abiba)** runs contracts from
|
||||||
|
**`/opt/contract-runner`** via `/etc/cron.d/contract-runner`. That clone is the
|
||||||
|
pinned execution copy for every scheduled contract, and it must be kept current
|
||||||
|
with `master` by fast-forward. Its `origin` is a local path to the upstream
|
||||||
|
working copy, not a network remote.
|
||||||
|
|
||||||
|
`daily-health-digest` is **not** in the table above because it has no contract
|
||||||
|
file and no mapping — it is dispatched by cron as
|
||||||
|
`fm-send.sh ops "run contract: daily-health-digest"` and was, until
|
||||||
|
2026-09-25, executed by hand from whichever clone the operator happened to be
|
||||||
|
in. Creating its contract file and pinning it to a clone is an open follow-up.
|
||||||
|
|
||||||
|
## How the check works
|
||||||
|
|
||||||
|
`scripts/revision-preflight.sh <script-path> <clone-path>`:
|
||||||
|
|
||||||
|
* resolves the **repo-relative** path of the executing script inside the clone;
|
||||||
|
* **fetches** the remote first, so a stale local ref cannot make a stale script
|
||||||
|
look current — bounded by `--fetch-timeout` (default 20s) so a hung remote
|
||||||
|
cannot block a scheduled contract;
|
||||||
|
* compares the script's sha256 against `<ref>:<repo-relative-path>`;
|
||||||
|
* **fails closed** — a path absent from the ref, an unresolvable ref, or a
|
||||||
|
failed fetch is a failure, never a warning.
|
||||||
|
|
||||||
|
### Exit codes and reason classes
|
||||||
|
|
||||||
|
The guard distinguishes **"I could not check"** from **"this copy is wrong"**,
|
||||||
|
and every non-zero exit prints a machine-readable `REASON=<class>` line before
|
||||||
|
the human text, because a warning nobody can classify is not actionable — and
|
||||||
|
the flip to `enforce` (below) depends on being able to read these apart.
|
||||||
|
|
||||||
|
| exit | `REASON=` | meaning |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| 0 | — | verified match |
|
||||||
|
| 2 | `cannot-verify:fetch-failed` | remote unreachable, failed, or timed out |
|
||||||
|
| 2 | `cannot-verify:ref-unresolvable` | `<ref>` does not exist in the clone |
|
||||||
|
| 1 | `mismatch:path-absent` | the script does not exist in `<ref>` |
|
||||||
|
| 1 | `mismatch:content` | the script differs from `<ref>` |
|
||||||
|
| 1 | `mismatch:detached-head` | the clone is on a detached HEAD |
|
||||||
|
| 1 | `mismatch:clone-ahead` | local HEAD is strictly ahead of `<ref>` (mid-review) |
|
||||||
|
|
||||||
|
`detached-head` and `clone-ahead` are named separately on purpose: they are
|
||||||
|
*legitimate* states that merely fail to be "the merged copy", and they are far
|
||||||
|
less alarming than a hand-edited file. `clone-ahead` requires HEAD to be
|
||||||
|
**strictly** ahead — an uncommitted edit on a commit that *is* the ref is a
|
||||||
|
plain `content` mismatch.
|
||||||
|
|
||||||
|
## Modes in `contract-run.sh`
|
||||||
|
|
||||||
|
| `CONTRACT_REVISION_PREFLIGHT` | Behaviour |
|
||||||
|
| --- | --- |
|
||||||
|
| unset / **`warn` (default)** | log the refusal and its class, then still report |
|
||||||
|
| `enforce` | withhold the verdict, alert, exit `2` |
|
||||||
|
| `off` | skip the check entirely |
|
||||||
|
|
||||||
|
**The default is `warn`, deliberately.** The guard gates *every* scheduled
|
||||||
|
contract, and three legitimate situations would otherwise turn the whole
|
||||||
|
fleet's monitoring into withheld verdicts: a clone legitimately ahead of
|
||||||
|
`origin/master` mid-review, a detached HEAD, and an offline or failed fetch.
|
||||||
|
That is a bigger risk than the staleness the guard exists to catch. `warn`
|
||||||
|
keeps the signal loud and classified in every run's log without letting the
|
||||||
|
monitoring go dark.
|
||||||
|
|
||||||
|
### Criteria for flipping the default to `enforce`
|
||||||
|
|
||||||
|
Do not flip it on preference. Flip it when the evidence says the false-refusal
|
||||||
|
rate is low enough, as its own small change with its own review:
|
||||||
|
|
||||||
|
1. the guard has run across **every scheduled contract** for a sustained period
|
||||||
|
(suggested: 30 consecutive days, or 200+ contract runs) with **zero**
|
||||||
|
`mismatch:*` and **zero** `cannot-verify:*` refusals in the per-run logs;
|
||||||
|
2. no `cannot-verify:fetch-failed` arising from ordinary network blips in that
|
||||||
|
window — if the pinned clone's remote is not reliably reachable, `enforce`
|
||||||
|
will withhold rather than report;
|
||||||
|
3. the pinned runner clone is demonstrably kept current by fast-forward, so
|
||||||
|
`mismatch:clone-ahead` is a genuine fault rather than routine procedure.
|
||||||
|
|
||||||
|
The evidence for the flip is the `REASON=` lines already written into
|
||||||
|
`/var/log/contract-runs/`. Until then the default stays `warn`.
|
||||||
|
|
||||||
|
## Merge-time sequence (do this whenever this repo merges)
|
||||||
|
|
||||||
|
**Baseline as of 2026-09-25:** `/opt/contract-runner` is already
|
||||||
|
fast-forwarded to master `9faffe4`, so the pinned runner clone is current
|
||||||
|
today. This sequence exists to keep it that way.
|
||||||
|
|
||||||
|
After any merge to `master`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. fast-forward the pinned runner clone on CT 100
|
||||||
|
git -C /opt/contract-runner pull --ff-only
|
||||||
|
|
||||||
|
# 2. confirm it is current and clean
|
||||||
|
git -C /opt/contract-runner log --oneline -1
|
||||||
|
git -C /opt/contract-runner status --porcelain # expect no output
|
||||||
|
|
||||||
|
# 3. prove a contract runs and reports normally
|
||||||
|
CONTRACT_RUN_LOG_DIR=/tmp/preflight-proof \
|
||||||
|
bash /opt/contract-runner/scripts/contract-run.sh search-stack-visibility
|
||||||
|
echo "EXIT=$?" # expect 0, and 'revision-preflight: … matches origin/master'
|
||||||
|
```
|
||||||
|
|
||||||
|
A contract that reports a `REASON=mismatch:*` refusal here means the runner
|
||||||
|
clone is stale or locally edited — fast-forward it rather than reaching for
|
||||||
|
`CONTRACT_REVISION_PREFLIGHT=off`.
|
||||||
|
|
||||||
|
**Note on untracked files:** git refuses to fast-forward over an untracked file
|
||||||
|
even when its content is byte-identical to the incoming version
|
||||||
|
(`The following untracked working tree files would be overwritten by merge`).
|
||||||
|
A dirty clone will therefore block step 1. Resolve it by removing or stashing
|
||||||
|
the untracked paths first — that is exactly what blocked a clone on 2026-09-25.
|
||||||
|
|
||||||
|
## Operating notes
|
||||||
|
|
||||||
|
* Under the default `warn`, a stale pinned clone still produces verdicts but
|
||||||
|
every run logs the refusal and its class. Read those lines; do not ignore
|
||||||
|
them.
|
||||||
|
* Under `enforce`, a stale pinned clone **withholds**. That is the intended
|
||||||
|
failure. Recover by fast-forwarding:
|
||||||
|
`git -C /opt/contract-runner pull --ff-only`.
|
||||||
|
* When a contract legitimately changes, land it through the normal branch + PR
|
||||||
|
path and fast-forward the pinned clone. Do not copy files into it by hand.
|
||||||
|
* `--no-fetch` exists for offline inspection; it prints that freshness is
|
||||||
|
assumed rather than verified, and it is not used by `contract-run.sh`.
|
||||||
@@ -1,5 +1,14 @@
|
|||||||
# Probe-drift round 2 — per-leg before/after evidence
|
# Probe-drift round 2 — per-leg before/after evidence
|
||||||
|
|
||||||
|
> **Historical record** — 2026-09-28: The lines below that describe tanko as
|
||||||
|
> "DSH (DeepSeek Harness)" only reflect what the check reported when it was
|
||||||
|
> running. Tanko's runtime was later found to be **hybrid (DSH + Hermes)** —
|
||||||
|
> the check had a `/root/` hardcoding bug that made it probe the wrong home
|
||||||
|
> directory and report `wrapper-missing:tanko` for an agent with a working
|
||||||
|
> wrapper. This document records the observed output, not the underlying
|
||||||
|
> truth; see `fix/agent-health-root-hardcoding-20260928` for the correction.
|
||||||
|
|
||||||
|
|
||||||
**Date:** 2026-09-10
|
**Date:** 2026-09-10
|
||||||
**Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`
|
**Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`
|
||||||
**Branch:** `fm/probe-drift-round2-20260909`
|
**Branch:** `fm/probe-drift-round2-20260909`
|
||||||
|
|||||||
@@ -174,6 +174,29 @@ Key notes:
|
|||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
|
### Executor, schedule and dead-man's-switch
|
||||||
|
|
||||||
|
This contract is implemented by a real script, scheduled on the inference host:
|
||||||
|
|
||||||
|
| | |
|
||||||
|
| --- | --- |
|
||||||
|
| **Executor** | `/opt/inference-harness/scripts/gpu-self-heal.py` on **CT 116** |
|
||||||
|
| **Schedule** | `/etc/cron.d/gpu-self-heal` on CT 116 — `2 */6 * * *` |
|
||||||
|
| **Log** | `/var/log/litellm/gpu-self-heal.log` |
|
||||||
|
| **Posting** | the script calls `gitea-logger.sh gpu {RUN_ID}.json <report>` → `SyslogSolution/health-logs/gpu/{RUN_ID}.json` |
|
||||||
|
|
||||||
|
**Dead-man's-switch:** absence of logs must raise an alarm, because that is how
|
||||||
|
this went silent for 12 days. That alarm cannot live on the producer — a stopped
|
||||||
|
job cannot report that it stopped — so it lives off-host as the
|
||||||
|
**`health-log-freshness` contract on CT 100**, which fails when
|
||||||
|
`health-logs/gpu/` is older than 12 h. A failed run is also visible in the log
|
||||||
|
above, but only the off-host check catches a *missing* run.
|
||||||
|
|
||||||
|
**History (2026-09-26, relay-785):** the executor was never lost — only its
|
||||||
|
schedule was, dropped during a CT 116 `/etc/cron.d` rework on 2026-09-21. The
|
||||||
|
`gpu/` log was silent from `2026-09-14T18:02:03Z`. The schedule was restored and
|
||||||
|
the off-host freshness check added; do not treat either as optional.
|
||||||
|
|
||||||
```prose
|
```prose
|
||||||
-- Phase 1: Fetch live GPU data
|
-- Phase 1: Fetch live GPU data
|
||||||
let fleet = call gpu-monitor
|
let fleet = call gpu-monitor
|
||||||
|
|||||||
@@ -0,0 +1,80 @@
|
|||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: health-log-freshness
|
||||||
|
description: >
|
||||||
|
Dead-man's-switch for the health-logs posting jobs. Fails when the newest
|
||||||
|
commit in a watched SyslogSolution/health-logs directory is older than that
|
||||||
|
directory's threshold.
|
||||||
|
|
||||||
|
Exists because absence of logs raises no alarm: the gpu/ directory went silent
|
||||||
|
for 12 days (2026-09-14T18:02:03Z to 2026-09-26) and nothing noticed, because
|
||||||
|
the only thing that would have noticed was the job that had stopped. The check
|
||||||
|
therefore runs on CT 100, a DIFFERENT host from the producers on CT 116, so a
|
||||||
|
dead producer host or a deleted schedule still raises an alarm.
|
||||||
|
|
||||||
|
Watched: gpu/ (12h, producer 2 */6 * * *), litellm/ (18h, producer 0 */6 * * *).
|
||||||
|
Not watched: pm2/ — that contract's health-logs claim was retired 2026-09-26;
|
||||||
|
see pm2-self-heal.prose.md step 6.
|
||||||
|
|
||||||
|
Verified 2026-09-26: would have caught the real gap (at 2026-09-20 the newest
|
||||||
|
gpu/ entry was 126h old against a 12h limit).
|
||||||
|
|
||||||
|
version: 1.0.0
|
||||||
|
---
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
Host-scheduled on CT 100 via `/etc/cron.d/contract-runner`:
|
||||||
|
|
||||||
|
```
|
||||||
|
20 */4 * * * root CONTRACT_RUN_LOG_DIR=/var/log/contract-runs /bin/bash /opt/contract-runner/scripts/contract-run.sh health-log-freshness >/dev/null 2>&1 || /root/abiba-workspace/bin/fm-inbox.sh note "contract-runner: health-log-freshness FAILED - see /var/log/contract-runs/" >/dev/null 2>&1
|
||||||
|
```
|
||||||
|
|
||||||
|
Every 4 hours, offset to `:20` to avoid the existing `:05`/`:15`/`:35` slots.
|
||||||
|
|
||||||
|
## Output shape
|
||||||
|
|
||||||
|
```
|
||||||
|
Health-log freshness — dead-man's-switch
|
||||||
|
==================================================================
|
||||||
|
✅ health-logs/gpu/ newest 2026-09-26T14:48:03Z (0.02h old, limit 12.0h)
|
||||||
|
last commit: gpu: gpu-self-heal-20260926-144802.json
|
||||||
|
producer: gpu-self-heal.py, CT116 cron 2 */6 * * *
|
||||||
|
==================================================================
|
||||||
|
VERDICT: PASS — every watched health-log directory is advancing
|
||||||
|
```
|
||||||
|
|
||||||
|
`--json` emits `{checked: {...}, failures: [...]}`.
|
||||||
|
|
||||||
|
## Exit codes
|
||||||
|
|
||||||
|
| exit | meaning |
|
||||||
|
| --- | --- |
|
||||||
|
| 0 | every watched directory is advancing |
|
||||||
|
| 1 | at least one is stale, or could not be read |
|
||||||
|
| 2 | the check could not run (no Gitea credential) |
|
||||||
|
|
||||||
|
A directory that **cannot be read** is a failure, not a skip: unreadable and
|
||||||
|
stopped are indistinguishable from the outside.
|
||||||
|
|
||||||
|
## Configuration
|
||||||
|
|
||||||
|
| variable | default | meaning |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `GITEA_URL` | `https://git.sysloggh.net` | Gitea base URL |
|
||||||
|
| `GITEA_TOKEN` / `GITEA_PAT` | — | API token; falls back to basic auth from `~/.git-credentials` |
|
||||||
|
| `HEALTH_LOG_MAX_AGE_GPU` | `12` | hours |
|
||||||
|
| `HEALTH_LOG_MAX_AGE_LITELLM` | `18` | hours |
|
||||||
|
|
||||||
|
## Thresholds
|
||||||
|
|
||||||
|
Sized for the producer cadence plus one missed run, so a single blip does not
|
||||||
|
page but a genuine stop does:
|
||||||
|
|
||||||
|
* `gpu/` — 6 h cadence, 12 h limit;
|
||||||
|
* `litellm/` — 6 h cadence, 18 h limit (proven healthy; a looser bound avoids noise).
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- health-logs-gpu-freshness: { status: "ok|stale", last_check: timestamp }
|
||||||
|
- health-logs-litellm-freshness: { status: "ok|stale", last_check: timestamp }
|
||||||
@@ -70,10 +70,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
|
|||||||
|
|
||||||
## Config Pattern — Mandatory Fields
|
## Config Pattern — Mandatory Fields
|
||||||
|
|
||||||
### For Hermes Agents (Mumuni, Koonimo)
|
### For Hermes Agents (Mumuni, Koonimo, Tanko-hybrid)
|
||||||
|
|
||||||
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||||
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
|
(Tanko is hybrid — runs both DSH and Hermes since 2026-08-27, so its Hermes config is also checked.)
|
||||||
|
|
||||||
### 1. Main Model
|
### 1. Main Model
|
||||||
```yaml
|
```yaml
|
||||||
@@ -297,7 +297,7 @@ Run the consolidated health check:
|
|||||||
```bash
|
```bash
|
||||||
python3 /root/scripts/agent-health-check.py
|
python3 /root/scripts/agent-health-check.py
|
||||||
```
|
```
|
||||||
This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes),
|
This validates each agent's live LiteLLM key against the gateway, including tanko, which runs HYBRID (DSH + Hermes) since 2026-08-27; detects GPU port conflicts (ghost processes),
|
||||||
verifies gateway liveness, confirms Zulip streaming (`edit_message` present),
|
verifies gateway liveness, confirms Zulip streaming (`edit_message` present),
|
||||||
and counts recent errors. Non-disruptive — never restarts anything.
|
and counts recent errors. Non-disruptive — never restarts anything.
|
||||||
|
|
||||||
|
|||||||
@@ -5,7 +5,13 @@ description: >
|
|||||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||||
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
|
UPDATED 2026-09-27: Clarified the Auxiliary Tasks policy — light aux (vision,
|
||||||
|
web_extract/browsing) -> gpu-vision (RTX 5070); context-heavy aux (compression) ->
|
||||||
|
syslog-auto (2026-07-23 decision, Rule 7). Removed the false "one model for all
|
||||||
|
auxiliary" / "never syslog-auto" claim; stated gpu-dense + strix-moe are the reasoning
|
||||||
|
hosts and aux should not be pinned to them. Now matches audit-hermes-config.py line-for-line.
|
||||||
|
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
|
||||||
|
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
|
||||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||||
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
||||||
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
||||||
@@ -145,6 +151,12 @@ mcp_servers:
|
|||||||
url: http://192.168.68.65:3100/mcp
|
url: http://192.168.68.65:3100/mcp
|
||||||
timeout: 120
|
timeout: 120
|
||||||
connect_timeout: 60
|
connect_timeout: 60
|
||||||
|
litellm:
|
||||||
|
url: https://litellm.sysloggh.net/mcp
|
||||||
|
headers:
|
||||||
|
x-litellm-api-key: "Bearer <AGENT_KEY>" # Rule 15: must be a REAL key (sk-...), not an env-var name
|
||||||
|
# Note: MCP endpoint requires Accept: application/json, text/event-stream header
|
||||||
|
# This is handled by the MCP client library; don't add to config
|
||||||
|
|
||||||
# ─── Compression ───
|
# ─── Compression ───
|
||||||
compression:
|
compression:
|
||||||
@@ -160,13 +172,16 @@ compression:
|
|||||||
abort_on_summary_failure: false
|
abort_on_summary_failure: false
|
||||||
|
|
||||||
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
||||||
# All auxiliary services MUST use identical model, base_url, and api_key_env:
|
# Auxiliary tasks split into TWO model classes — do NOT assume one model for all:
|
||||||
# model: gpu-vision # stable alias (NOT a raw model name)
|
# Light auxiliary (vision, web_extract/browsing) -> model: gpu-vision # RTX 5070
|
||||||
|
# Keeps the reasoning hosts (gpu-dense / strix-moe) free for agent prompts.
|
||||||
|
# Context-heavy auxiliary (compression) -> model: syslog-auto # weighted pool
|
||||||
|
# Deliberate per the 2026-07-23 OPERATIONAL DECISION in Rule 7: summarization
|
||||||
|
# runs against long histories and must be able to use the pool.
|
||||||
|
# Do NOT pin auxiliary work to the reasoning hosts (gpu-dense / strix-moe).
|
||||||
|
# All auxiliary services share identical ROUTING (base_url + api_key_env), not model:
|
||||||
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||||
# api_key_env: LITELLM_API_KEY
|
# api_key_env: LITELLM_API_KEY
|
||||||
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
|
||||||
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
|
||||||
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
|
|
||||||
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
||||||
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
|
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
|
||||||
auxiliary:
|
auxiliary:
|
||||||
@@ -217,6 +232,40 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
|||||||
3. **After update**: Restart Hermes on the agent host
|
3. **After update**: Restart Hermes on the agent host
|
||||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||||
|
|
||||||
|
## MCP Server Configuration
|
||||||
|
|
||||||
|
MCP server entries in `mcp_servers:` must follow the format shown in the Template section.
|
||||||
|
|
||||||
|
**Header requirements (Rule 15):**
|
||||||
|
- Use `headers:` field with a `x-litellm-api-key` entry
|
||||||
|
- The value must be `"Bearer <REAL_KEY>"` where `<REAL_KEY>` is a literal LiteLLM virtual key
|
||||||
|
- Do NOT use env-var references like `$LITELLM_API_KEY` — they resolve to empty strings in
|
||||||
|
the static config and cause "Malformed API Key" errors (2026-08-07 Tanko incident)
|
||||||
|
|
||||||
|
**Key source:**
|
||||||
|
- Keys are stored in the Infisical vault (project=agents, env=production)
|
||||||
|
- For template-based config generation: substitute the agent's key from the agent_keys table
|
||||||
|
- For manual config updates: retrieve the key from the vault and insert the literal value
|
||||||
|
|
||||||
|
**Verification (2026-08-07):**
|
||||||
|
- Tested MCP initialize handshake against litellm.sysloggh.net/mcp with agent virtual key
|
||||||
|
- Confirmed: 200 response with `serverInfo.name: "litellm-mcp-server"`
|
||||||
|
- Confirmed: tools/list returns 200 (MCP endpoint accessible with virtual keys)
|
||||||
|
- Note: Per-key MCP grants are now supported (verified 2026-09-18), resolving the earlier
|
||||||
|
contradiction with infrastructure-update.prose.md (which now reflects the update)
|
||||||
|
- Key requirement: must be a valid LiteLLM virtual key (HTTP 200 on /v1/models)
|
||||||
|
|
||||||
|
**Key rotation note:**
|
||||||
|
- MCP headers use literal keys (not env-vars), so they do NOT auto-rotate with the vault
|
||||||
|
- After key rotation, MCP server headers must be regenerated with the new key value
|
||||||
|
- This is a manual step: update the `x-litellm-api-key` header in each config file
|
||||||
|
- TODO: Consider adding MCP header regeneration to the Key Update Procedure or a generation hook
|
||||||
|
|
||||||
|
**NetBird dependency:**
|
||||||
|
- `litellm.sysloggh.net` is a NetBird endpoint (see Rule 5 and Infrastructure Stack table)
|
||||||
|
- NetBird outages cause 502 errors on MCP requests, not auth failures
|
||||||
|
- Diagnose: if MCP requests fail with 502, check NetBird status before investigating keys
|
||||||
|
|
||||||
## Violation Classification
|
## Violation Classification
|
||||||
|
|
||||||
When reporting findings, separate POLICY observations from FAULT findings:
|
When reporting findings, separate POLICY observations from FAULT findings:
|
||||||
@@ -455,14 +504,28 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
|||||||
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
||||||
before and after any config change to catch this and all other rule violations.
|
before and after any config change to catch this and all other rule violations.
|
||||||
|
|
||||||
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
|
### Rule 15: MCP Endpoint and Header Validation (UPDATED 2026-08-07)
|
||||||
- Every MCP server entry must point at the correct endpoint:
|
|
||||||
- ra-h-os = http://192.168.68.65:3100/mcp
|
**Endpoint validation:**
|
||||||
- litellm = https://litellm.sysloggh.net/mcp
|
- ra-h-os must point to `http://192.168.68.65:3100/mcp`
|
||||||
- MCP entries must carry a REAL key value in the header.
|
- litellm must point to `https://litellm.sysloggh.net/mcp`
|
||||||
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
|
- Mismatched endpoints cause silent failures (e.g., 2026-08-07 incident: Tanko's config had
|
||||||
endpoints and result in "Malformed API Key" floods.
|
ra-h-os pointing to litellm's endpoint)
|
||||||
- Ensure the header value is the actual key (e.g., `sk-...`).
|
|
||||||
|
**Header validation:**
|
||||||
|
- Every MCP entry with authentication must carry a `headers:` field
|
||||||
|
- The header value must be a REAL key (e.g., `Bearer sk-abc123...`), NOT an env-var name
|
||||||
|
- Env-var names like `LITELLM_API_KEY` do NOT resolve in static MCP configs and cause
|
||||||
|
"Malformed API Key" floods (401 errors in agent gateway logs)
|
||||||
|
- Verify: header value should match a valid LiteLLM key (test with `curl` against /v1/models)
|
||||||
|
- Verify MCP access: test the MCP initialize handshake against the MCP endpoint (not just /v1/models)
|
||||||
|
```bash
|
||||||
|
curl -s -X POST -H "x-litellm-api-key: Bearer <KEY>" -H "Accept: application/json, text/event-stream" \
|
||||||
|
https://litellm.sysloggh.net/mcp -d '{"jsonrpc":"2.0","id":1,"method":"initialize",...}' \
|
||||||
|
| jq '.data.result.serverInfo' # should show serverInfo.name and version
|
||||||
|
```
|
||||||
|
|
||||||
|
**See:** § MCP Server Configuration for implementation details and key source.
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
|
|||||||
@@ -16,7 +16,26 @@ author: Abiba (pi agent)
|
|||||||
|
|
||||||
## Rule (One Sentence)
|
## Rule (One Sentence)
|
||||||
|
|
||||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with authenticated path `http://192.168.68.116/litellm/v1/responses` — hardcoded keys AND unauthenticated `/v1` direct access are both forbidden.**
|
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
|
||||||
|
|
||||||
|
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
|
||||||
|
|
||||||
|
## Model Access Tiers (2026-09-20)
|
||||||
|
|
||||||
|
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
|
||||||
|
|
||||||
|
| Tier | Models | Who gets it |
|
||||||
|
|------|--------|-------------|
|
||||||
|
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
|
||||||
|
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
|
||||||
|
|
||||||
|
**Rules:**
|
||||||
|
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
|
||||||
|
2. Agent keys MUST carry an **explicit local-only** `models` list.
|
||||||
|
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
|
||||||
|
4. The master key always bypasses scoping — it is admin-only, never for inference.
|
||||||
|
|
||||||
|
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
|
||||||
|
|
||||||
## Scope
|
## Scope
|
||||||
|
|
||||||
@@ -38,8 +57,8 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
|
|||||||
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
|
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
|
||||||
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
|
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
|
||||||
|
|
||||||
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid).
|
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` non-canonical but working).
|
||||||
Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**.
|
The public host `https://litellm.sysloggh.net` serves `/v1` ONLY (404 on `/litellm/v1`).
|
||||||
|
|
||||||
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
|
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
|
||||||
|
|
||||||
@@ -101,7 +120,7 @@ auxiliary:
|
|||||||
fallback_providers:
|
fallback_providers:
|
||||||
- provider: deepseek
|
- provider: deepseek
|
||||||
base_url: https://api.deepseek.com
|
base_url: https://api.deepseek.com
|
||||||
api_key: sk-b7d9... # ← hardcoded OK (external)
|
api_key: sk-synthetic-external-example # ← hardcoded OK (external, synthetic example)
|
||||||
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
|
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -109,11 +128,11 @@ fallback_providers:
|
|||||||
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
|
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
|
||||||
model:
|
model:
|
||||||
provider: harness
|
provider: harness
|
||||||
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION: hardcoded key
|
api_key: sk-synthetic-example-12345 # ← RULE VIOLATION: hardcoded key (synthetic example)
|
||||||
|
|
||||||
model:
|
model:
|
||||||
provider: harness
|
provider: harness
|
||||||
base_url: http://192.168.68.116/v1 # ← RULE VIOLATION: unauthenticated path
|
base_url: http://192.168.68.116/v1 # ← NON-CANONICAL but WORKING (authenticated via nginx, WARN not FAIL)
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -139,6 +158,16 @@ scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
|
|||||||
|
|
||||||
When reporting findings, separate POLICY observations from FAULT findings:
|
When reporting findings, separate POLICY observations from FAULT findings:
|
||||||
|
|
||||||
|
### ACCEPTABLE PATTERN
|
||||||
|
Agent keys live in `.env` or `.env.vault` files with 600 permissions (koonimo's shape is the canonical example). A plaintext key inside a `config.yaml` or any `config.yaml.bak-*` file is a violation — the backup files are not part of the runtime credential path and are not watched by the scanner, so a key in them is stale clutter that a future reader can mistake for a working key.
|
||||||
|
|
||||||
|
**Fix procedure** (when a backup file is found with a plaintext key):
|
||||||
|
1. Move the file out of the scanned tree (e.g., `mv /root/.hermes/config.yaml.bak-* /root/hermes-config-backups/`) — do NOT delete the file, just move it so the scanner pattern no longer matches.
|
||||||
|
2. Re-run the reachability check to confirm COMPLIANT.
|
||||||
|
3. Report the before/after check output and the commands you ran.
|
||||||
|
|
||||||
|
**Rationale**: Moving the file preserves history without leaving a credential where a scanner trips over it. Deleting the file loses the historical context. Keeping it in place means the next scan will report it as a finding and waste time re-deciding.
|
||||||
|
|
||||||
### POLICY (observation only, not a fault)
|
### POLICY (observation only, not a fault)
|
||||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||||
@@ -172,7 +201,7 @@ grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
|
|||||||
|
|
||||||
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
||||||
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
|
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
|
||||||
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /root/.config/systemd/ 2>/dev/null
|
grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-…' /root/.config/systemd/ 2>/dev/null
|
||||||
|
|
||||||
# 3. Verify running process env matches dedicated key
|
# 3. Verify running process env matches dedicated key
|
||||||
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
||||||
|
|||||||
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
|
|
||||||
| Param | Type | Required | Default | Description |
|
| Param | Type | Required | Default | Description |
|
||||||
|-------|------|----------|---------|-------------|
|
|-------|------|----------|---------|-------------|
|
||||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) |
|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — hybrid (DSH + Hermes) since 2026-08-27, no Hermes plugin) |
|
||||||
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -55,7 +55,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
|
|
||||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||||
|------|-----|---------|-------------|-------------|------|
|
|------|-----|---------|-------------|-------------|------|
|
||||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
|
| Tanko | CT112 | minipve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(hybrid (DSH + Hermes) since 2026-08-27 — historical, plugin retired on this host)* |
|
||||||
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
|
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
|
||||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||||
|
|
||||||
@@ -72,7 +72,7 @@ connectivity recovery including end-to-end DM validation.
|
|||||||
### Step 1: Resolve Target
|
### Step 1: Resolve Target
|
||||||
|
|
||||||
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
|
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
|
||||||
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
For CT112 route through `ssh root@minipve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||||
|
|
||||||
### Step 2: Pull Latest Plugin Source
|
### Step 2: Pull Latest Plugin Source
|
||||||
|
|
||||||
@@ -121,7 +121,7 @@ cp plugins/platforms/zulip/adapter.py \
|
|||||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
# Fix ownership (was Tanko-only, runs as jerome user)
|
# Fix ownership (was Tanko-only, runs as jerome user)
|
||||||
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
|
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (hybrid: DSH + Hermes).
|
||||||
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
||||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
|
|||||||
@@ -24,7 +24,7 @@ gateway restart, and connection validation.
|
|||||||
|
|
||||||
| Param | Type | Required | Default | Description |
|
| Param | Type | Required | Default | Description |
|
||||||
|-------|------|----------|---------|-------------|
|
|-------|------|----------|---------|-------------|
|
||||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) |
|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — hybrid (DSH + Hermes) since 2026-08-27) |
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
@@ -67,7 +67,7 @@ gateway restart, and connection validation.
|
|||||||
### Step 1: Locate Target
|
### Step 1: Locate Target
|
||||||
|
|
||||||
Map `target` to connectivity parameters from the live-state table above.
|
Map `target` to connectivity parameters from the live-state table above.
|
||||||
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
For CT112 route through `ssh root@minipve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||||
|
|
||||||
### Step 2: Deploy Zulip Adapter
|
### Step 2: Deploy Zulip Adapter
|
||||||
|
|
||||||
@@ -93,7 +93,7 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
|
|||||||
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
||||||
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
||||||
|
|
||||||
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin)
|
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on hybrid (DSH + Hermes), no Hermes plugin)
|
||||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
|
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
|
||||||
|
|
||||||
# Clean up
|
# Clean up
|
||||||
|
|||||||
@@ -105,8 +105,8 @@ description: >
|
|||||||
|
|
||||||
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
||||||
|------|----|-----|-----|---------|------|
|
|------|----|-----|-----|---------|------|
|
||||||
| minipve | .12 | 16C | 30GB | abiba, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
| minipve | .12 | 16C | 30GB | abiba, tanko, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||||
| amdpve | .15 | 32C | 62GB | kagentz, tanko, baggy, scottdenya, adguard2 | Agents, compute |
|
| amdpve | .15 | 32C | 62GB | kagentz, baggy, scottdenya, adguard2 | Agents, compute |
|
||||||
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
|
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
|
||||||
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
|
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
|
||||||
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
||||||
@@ -181,8 +181,8 @@ description: >
|
|||||||
**Stirling-PDF** (deployed 2026-07-03, Authentik SSO 2026-07-03):
|
**Stirling-PDF** (deployed 2026-07-03, Authentik SSO 2026-07-03):
|
||||||
- URL: `https://pdf.sysloggh.net` (public) / `http://192.168.68.7:8989` (direct)
|
- URL: `https://pdf.sysloggh.net` (public) / `http://192.168.68.7:8989` (direct)
|
||||||
- Swagger: `http://192.168.68.7:8989/swagger-ui.html`
|
- Swagger: `http://192.168.68.7:8989/swagger-ui.html`
|
||||||
- Admin credentials: `admin` / `kakashi20stirling`
|
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_USER»` / `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||||
- API key: `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88`
|
- API key: `«vault: infrastructure/production STIRLING_API_KEY»`
|
||||||
- Authentik OAuth2: configured but disabled (requires paid Server license). Ready to enable: set `SECURITY_OAUTH2_ENABLED=true` + `SECURITY_LOGINMETHOD=all`
|
- Authentik OAuth2: configured but disabled (requires paid Server license). Ready to enable: set `SECURITY_OAUTH2_ENABLED=true` + `SECURITY_LOGINMETHOD=all`
|
||||||
- Compose: `/opt/home_stack/docker-compose.yml`
|
- Compose: `/opt/home_stack/docker-compose.yml`
|
||||||
- Control script: `/opt/home_stack/infra-control.sh`
|
- Control script: `/opt/home_stack/infra-control.sh`
|
||||||
@@ -636,7 +636,7 @@ monitor, or integration breaks.
|
|||||||
```bash
|
```bash
|
||||||
# Full cluster status
|
# Full cluster status
|
||||||
PVE="https://minipve.sysloggh.net"
|
PVE="https://minipve.sysloggh.net"
|
||||||
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
AUTH="Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
|
||||||
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
|
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
|
||||||
|
|
||||||
# Docker health from Abiba
|
# Docker health from Abiba
|
||||||
@@ -682,7 +682,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
|||||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||||
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
||||||
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
|
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
|
||||||
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
|
| 112 | tanko | minipve | .122 | hybrid (DSH + Hermes) agent | ✅ |
|
||||||
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
||||||
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
||||||
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
||||||
@@ -712,7 +712,7 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
|
|||||||
| 100 | abiba | minipve | `pct-run 100` |
|
| 100 | abiba | minipve | `pct-run 100` |
|
||||||
| 105 | kagentz | amdpve | `pct-run 105` |
|
| 105 | kagentz | amdpve | `pct-run 105` |
|
||||||
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
|
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
|
||||||
| 112 | tanko | amdpve | `pct-run 112` |
|
| 112 | tanko | minipve | `pct-run 112` |
|
||||||
| 113 | baggy | amdpve | `pct-run 113` |
|
| 113 | baggy | amdpve | `pct-run 113` |
|
||||||
| 115 | scottdenya | amdpve | `pct-run 115` |
|
| 115 | scottdenya | amdpve | `pct-run 115` |
|
||||||
| 104 | authentik | minipve | `pct-run 104` |
|
| 104 | authentik | minipve | `pct-run 104` |
|
||||||
@@ -759,4 +759,4 @@ which was kill+nohup outside systemd) are banned by policy.
|
|||||||
| Script | Why Disabled |
|
| Script | Why Disabled |
|
||||||
|--------|-------------|
|
|--------|-------------|
|
||||||
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
|
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
|
||||||
| `zulip-monitor.sh` (Abiba) | Replaced by agent-health-check.py + PM2 auto-restart |
|
| `zulip-monitor.sh` (Abiba) | Standalone cron replaced by agent-health-check.py + PM2 auto-restart; the script itself remains active as the execution step of the `zulip-health` responsibility contract (see `zulip-health.prose.md`) |
|
||||||
|
|||||||
@@ -102,6 +102,13 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
|||||||
- Stack persists across reboots (systemd for exporters, Docker restart policy)
|
- Stack persists across reboots (systemd for exporters, Docker restart policy)
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||||
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||||
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||||
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||||
|
execution — it only proves the agent read the result and reported it. The actual
|
||||||
|
monitoring work happens in the host cron job.
|
||||||
|
|
||||||
|
|
||||||
### Liveness rule (scoped)
|
### Liveness rule (scoped)
|
||||||
|
|
||||||
@@ -138,11 +145,19 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
|||||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
||||||
tool calls; never repeat a prior report unless a live probe fails.**
|
tool calls; never repeat a prior report unless a live probe fails.**
|
||||||
|
|
||||||
|
**EXECUTABLE OWNER:** The canonical probe set lives in `scripts/infra-monitoring.sh`.
|
||||||
|
A check run is a single command: `bash scripts/infra-monitoring.sh` (from the
|
||||||
|
repository root). Paste its raw output verbatim into the report. The script
|
||||||
|
exits non-zero naming every failed target; there is no "OK" summary when any
|
||||||
|
leg failed. Port drift is caught by `scripts/test_infra_monitoring.sh` which
|
||||||
|
asserts every probed port matches the documented value.
|
||||||
|
|
||||||
**PROBE SHAPE (per standing rules above):**
|
**PROBE SHAPE (per standing rules above):**
|
||||||
- Every probe prints the target name + URL + HTTP code (or failure kind)
|
- Every probe prints: `✅ <name>: alive` on success, or `🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)` on failure
|
||||||
- Retry once on connection failure at longer timeout
|
- PVE API failures include `(any-HTTP liveness, -k for self-signed)` to distinguish TLS vs connection
|
||||||
|
- Retry once on connection failure at longer timeout (25s connect, 30s max)
|
||||||
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
||||||
- Report the actual probe command and its result, not a summary verdict
|
- Report the actual probe output, not a summary verdict
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Provenance — run first; paste the absolute path into the report
|
# Provenance — run first; paste the absolute path into the report
|
||||||
@@ -269,6 +284,27 @@ code (or failure kind with retry details). Apply the standing probe rules: any
|
|||||||
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
||||||
a failure.
|
a failure.
|
||||||
|
|
||||||
|
### Docker Stats and PVE Exporter Ports
|
||||||
|
|
||||||
|
These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:
|
||||||
|
|
||||||
|
| Exporter | Port | Container | Metrics |
|
||||||
|
|----------|------|-----------|---------|
|
||||||
|
| **Docker Stats** | **9324** | harness-docker-stats | `docker_container_*` (per-container CPU/mem/network) |
|
||||||
|
| **PVE Exporter** | **9221** | harness-pve-exporter | `pve_*` (5 cluster-level metrics) |
|
||||||
|
|
||||||
|
**IMPORTANT**: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (`builder_builds_*`, `containerd_build_info_*`).
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Docker Stats (harness-docker-stats)
|
||||||
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
|
||||||
|
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||||
|
|
||||||
|
# PVE Exporter (harness-pve-exporter)
|
||||||
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
|
||||||
|
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||||
|
```
|
||||||
|
|
||||||
### Phase 1: GPU Exporters
|
### Phase 1: GPU Exporters
|
||||||
|
|
||||||
**NVIDIA (.8 and .110)**:
|
**NVIDIA (.8 and .110)**:
|
||||||
|
|||||||
@@ -59,7 +59,7 @@ Before ANY update wave:
|
|||||||
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||||
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
| CT 112 (tanko, minipve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||||
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
||||||
@@ -208,19 +208,19 @@ mcp_servers:
|
|||||||
| Key | MCP Access |
|
| Key | MCP Access |
|
||||||
|-----|-----------|
|
|-----|-----------|
|
||||||
| Master key | ✅ Full — 90 tools (vault-injected) |
|
| Master key | ✅ Full — 90 tools (vault-injected) |
|
||||||
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.99.1 |
|
| Agent keys (mumuni, tanko, etc.) | ✅ Per-key grants supported (as of 2026-09-18 verification) |
|
||||||
|
|
||||||
### Known Limitations
|
### Known Limitations
|
||||||
- Per-key MCP server grants not functional — only master key has access
|
- ~~Per-key MCP server grants not functional — only master key has access~~ (resolved 2026-09-18: per-key grants now work)
|
||||||
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
|
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
|
||||||
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
|
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
|
||||||
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
|
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
|
||||||
|
|
||||||
### Migration Path
|
### Migration Path (COMPLETED 2026-09-18)
|
||||||
When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
Per-key MCP grants are now supported:
|
||||||
1. Grant agent keys `mcp_servers: ["ra_h_os"]`
|
1. ✅ Agent keys granted MCP access via `allowed_mcp_servers` field
|
||||||
2. Update Hermes `mcp_servers.ra-h-os.url` from `http://192.168.68.65:3100/mcp` → `http://192.168.68.116:4000/mcp/`
|
2. ✅ Hermes `mcp_servers.litellm.url` set to `https://litellm.sysloggh.net/mcp`
|
||||||
3. Add `headers: {x-litellm-api-key: "Bearer $LITELLM_API_KEY"}` to MCP config
|
3. ✅ `headers: {x-litellm-api-key: "Bearer <literal_key>"}` added to MCP config
|
||||||
|
|
||||||
## Security-Specific Updates
|
## Security-Specific Updates
|
||||||
|
|
||||||
|
|||||||
@@ -71,6 +71,10 @@ description: >
|
|||||||
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
||||||
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
||||||
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
||||||
|
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
|
||||||
|
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
|
||||||
|
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
|
||||||
|
In LiteLLM Community both silently grant access to EVERY model, including cloud.
|
||||||
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
||||||
- Return the new key
|
- Return the new key
|
||||||
5. **If action == "rotate"**:
|
5. **If action == "rotate"**:
|
||||||
@@ -87,6 +91,66 @@ description: >
|
|||||||
- Confirm key alias matches agent_name in LiteLLM key list
|
- Confirm key alias matches agent_name in LiteLLM key list
|
||||||
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
||||||
|
|
||||||
|
## Cloud Provider Consolidation (2026-09-20)
|
||||||
|
|
||||||
|
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
|
||||||
|
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
|
||||||
|
|
||||||
|
### Provider map (per-account namespacing)
|
||||||
|
|
||||||
|
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
|
||||||
|
stay separate:
|
||||||
|
|
||||||
|
| Prefix | Upstream | Auth | Vault secret |
|
||||||
|
|--------|----------|------|--------------|
|
||||||
|
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
|
||||||
|
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
|
||||||
|
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
|
||||||
|
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
|
||||||
|
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
|
||||||
|
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
|
||||||
|
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
|
||||||
|
|
||||||
|
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
|
||||||
|
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
|
||||||
|
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
|
||||||
|
|
||||||
|
### Access tiers (MUST be enforced per key)
|
||||||
|
|
||||||
|
| Tier | Model names | Granted to |
|
||||||
|
|------|-------------|------------|
|
||||||
|
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
|
||||||
|
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
|
||||||
|
|
||||||
|
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
|
||||||
|
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
|
||||||
|
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
|
||||||
|
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
|
||||||
|
|
||||||
|
### Creating the cloud-enabled key
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# ALWAYS read the live roster first (key-scoped):
|
||||||
|
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
|
||||||
|
| jq -r '.data[].id'
|
||||||
|
|
||||||
|
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
|
||||||
|
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
|
||||||
|
```
|
||||||
|
|
||||||
|
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
|
||||||
|
any `<prefix>/` cloud model.
|
||||||
|
|
||||||
|
### Adding a new cloud provider
|
||||||
|
|
||||||
|
1. Add the upstream key to Infisical `infrastructure/production/root`.
|
||||||
|
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
|
||||||
|
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
|
||||||
|
3. Restart the `harness-litellm` container.
|
||||||
|
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
|
||||||
|
captain approval.
|
||||||
|
5. Update this table and the access-tier section.
|
||||||
|
|
||||||
## Production Vault Access Process (canonical, 2026-07-17)
|
## Production Vault Access Process (canonical, 2026-07-17)
|
||||||
|
|
||||||
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
||||||
@@ -144,8 +208,8 @@ through its agent wrapper.
|
|||||||
safety net for vault outage or token revocation. Must be kept in sync on rotation.
|
safety net for vault outage or token revocation. Must be kept in sync on rotation.
|
||||||
Example:
|
Example:
|
||||||
```bash
|
```bash
|
||||||
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
|
MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY»
|
||||||
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
|
MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY»
|
||||||
```
|
```
|
||||||
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
|
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
|
||||||
```ini
|
```ini
|
||||||
@@ -191,7 +255,7 @@ through its agent wrapper.
|
|||||||
### Tanko migration (COMPLETED 2026-07-17)
|
### Tanko migration (COMPLETED 2026-07-17)
|
||||||
|
|
||||||
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
|
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
|
||||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
|
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-synthetic-tanko-example…`)
|
||||||
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
|
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
|
||||||
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
|
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
|
||||||
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
|
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
|
||||||
@@ -271,13 +335,13 @@ not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update ma
|
|||||||
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
||||||
|
|
||||||
**Key Storage:**
|
**Key Storage:**
|
||||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
|
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»…`)
|
||||||
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
||||||
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
||||||
(unlike fleet agents which require vault injection)
|
(unlike fleet agents which require vault injection)
|
||||||
|
|
||||||
**Current Key (2026-09-01):**
|
**Current Key (2026-09-01):**
|
||||||
- **Prefix**: `sk-or-v1-0af3f3…`
|
- **Prefix**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||||
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||||
- **Plan**: Paid (not free tier)
|
- **Plan**: Paid (not free tier)
|
||||||
- **Usage**: 0 (as of 2026-09-01)
|
- **Usage**: 0 (as of 2026-09-01)
|
||||||
|
|||||||
@@ -19,7 +19,7 @@ description: >
|
|||||||
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||||
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||||
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||||
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
|
- Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).
|
||||||
---
|
---
|
||||||
|
|
||||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||||
@@ -111,6 +111,13 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
|||||||
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
|
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||||
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||||
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||||
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||||
|
execution — it only proves the agent read the result and reported it. The actual
|
||||||
|
monitoring work happens in the host cron job.
|
||||||
|
|
||||||
|
|
||||||
1. **Read parameters** — Use provided values or defaults
|
1. **Read parameters** — Use provided values or defaults
|
||||||
|
|
||||||
|
|||||||
@@ -116,7 +116,7 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
|||||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (hardcoded key → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM (deprecated key, no live usage).
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
|
|
||||||
|
|||||||
+44
-6
@@ -6,13 +6,19 @@ name: memory-fixer
|
|||||||
description: >
|
description: >
|
||||||
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
||||||
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||||
version: 2.1.0
|
version: 2.2.0
|
||||||
---
|
---
|
||||||
---
|
---
|
||||||
|
|
||||||
# Memory Fixer
|
# Memory Fixer
|
||||||
|
|
||||||
> **Canonical copy:** `/root/.hermes/contracts/memory-fixer-v3.md` (used by the `memory-fixer-daily` cron job). This file is the institutional record of the same contract. When the two diverge, treat the v3 source in `/root/.hermes/contracts/` as executable truth.
|
> **Executable copy:** the `okyeame-memory-fixer` cron job on kagentz (`hermes cron list`) holds its instruction
|
||||||
|
> set **inline in `~/.hermes/cron/jobs.json`** (`hermes cron edit <id> --prompt …`; there is no `--prompt-file`, and
|
||||||
|
> `~/.hermes/cron/memory-fixer-prompt.md` is a synced draft, not the live instruction). This file is the institutional
|
||||||
|
> record of the same contract; when the two diverge, the job prompt is what actually runs — diff it against this file
|
||||||
|
> before claiming a prompt change landed.
|
||||||
|
> ⚠️ Corrected 2026-09-26: the previous pointer (`/root/.hermes/contracts/memory-fixer-v3.md`) does not exist on
|
||||||
|
> kagentz — no `/root` access from this container — and was verified unreachable, not merely stale.
|
||||||
|
|
||||||
## Purpose
|
## Purpose
|
||||||
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever.
|
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever.
|
||||||
@@ -125,15 +131,44 @@ updateNode(id, {
|
|||||||
|
|
||||||
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
|
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
|
||||||
|
|
||||||
|
### 5. Duplicate-Node Detection (Level 1 — read-only, every run)
|
||||||
|
|
||||||
|
The graph's duplicate problem is rarely an agent mistyping a title: it is **recurring writers creating a new
|
||||||
|
node per run instead of updating one**. This phase detects that class and reports it. It is read-only and
|
||||||
|
**never merges**.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 /home/hermes/.hermes/scripts/memory_dup_detect.py --json
|
||||||
|
```
|
||||||
|
Read-only, ~15s over the whole graph, exit 0. That script is the source of truth for the clustering logic —
|
||||||
|
do not re-implement it in the prompt or hand-count "duplicates" from titles.
|
||||||
|
|
||||||
|
Consume each `items[]` entry's `verdict` field; do not invent your own:
|
||||||
|
|
||||||
|
| `verdict` | Meaning | Required action |
|
||||||
|
|---|---|---|
|
||||||
|
| `WRITER-DEFECT` (`run_family: true`) | ONE scheduled task writes a new node per run | Report the ids, the `agents` (the writer) and `span_days`. **Never merge** — each node is that run's audit record. If the family grew since the last report, say `UNFIXED` and name the writer. |
|
||||||
|
| `SAFE-MERGE` | Bodies identical | Still requires an explicit `merge #A into #B` decision from Kwame. |
|
||||||
|
| `HUMAN-DECISION` | Same subject, bodies differ | Propose **connect (an edge)**, never merge. |
|
||||||
|
|
||||||
|
- **Title overlap alone is not duplication.** Four distinct client workflows of one family (#357-#361) and two
|
||||||
|
different machines' migrations (#1792/#1793) both score high on title tokens while their bodies sit 0.1-0.3
|
||||||
|
apart. Confirm against body similarity before calling anything a duplicate.
|
||||||
|
- Report clusters as **candidates for Kwame's decision**, never as established duplicates — a wrong auto-merge
|
||||||
|
destroys distinct content irrecoverably.
|
||||||
|
- Per-run history nodes are kept deliberately. Bulk-merging a run family destroys the audit trail the family exists for.
|
||||||
|
|
||||||
## Level 2 Escalations (Kwame Decision Required)
|
## Level 2 Escalations (Kwame Decision Required)
|
||||||
|
|
||||||
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
|
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
|
||||||
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
|
2. **Duplicate Nodes** — as detected by fix 5, by `verdict`, never by raw title overlap. `WRITER-DEFECT` is a writer fix (update one canonical node), not a merge decision; `SAFE-MERGE` and `HUMAN-DECISION` clusters are escalated for merge-or-connect.
|
||||||
3. **Orphan Nodes >90 days old** — Archive or connect?
|
3. **Orphan Nodes >90 days old** — Archive or connect?
|
||||||
|
|
||||||
## Reporting Format
|
## Reporting Format
|
||||||
|
|
||||||
The fixer reports to Kwame via this Zulip DM:
|
The fixer does **not** send anything. Under the single-egress model (2026-09-21) every report leaves the node
|
||||||
|
through Mumuni's gate (`comms_drop.py` for the queue, `comms_gate.py` to release and read-back verify), so
|
||||||
|
exit 0 means QUEUED, never delivered. A report body is written to a file and handed to the outbox helper:
|
||||||
|
|
||||||
```
|
```
|
||||||
🦅 Memory Fixer — [HH:MM UTC]
|
🦅 Memory Fixer — [HH:MM UTC]
|
||||||
@@ -147,8 +182,10 @@ Stale nodes needing review (max 10):
|
|||||||
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
|
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
|
||||||
...
|
...
|
||||||
|
|
||||||
Duplicates needing decision:
|
Duplicate clusters (candidates — Kwame decides; the fixer never merges unilaterally):
|
||||||
1. [Node #AAA] vs [Node #BBB] — Same title
|
1. [WRITER-DEFECT] #AAA/#BBB/#CCC — writer <agent>, N nodes, span Nd (UNFIXED if it grew since the last report)
|
||||||
|
2. [HUMAN-DECISION] #DDD/#EEE — same subject, bodies differ, SUGGEST: connect
|
||||||
|
3. "none" when the scan returned no clusters
|
||||||
|
|
||||||
Orphans >90 days:
|
Orphans >90 days:
|
||||||
1. [Node #EEE] Title — X days stale, orphaned
|
1. [Node #EEE] Title — X days stale, orphaned
|
||||||
@@ -195,6 +232,7 @@ The result must be 0 rows when all decisions are executed. Report what was done.
|
|||||||
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
||||||
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
|
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
|
||||||
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
||||||
|
- **Duplicate scan ran:** every report carries the fix 5 block (`none` when there were no clusters). A report with no duplicate section means phase 5 was skipped — a silently skipped detection phase is the failure this phase exists to prevent.
|
||||||
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
|
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
|
||||||
|
|
||||||
## Logging
|
## Logging
|
||||||
|
|||||||
+39
-5
@@ -2,6 +2,10 @@
|
|||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: pm2-self-heal
|
name: pm2-self-heal
|
||||||
description: >
|
description: >
|
||||||
|
Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
|
||||||
|
and auto-restarts any that are stopped or errored. Logs every action to
|
||||||
|
Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary).
|
||||||
|
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -12,6 +16,8 @@ description: >
|
|||||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||||
- last_check: timestamp
|
- last_check: timestamp
|
||||||
|
|
||||||
|
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed.
|
||||||
|
|
||||||
|
|
||||||
## Continuity
|
## Continuity
|
||||||
|
|
||||||
@@ -35,6 +41,9 @@ description: >
|
|||||||
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
||||||
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
||||||
spoton incident). Alerts include the restart count.
|
spoton incident). Alerts include the restart count.
|
||||||
|
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
|
||||||
|
four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton
|
||||||
|
reference above is historical context for the crash-loop guard, not a live process.
|
||||||
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
||||||
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
||||||
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
||||||
@@ -42,19 +51,44 @@ description: >
|
|||||||
and PM2 counter reset on 2026-06-28.
|
and PM2 counter reset on 2026-06-28.
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||||
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||||
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||||
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||||
|
execution — it only proves the agent read the result and reported it. The actual
|
||||||
|
monitoring work happens in the host cron job.
|
||||||
|
|
||||||
|
|
||||||
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
||||||
2. **Check abiba-telegram**:
|
2. **Check abiba-telegram** (safe to auto-restart):
|
||||||
- If status is "online" → pass
|
- If status is "online" → pass
|
||||||
- If status is "stopped" or "errored" → apply Rule 1
|
- If status is "stopped" or "errored" → apply Rule 1
|
||||||
- If restarts > 5 → alert owner
|
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
|
||||||
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
||||||
- If status is "online" → pass, log restarts count
|
- If status is "online" → pass, log restarts count
|
||||||
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
||||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||||
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
4. **Check gitea-runner**:
|
||||||
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
- If status is "online" → pass
|
||||||
6. **Wait 5 min** → repeat from step 1
|
- If status is "stopped" or "errored" → apply Rule 1
|
||||||
|
- If restarts > 5 → alert owner
|
||||||
|
5. **Check zulip-watchdog**:
|
||||||
|
- If status is "online" → pass
|
||||||
|
- If status is "stopped" or "errored" → apply Rule 1
|
||||||
|
- If restarts > 5 → alert owner
|
||||||
|
6. **Log results** — the durable per-run record is `/var/log/contract-runs/pm2-self-heal-<UTCstamp>.log` on CT 100, written by `scripts/contract-run.sh` from `/etc/cron.d/contract-runner` every 4 hours, with a firstmate inbox note raised on any non-zero exit.
|
||||||
|
|
||||||
|
**CORRECTED 2026-09-26 (relay-785):** this step previously required appending to
|
||||||
|
`SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea and called it a "hard rule".
|
||||||
|
That posting was **never implemented** — `scripts/pm2-self-heal.sh` contains no
|
||||||
|
Gitea or git-push code — so the requirement was a coverage claim the executor did
|
||||||
|
not honour, and `health-logs/pm2/` has held only its init commit since 2026-07-28.
|
||||||
|
The claim is retired rather than implemented: the contract-runner's per-run logs
|
||||||
|
plus its failure note already give a durable record and a working alarm, and a
|
||||||
|
second posting path would add work without adding a signal. `health-logs/pm2/`
|
||||||
|
is left as historical evidence, not as a live obligation.
|
||||||
|
7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||||
|
8. **Wait 5 min** → repeat from step 1
|
||||||
|
|
||||||
## Example Output (when healthy)
|
## Example Output (when healthy)
|
||||||
|
|
||||||
|
|||||||
@@ -89,6 +89,48 @@ agent: abiba
|
|||||||
| minipve | 192.168.68.12 | PVE |
|
| minipve | 192.168.68.12 | PVE |
|
||||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
||||||
|
|
||||||
|
## PBS GC (Proxmox Backup Server)
|
||||||
|
|
||||||
|
### Schedule
|
||||||
|
Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**.
|
||||||
|
|
||||||
|
**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.
|
||||||
|
|
||||||
|
### What Actually Runs
|
||||||
|
Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`.
|
||||||
|
|
||||||
|
**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC.
|
||||||
|
|
||||||
|
### Datastore Location
|
||||||
|
- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free)
|
||||||
|
- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume)
|
||||||
|
|
||||||
|
### Liveness Check
|
||||||
|
The new `proxmox-monitor.sh` leg checks storepve-datastore GC health:
|
||||||
|
- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json`
|
||||||
|
- **FAILS** if `last-run-endtime` is older than 48 hours
|
||||||
|
- Reports age in hours and pending-bytes
|
||||||
|
|
||||||
|
**All six verdict shapes** (exactly as emitted by the script):
|
||||||
|
|
||||||
|
1. **Healthy** (fresh GC, 0 B pending):
|
||||||
|
`✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)`
|
||||||
|
|
||||||
|
2. **Stale** (GC ran >48h ago):
|
||||||
|
`🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)`
|
||||||
|
|
||||||
|
3. **Probe-failed: empty read** (000/timeout/unreadable):
|
||||||
|
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)`
|
||||||
|
|
||||||
|
4. **Probe-failed: unparseable** (non-empty but invalid JSON — the "command not found" case):
|
||||||
|
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)`
|
||||||
|
|
||||||
|
5. **Never-run: datastore absent** (valid JSON but storepve-datastore not in list):
|
||||||
|
`🔴 PBS GC: never-run (storepve-datastore not found in GC list)`
|
||||||
|
|
||||||
|
6. **Never-run: no endtime** (valid JSON with datastore present but last-run-endtime is null/0):
|
||||||
|
`🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)`
|
||||||
|
|
||||||
## Operations
|
## Operations
|
||||||
|
|
||||||
### view-dashboards
|
### view-dashboards
|
||||||
|
|||||||
@@ -50,6 +50,11 @@ Changelog:
|
|||||||
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
|
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
|
||||||
from her side. This script must not probe mumuni or .24 — the v2 changelog
|
from her side. This script must not probe mumuni or .24 — the v2 changelog
|
||||||
roster line was the last reference still placing her at .24 / CT100.
|
roster line was the last reference still placing her at .24 / CT100.
|
||||||
|
v6 (2026-09-28): .8 GPU health probe now runs as `llmuser` instead of `root`.
|
||||||
|
Root SSH to .8 was lost when the guest was rebuilt, so every .8 leg read as
|
||||||
|
UNREACHABLE for a healthy host. llmuser owns llama-server and can read
|
||||||
|
`systemctl is-active`, `systemctl show -p MainPID`, and the :8080 pid.
|
||||||
|
.110 and .15 keep the default `root` user.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
import subprocess, json, sys, os, time, re, io, contextlib
|
import subprocess, json, sys, os, time, re, io, contextlib
|
||||||
@@ -70,7 +75,7 @@ PVE_NODES = {
|
|||||||
|
|
||||||
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
||||||
AGENTS = {
|
AGENTS = {
|
||||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
|
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "minipve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "hybrid"},
|
||||||
# abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its
|
# abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its
|
||||||
# local env file (key_env below), not from the shared vault or .bashrc.
|
# local env file (key_env below), not from the shared vault or .bashrc.
|
||||||
# runtime=pi: abiba has run pi-only since the harness purge. There is no
|
# runtime=pi: abiba has run pi-only since the harness purge. There is no
|
||||||
@@ -95,7 +100,7 @@ AGENTS = {
|
|||||||
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
|
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
|
||||||
# .15 strixhalo (amdpve) -> strix-server.service (active)
|
# .15 strixhalo (amdpve) -> strix-server.service (active)
|
||||||
GPU_HOSTS = {
|
GPU_HOSTS = {
|
||||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"},
|
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service", "user": "llmuser"},
|
||||||
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
|
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
|
||||||
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
|
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
|
||||||
}
|
}
|
||||||
@@ -122,10 +127,8 @@ def _fail(key, agent_name=None):
|
|||||||
|
|
||||||
|
|
||||||
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
|
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
|
||||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
|
||||||
|
|
||||||
# Fallback: if no env token, read the shared vault token file
|
|
||||||
if not INFISICAL_TOKEN:
|
if not INFISICAL_TOKEN:
|
||||||
|
# Fallback: read the shared vault token file
|
||||||
_token_path = os.path.expanduser("~/.infisical-token")
|
_token_path = os.path.expanduser("~/.infisical-token")
|
||||||
if os.path.isfile(_token_path):
|
if os.path.isfile(_token_path):
|
||||||
try:
|
try:
|
||||||
@@ -133,6 +136,7 @@ if not INFISICAL_TOKEN:
|
|||||||
INFISICAL_TOKEN = _f.read().strip()
|
INFISICAL_TOKEN = _f.read().strip()
|
||||||
except (OSError, UnicodeDecodeError):
|
except (OSError, UnicodeDecodeError):
|
||||||
pass
|
pass
|
||||||
|
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||||
|
|
||||||
# ── Helpers ──────────────────────────────────────────────────────────
|
# ── Helpers ──────────────────────────────────────────────────────────
|
||||||
|
|
||||||
@@ -148,6 +152,18 @@ def ssh(host, cmd, user="root"):
|
|||||||
except:
|
except:
|
||||||
return None
|
return None
|
||||||
|
|
||||||
|
def get_user_home(user):
|
||||||
|
"""Resolve the home directory for a user.
|
||||||
|
|
||||||
|
For 'root', returns '/root'. For any other user, returns '/home/<user>'.
|
||||||
|
This is used to construct paths that reference a user's home directory
|
||||||
|
(e.g., ~/.local/bin/hermes, ~/.hermes/config.yaml) instead of hardcoding /root/.
|
||||||
|
"""
|
||||||
|
if user == "root":
|
||||||
|
return "/root"
|
||||||
|
else:
|
||||||
|
return f"/home/{user}"
|
||||||
|
|
||||||
def http_get(url, headers=None, timeout=5):
|
def http_get(url, headers=None, timeout=5):
|
||||||
"""Return HTTP status code as string."""
|
"""Return HTTP status code as string."""
|
||||||
try:
|
try:
|
||||||
@@ -301,13 +317,14 @@ def check_gpu_ports():
|
|||||||
host = gpu["host"]
|
host = gpu["host"]
|
||||||
port = gpu["port"]
|
port = gpu["port"]
|
||||||
svc = gpu["service"]
|
svc = gpu["service"]
|
||||||
|
user = gpu.get("user", "root") # default root, overridden per-host where needed
|
||||||
|
|
||||||
# `systemctl is-active` exits non-zero when the unit is inactive or
|
# `systemctl is-active` exits non-zero when the unit is inactive or
|
||||||
# missing, which the ssh() helper would swallow as an SSH failure and
|
# missing, which the ssh() helper would swallow as an SSH failure and
|
||||||
# report as UNREACHABLE. `|| true` keeps the real state word so we can
|
# report as UNREACHABLE. `|| true` keeps the real state word so we can
|
||||||
# tell "unit inactive" from "host unreachable".
|
# tell "unit inactive" from "host unreachable".
|
||||||
svc_status = ssh(host, f"systemctl is-active {svc} || true")
|
svc_status = ssh(host, f"systemctl is-active {svc} || true", user=user)
|
||||||
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
|
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1", user=user)
|
||||||
|
|
||||||
if not svc_status:
|
if not svc_status:
|
||||||
print(f" ❌ {label}: UNREACHABLE")
|
print(f" ❌ {label}: UNREACHABLE")
|
||||||
@@ -318,14 +335,14 @@ def check_gpu_ports():
|
|||||||
print(f" ❌ {label}: PORT {port} NOT LISTENING (svc={svc_status})")
|
print(f" ❌ {label}: PORT {port} NOT LISTENING (svc={svc_status})")
|
||||||
FAIL.append(f"gpu-no-port:{label}")
|
FAIL.append(f"gpu-no-port:{label}")
|
||||||
elif svc_status != "active":
|
elif svc_status != "active":
|
||||||
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2")
|
svc_pid = ssh(host, f"systemctl show {svc} -p MainPID 2>/dev/null | cut -d= -f2", user=user)
|
||||||
if svc_pid and port_owner != svc_pid:
|
if svc_pid and port_owner != svc_pid:
|
||||||
print(f" ❌ {label}: GHOST PROCESS — port owned by pid {port_owner}, svc pid {svc_pid} (svc={svc_status})")
|
print(f" ❌ {label}: GHOST PROCESS — port owned by pid {port_owner}, svc pid {svc_pid} (svc={svc_status})")
|
||||||
FAIL.append(f"gpu-ghost:{label}:{port_owner}")
|
FAIL.append(f"gpu-ghost:{label}:{port_owner}")
|
||||||
else:
|
else:
|
||||||
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
|
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
|
||||||
else:
|
else:
|
||||||
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health")
|
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health", user=user)
|
||||||
if health and '"status":"ok"' in health:
|
if health and '"status":"ok"' in health:
|
||||||
print(f" ✅ {label}: healthy (pid={port_owner})")
|
print(f" ✅ {label}: healthy (pid={port_owner})")
|
||||||
elif health and '"status":"no slot available"' in health:
|
elif health and '"status":"no slot available"' in health:
|
||||||
@@ -377,10 +394,10 @@ def check_agents():
|
|||||||
ct = agent["ct"]
|
ct = agent["ct"]
|
||||||
report_only = agent.get("report_only", False)
|
report_only = agent.get("report_only", False)
|
||||||
|
|
||||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
|
# Tanko runs hybrid (DSH + Hermes) since 2026-08-27 — it runs both DSH and Hermes gateway.
|
||||||
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
|
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
|
||||||
# Non-Hermes runtimes have no gateway to probe. dsh = Tanko since
|
# Non-Hermes runtimes have no gateway to probe. dsh/pi-only skip the check;
|
||||||
# 2026-08-27; pi = abiba since the harness purge (.24 is pi-only).
|
# hybrid runs both DSH and Hermes and is checked normally.
|
||||||
if agent.get("runtime") in ("dsh", "pi"):
|
if agent.get("runtime") in ("dsh", "pi"):
|
||||||
is_dsh = agent.get("runtime") == "dsh"
|
is_dsh = agent.get("runtime") == "dsh"
|
||||||
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
|
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
|
||||||
@@ -501,7 +518,7 @@ def check_ct_liveness():
|
|||||||
def check_config_integrity():
|
def check_config_integrity():
|
||||||
"""Verify agent config.yaml parses as valid YAML."""
|
"""Verify agent config.yaml parses as valid YAML."""
|
||||||
for name, agent in AGENTS.items():
|
for name, agent in AGENTS.items():
|
||||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
|
# DSH/pi-only runtimes have no Hermes config.yaml; hybrid has both.
|
||||||
if agent.get("runtime") == "dsh":
|
if agent.get("runtime") == "dsh":
|
||||||
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
|
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
|
||||||
continue
|
continue
|
||||||
@@ -514,11 +531,11 @@ def check_config_integrity():
|
|||||||
print(f" ⬜ {name}: cannot SSH — skip config check")
|
print(f" ⬜ {name}: cannot SSH — skip config check")
|
||||||
continue
|
continue
|
||||||
|
|
||||||
|
home = get_user_home(user)
|
||||||
|
|
||||||
# Check YAML parses
|
# Check YAML parses
|
||||||
yaml_ok = ssh(host,
|
yaml_ok = ssh(host,
|
||||||
"python3 -c "
|
f"python3 -c \"import yaml; yaml.safe_load(open('{home}/.hermes/config.yaml')); print('OK')\" 2>&1 || echo 'FAIL'",
|
||||||
'"import yaml; yaml.safe_load(open(\'/root/.hermes/config.yaml\')); print(\'OK\')" '
|
|
||||||
"2>&1 || echo 'FAIL'",
|
|
||||||
user=user)
|
user=user)
|
||||||
if not yaml_ok:
|
if not yaml_ok:
|
||||||
print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)")
|
print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)")
|
||||||
@@ -557,7 +574,7 @@ def _infisical_invocation_paths(wrapper_body):
|
|||||||
def check_wrapper_integrity():
|
def check_wrapper_integrity():
|
||||||
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
||||||
for name, agent in AGENTS.items():
|
for name, agent in AGENTS.items():
|
||||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
|
# DSH/pi-only runtimes have no hermes CLI wrapper; hybrid has both.
|
||||||
if agent.get("runtime") == "dsh":
|
if agent.get("runtime") == "dsh":
|
||||||
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
|
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
|
||||||
continue
|
continue
|
||||||
@@ -571,7 +588,8 @@ def check_wrapper_integrity():
|
|||||||
continue
|
continue
|
||||||
|
|
||||||
# Check wrapper exists
|
# Check wrapper exists
|
||||||
wrapper = ssh(host, "ls -la /root/.local/bin/hermes 2>/dev/null", user=user)
|
home = get_user_home(user)
|
||||||
|
wrapper = ssh(host, f"ls -la {home}/.local/bin/hermes 2>/dev/null", user=user)
|
||||||
if not wrapper:
|
if not wrapper:
|
||||||
# Check alternate wrapper locations
|
# Check alternate wrapper locations
|
||||||
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
|
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
|
||||||
@@ -594,7 +612,7 @@ def check_wrapper_integrity():
|
|||||||
# a removed path (litellm-api-keys.prose.md documents
|
# a removed path (litellm-api-keys.prose.md documents
|
||||||
# `rm -f /usr/local/bin/infisical`) must neither produce a dangling path
|
# `rm -f /usr/local/bin/infisical`) must neither produce a dangling path
|
||||||
# nor trigger the PATH check — it is not an invocation.
|
# nor trigger the PATH check — it is not an invocation.
|
||||||
wrapper_body = ssh(host, "cat /root/.local/bin/hermes 2>/dev/null", user=user) or ""
|
wrapper_body = ssh(host, f"cat {home}/.local/bin/hermes 2>/dev/null", user=user) or ""
|
||||||
wrapper_code = "\n".join(line.split("#", 1)[0] for line in wrapper_body.splitlines())
|
wrapper_code = "\n".join(line.split("#", 1)[0] for line in wrapper_body.splitlines())
|
||||||
invoked_paths = _infisical_invocation_paths(wrapper_body)
|
invoked_paths = _infisical_invocation_paths(wrapper_body)
|
||||||
if "infisical" in wrapper_code:
|
if "infisical" in wrapper_code:
|
||||||
@@ -628,24 +646,35 @@ def check_wrapper_integrity():
|
|||||||
else:
|
else:
|
||||||
print(f" ℹ️ {name}: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK")
|
print(f" ℹ️ {name}: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK")
|
||||||
|
|
||||||
# Check hermes-real exists
|
# Check that the wrapper's target resolves. The fleet's wrappers do NOT
|
||||||
|
# all use a hermes-real indirection — some exec the venv module directly.
|
||||||
|
# Verify the wrapper actually points to something runnable.
|
||||||
hermes_real = ssh(host,
|
hermes_real = ssh(host,
|
||||||
"ls -la /root/.local/bin/hermes-real 2>/dev/null || echo MISS",
|
f"ls -la {home}/.local/bin/hermes-real 2>/dev/null || echo MISS",
|
||||||
user=user)
|
user=user)
|
||||||
if not hermes_real or hermes_real.strip() == "MISS":
|
if hermes_real and hermes_real.strip() != "MISS":
|
||||||
# Check venv path
|
print(f" ✅ {name}: wrapper shape: hermes-real at {home}/.local/bin/hermes-real")
|
||||||
hermes_real = ssh(host,
|
else:
|
||||||
"ls -la /usr/local/lib/hermes-agent/venv/bin/hermes 2>/dev/null || echo MISS",
|
# Try the venv under home
|
||||||
|
venv_home = ssh(host,
|
||||||
|
f"test -x {home}/.hermes/hermes-agent/venv/bin/python && echo OK || echo MISS",
|
||||||
user=user)
|
user=user)
|
||||||
if not hermes_real or hermes_real.strip() == "MISS":
|
if venv_home and venv_home.strip().splitlines()[-1] == "OK":
|
||||||
print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)")
|
print(f" ✅ {name}: wrapper shape: direct venv exec ({home}/.hermes/hermes-agent/venv/bin/python)")
|
||||||
_fail(f"wrapper-no-hermes-real:{name}", name)
|
|
||||||
else:
|
else:
|
||||||
print(f" ✅ {name}: hermes-real at alt path")
|
# Try the system-wide venv
|
||||||
|
venv_sys = ssh(host,
|
||||||
|
"test -x /usr/local/lib/hermes-agent/venv/bin/python && echo OK || echo MISS",
|
||||||
|
user=user)
|
||||||
|
if venv_sys and venv_sys.strip().splitlines()[-1] == "OK":
|
||||||
|
print(f" ✅ {name}: wrapper shape: system venv (/usr/local/lib/hermes-agent/venv/bin/python)")
|
||||||
|
else:
|
||||||
|
print(f" ❌ {name}: wrapper target NOT RESOLVABLE (no hermes-real, no venv)")
|
||||||
|
_fail(f"wrapper-no-hermes-real:{name}", name)
|
||||||
|
|
||||||
# Check the .env file has the key
|
# Check the .env file has the key
|
||||||
env_has_key = ssh(host,
|
env_has_key = ssh(host,
|
||||||
"grep -c 'LITELLM_API_KEY' /root/.hermes/.env 2>/dev/null || echo 0",
|
f"grep -c 'LITELLM_API_KEY' {home}/.hermes/.env 2>/dev/null || echo 0",
|
||||||
user=user)
|
user=user)
|
||||||
if env_has_key and env_has_key.strip() not in ("", "0"):
|
if env_has_key and env_has_key.strip() not in ("", "0"):
|
||||||
print(f" ✅ {name}: wrapper + .env key present")
|
print(f" ✅ {name}: wrapper + .env key present")
|
||||||
|
|||||||
Executable
+225
@@ -0,0 +1,225 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# contract-run.sh — Deterministic contract execution from machine scheduler
|
||||||
|
#
|
||||||
|
# Takes a contract name, resolves its script, runs it with timeout,
|
||||||
|
# logs output to $CONTRACT_RUN_LOG_DIR (default: /var/log/contract-runs/),
|
||||||
|
# and alerts on failure.
|
||||||
|
#
|
||||||
|
# Environment:
|
||||||
|
# CONTRACT_RUN_LOG_DIR Override the log directory (default: /var/log/contract-runs)
|
||||||
|
#
|
||||||
|
# Usage: bash scripts/contract-run.sh <contract-name>
|
||||||
|
#
|
||||||
|
# Contract names map to scripts as follows:
|
||||||
|
# infrastructure-monitoring -> scripts/infra-monitoring.sh
|
||||||
|
# proxmox-monitor -> scripts/proxmox-monitor.sh
|
||||||
|
# zulip-health -> scripts/zulip-monitor.sh
|
||||||
|
# agent-health-check -> scripts/agent-health-check.py
|
||||||
|
# litellm-health -> scripts/litellm-health-check.py
|
||||||
|
# disk-gc-threat-response -> scripts/disk-gc-scan.py
|
||||||
|
# pm2-self-heal -> scripts/pm2-self-heal.sh
|
||||||
|
# search-stack-visibility -> scripts/search-stack-check.py
|
||||||
|
# health-log-freshness -> scripts/health-log-freshness.py
|
||||||
|
#
|
||||||
|
# Execution copy: every contract pins the clone this script lives in (see
|
||||||
|
# docs/contract-execution-pinning.md). Before a contract runs, this wrapper
|
||||||
|
# proves the script it is about to execute byte-matches origin/master:
|
||||||
|
# CONTRACT_REVISION_PREFLIGHT=warn (default) log a refusal, still report
|
||||||
|
# CONTRACT_REVISION_PREFLIGHT=enforce refuse to report on a mismatch
|
||||||
|
# CONTRACT_REVISION_PREFLIGHT=off skip the check entirely
|
||||||
|
# A refusal names its class: cannot-verify:fetch-failed|ref-unresolvable,
|
||||||
|
# or mismatch:content|path-absent|detached-head|clone-ahead.
|
||||||
|
#
|
||||||
|
# Exit codes:
|
||||||
|
# 0 = contract passed
|
||||||
|
# 1 = contract failed (alert sent)
|
||||||
|
# 2 = probe failed (script missing, timeout, etc.)
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
CONTRACT_NAME="$1"
|
||||||
|
SCRIPTS_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
LOG_DIR="${CONTRACT_RUN_LOG_DIR:-/var/log/contract-runs}"
|
||||||
|
TIMESTAMP=$(date -u '+%Y%m%d-%H%M%S')
|
||||||
|
LOG_FILE="${LOG_DIR}/${CONTRACT_NAME}-${TIMESTAMP}.log"
|
||||||
|
|
||||||
|
# Ensure log directory exists
|
||||||
|
mkdir -p "$LOG_DIR"
|
||||||
|
|
||||||
|
# Map contract name to script path
|
||||||
|
case "$CONTRACT_NAME" in
|
||||||
|
infrastructure-monitoring)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/infra-monitoring.sh"
|
||||||
|
INTERPRETER="bash"
|
||||||
|
;;
|
||||||
|
proxmox-monitor)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/proxmox-monitor.sh"
|
||||||
|
INTERPRETER="bash"
|
||||||
|
;;
|
||||||
|
zulip-health)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/zulip-monitor.sh"
|
||||||
|
INTERPRETER="bash"
|
||||||
|
;;
|
||||||
|
agent-health-check)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/agent-health-check.py"
|
||||||
|
INTERPRETER="python3"
|
||||||
|
;;
|
||||||
|
litellm-health)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/litellm-health-check.py"
|
||||||
|
INTERPRETER="python3"
|
||||||
|
;;
|
||||||
|
pm2-self-heal)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/pm2-self-heal.sh"
|
||||||
|
INTERPRETER="bash"
|
||||||
|
;;
|
||||||
|
disk-gc-threat-response)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/disk-gc-scan.py"
|
||||||
|
INTERPRETER="python3"
|
||||||
|
;;
|
||||||
|
search-stack-visibility)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/search-stack-check.py"
|
||||||
|
INTERPRETER="python3"
|
||||||
|
;;
|
||||||
|
health-log-freshness)
|
||||||
|
SCRIPT_PATH="${SCRIPTS_DIR}/health-log-freshness.py"
|
||||||
|
INTERPRETER="python3"
|
||||||
|
;;
|
||||||
|
*)
|
||||||
|
echo "Unknown contract: $CONTRACT_NAME" | tee -a "$LOG_FILE"
|
||||||
|
# Send alert for unknown contract
|
||||||
|
ALERT_MSG="🔴 Contract $CONTRACT_NAME: unknown contract name. Log: $LOG_FILE"
|
||||||
|
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
||||||
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||||
|
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
||||||
|
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
||||||
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||||
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||||
|
-d "type=private" \
|
||||||
|
-d "to=9" \
|
||||||
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
||||||
|
fi
|
||||||
|
exit 2
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
# Check if script exists
|
||||||
|
if [ ! -f "$SCRIPT_PATH" ]; then
|
||||||
|
echo "Script not found: $SCRIPT_PATH" | tee -a "$LOG_FILE"
|
||||||
|
# Send alert for missing script
|
||||||
|
ALERT_MSG="🔴 Contract $CONTRACT_NAME: script not found at $SCRIPT_PATH. Log: $LOG_FILE"
|
||||||
|
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
||||||
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||||
|
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
||||||
|
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
||||||
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||||
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||||
|
-d "type=private" \
|
||||||
|
-d "to=9" \
|
||||||
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
||||||
|
fi
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Run the script with timeout and capture output
|
||||||
|
echo "=== Contract: $CONTRACT_NAME ===" | tee "$LOG_FILE"
|
||||||
|
echo "Started: $(date -u '+%Y-%m-%d %H:%M:%S UTC')" | tee -a "$LOG_FILE"
|
||||||
|
echo "Script: $SCRIPT_PATH" | tee -a "$LOG_FILE"
|
||||||
|
echo "" | tee -a "$LOG_FILE"
|
||||||
|
|
||||||
|
# ── Revision preflight ───────────────────────────────────────────────────────
|
||||||
|
# A verdict is only meaningful if it came from the merged copy. This reports
|
||||||
|
# whether the executing copy matches, and distinguishes "could not check" from
|
||||||
|
# "this copy is wrong" so an operator can tell them apart.
|
||||||
|
# See docs/contract-execution-pinning.md.
|
||||||
|
#
|
||||||
|
# Default is WARN, not enforce: the guard gates every scheduled contract, and a
|
||||||
|
# legitimate state (branch mid-review, detached HEAD, briefly offline) would
|
||||||
|
# otherwise turn the whole fleet's monitoring into withheld verdicts. The
|
||||||
|
# criteria for flipping the default to enforce are written down in the doc.
|
||||||
|
REVISION_PREFLIGHT_MODE="${CONTRACT_REVISION_PREFLIGHT:-warn}"
|
||||||
|
REPO_ROOT="$(cd "${SCRIPTS_DIR}/.." && pwd)"
|
||||||
|
if [ "$REVISION_PREFLIGHT_MODE" != "off" ] && [ -x "${SCRIPTS_DIR}/revision-preflight.sh" ]; then
|
||||||
|
PREFLIGHT_OUT="$(mktemp)"
|
||||||
|
if "${SCRIPTS_DIR}/revision-preflight.sh" "$SCRIPT_PATH" "$REPO_ROOT" >"$PREFLIGHT_OUT" 2>&1; then
|
||||||
|
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
|
||||||
|
else
|
||||||
|
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
|
||||||
|
PREFLIGHT_REASON="$(grep -m1 '^REASON=' "$PREFLIGHT_OUT" | cut -d= -f2-)"
|
||||||
|
[ -n "$PREFLIGHT_REASON" ] || PREFLIGHT_REASON="unclassified"
|
||||||
|
if [ "$REVISION_PREFLIGHT_MODE" = "enforce" ]; then
|
||||||
|
echo "🚫 VERDICT WITHHELD: $PREFLIGHT_REASON" | tee -a "$LOG_FILE"
|
||||||
|
ALERT_MSG="🔴 Contract $CONTRACT_NAME: revision preflight REFUSED ($PREFLIGHT_REASON) — verdict withheld. Log: $LOG_FILE"
|
||||||
|
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
||||||
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||||
|
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
||||||
|
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
||||||
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||||
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||||
|
-d "type=private" \
|
||||||
|
-d "to=9" \
|
||||||
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
||||||
|
fi
|
||||||
|
rm -f "$PREFLIGHT_OUT"
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
echo "⚠️ revision preflight: $PREFLIGHT_REASON — continuing because CONTRACT_REVISION_PREFLIGHT=$REVISION_PREFLIGHT_MODE" | tee -a "$LOG_FILE"
|
||||||
|
fi
|
||||||
|
rm -f "$PREFLIGHT_OUT"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Use timeout to prevent hangs (10 minutes default)
|
||||||
|
TIMEOUT=600
|
||||||
|
timeout "$TIMEOUT" $INTERPRETER "$SCRIPT_PATH" 2>&1 | tee -a "$LOG_FILE"
|
||||||
|
EXIT_CODE=${PIPESTATUS[0]}
|
||||||
|
|
||||||
|
# If timeout killed the process, EXIT_CODE will be 124
|
||||||
|
if [ $EXIT_CODE -eq 124 ]; then
|
||||||
|
echo "⏰ TIMEOUT: script exceeded ${TIMEOUT}s limit" | tee -a "$LOG_FILE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "" | tee -a "$LOG_FILE"
|
||||||
|
if [ $EXIT_CODE -eq 0 ]; then
|
||||||
|
echo "✅ VERDICT: PASS" | tee -a "$LOG_FILE"
|
||||||
|
exit 0
|
||||||
|
else
|
||||||
|
echo "🔴 VERDICT: FAIL (exit code $EXIT_CODE)" | tee -a "$LOG_FILE"
|
||||||
|
|
||||||
|
# Send alert (Zulip DM to user 9 + stream agent-hub topic alerts-infra)
|
||||||
|
# Using the same alert path as other monitors
|
||||||
|
ALERT_MSG="🔴 Contract $CONTRACT_NAME failed (exit $EXIT_CODE). Log: $LOG_FILE"
|
||||||
|
ALERT_SENT=false
|
||||||
|
|
||||||
|
# Take credentials from environment (ZULIP_API_KEY required)
|
||||||
|
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
||||||
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||||
|
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
||||||
|
|
||||||
|
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
||||||
|
# DM to user 9
|
||||||
|
DM_EXIT=0
|
||||||
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||||
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||||
|
-d "type=private" \
|
||||||
|
-d "to=9" \
|
||||||
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || DM_EXIT=$?
|
||||||
|
|
||||||
|
# Stream agent-hub topic alerts-infra
|
||||||
|
STREAM_EXIT=0
|
||||||
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
||||||
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
||||||
|
-d "type=stream" \
|
||||||
|
-d "to=agent-hub" \
|
||||||
|
-d "topic=alerts-infra" \
|
||||||
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || STREAM_EXIT=$?
|
||||||
|
|
||||||
|
if [ $DM_EXIT -eq 0 ] || [ $STREAM_EXIT -eq 0 ]; then
|
||||||
|
ALERT_SENT=true
|
||||||
|
else
|
||||||
|
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT FAILURE: DM exit=$DM_EXIT, stream exit=$STREAM_EXIT" >> "$LOG_FILE"
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT SKIPPED: no ZULIP_API_KEY or curl" >> "$LOG_FILE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
+225
-35
@@ -16,14 +16,42 @@ from email.mime.text import MIMEText
|
|||||||
from email.mime.multipart import MIMEMultipart
|
from email.mime.multipart import MIMEMultipart
|
||||||
|
|
||||||
PVE = "https://192.168.68.12:8006"
|
PVE = "https://192.168.68.12:8006"
|
||||||
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
|
||||||
|
# The HTTP header prefix is a protocol constant, not a credential. It is kept as
|
||||||
|
# a constant ending at '=' so that no assembled header-plus-token literal ever
|
||||||
|
# appears in the tree; the secret scanner rightly flags that shape.
|
||||||
|
PVE_AUTH_HEADER = "PVEAPIToken="
|
||||||
|
|
||||||
|
|
||||||
|
def pve_auth():
|
||||||
|
"""PVE API auth header, resolved at call time from the injected environment.
|
||||||
|
|
||||||
|
The token is injected by ``infisical run --env=prod`` as ``PVE_TOKEN``
|
||||||
|
(format ``user@realm!tokenid=secret``). It must never be hardcoded: a
|
||||||
|
placeholder literal authenticates as nobody, which is how this probe
|
||||||
|
reported zero nodes while still exiting 0. Raise loudly instead.
|
||||||
|
"""
|
||||||
|
token = os.environ.get("PVE_TOKEN")
|
||||||
|
if not token:
|
||||||
|
raise RuntimeError("PVE_TOKEN is not set (run under `infisical run --env=prod`)")
|
||||||
|
return f"Authorization: {PVE_AUTH_HEADER}{token}"
|
||||||
|
|
||||||
# ── Shared credentials —─
|
# ── Shared credentials —─
|
||||||
|
|
||||||
ZULIP_SITE = "https://chat.sysloggh.net"
|
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||||
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||||
ZULIP_KEY = "cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
# Note: /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential).
|
||||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_KEY}"
|
# No Zulip API key is required for this call. If a future leg genuinely needs abiba-bot's key,
|
||||||
|
# it must prove it with a 200 from /api/v1/users/me as abiba-bot and label itself degraded when it cannot.
|
||||||
|
# Never fall back to the vault's shared ZULIP_API_KEY.
|
||||||
|
ZULIP_AUTH = None
|
||||||
|
DEGRADED_LEGS = []
|
||||||
|
|
||||||
|
# Probe failures are different from degraded legs. A missing credential is an
|
||||||
|
# expected, survivable state (stays exit 0). A probe that cannot reach the API
|
||||||
|
# means the report has NO data for that section, which is a monitoring loss and
|
||||||
|
# must exit non-zero so it cannot pass unnoticed.
|
||||||
|
PROBE_FAILURES = []
|
||||||
|
|
||||||
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||||
LITELLM_BACKEND = "192.168.68.116"
|
LITELLM_BACKEND = "192.168.68.116"
|
||||||
@@ -44,9 +72,13 @@ TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
|
|||||||
# ── Helpers ──
|
# ── Helpers ──
|
||||||
|
|
||||||
def pve_get(path):
|
def pve_get(path):
|
||||||
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list)."""
|
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list).
|
||||||
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
|
||||||
|
A missing PVE_TOKEN is caught here and reported as ``None`` so the caller
|
||||||
|
records a probe failure; it must not escape as an unhandled exception.
|
||||||
|
"""
|
||||||
try:
|
try:
|
||||||
|
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{pve_auth()}"'
|
||||||
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
|
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
|
||||||
if r.returncode != 0:
|
if r.returncode != 0:
|
||||||
return None
|
return None
|
||||||
@@ -114,6 +146,7 @@ def collect():
|
|||||||
report["node_count"] = 0
|
report["node_count"] = 0
|
||||||
report["nodes_online"] = 0
|
report["nodes_online"] = 0
|
||||||
report["pve_probe_status"] = "unreachable"
|
report["pve_probe_status"] = "unreachable"
|
||||||
|
PROBE_FAILURES.append("proxmox: node list unreachable (PVE_TOKEN missing or API down)")
|
||||||
else:
|
else:
|
||||||
report["nodes"] = {n["node"]: {
|
report["nodes"] = {n["node"]: {
|
||||||
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||||
@@ -133,6 +166,7 @@ def collect():
|
|||||||
if resources is None:
|
if resources is None:
|
||||||
vms = []
|
vms = []
|
||||||
report["resources_probe_status"] = "unreachable"
|
report["resources_probe_status"] = "unreachable"
|
||||||
|
PROBE_FAILURES.append("proxmox: cluster resources unreachable")
|
||||||
else:
|
else:
|
||||||
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||||
report["resources_probe_status"] = "ok"
|
report["resources_probe_status"] = "ok"
|
||||||
@@ -156,7 +190,7 @@ def collect():
|
|||||||
# ── Storage ──
|
# ── Storage ──
|
||||||
storages = pve_get("/api2/json/nodes/storepve/storage")
|
storages = pve_get("/api2/json/nodes/storepve/storage")
|
||||||
report["storage"] = []
|
report["storage"] = []
|
||||||
for s in storages:
|
for s in (storages or []):
|
||||||
total = s.get("total",0) or 1
|
total = s.get("total",0) or 1
|
||||||
used = s.get("used",0)
|
used = s.get("used",0)
|
||||||
pct = used/total*100
|
pct = used/total*100
|
||||||
@@ -209,7 +243,7 @@ def collect():
|
|||||||
("Pulse", "https://pulse.sysloggh.net"),
|
("Pulse", "https://pulse.sysloggh.net"),
|
||||||
("Proxmox", "https://192.168.68.12:8006"),
|
("Proxmox", "https://192.168.68.12:8006"),
|
||||||
("SearXNG", "http://192.168.68.7:8888"),
|
("SearXNG", "http://192.168.68.7:8888"),
|
||||||
("Firecrawl", "http://192.168.68.7:3002/health"),
|
("Firecrawl", "http://192.168.68.7:3002/"), # Firecrawl serves no /health - the root is its liveness endpoint
|
||||||
]
|
]
|
||||||
report["endpoints"] = []
|
report["endpoints"] = []
|
||||||
for name, url in endpoints:
|
for name, url in endpoints:
|
||||||
@@ -355,6 +389,28 @@ def collect():
|
|||||||
|
|
||||||
# ── HTML Dashboard ──
|
# ── HTML Dashboard ──
|
||||||
|
|
||||||
|
def classify_endpoint(code):
|
||||||
|
"""Classify an endpoint probe per the fleet's probe policy.
|
||||||
|
|
||||||
|
Codified 2026-09-14 in the monitoring contracts: ANY HTTP status proves the
|
||||||
|
service answered, so the service is ALIVE - 200/301/302/401/403/404 alike.
|
||||||
|
Only a failed CONNECTION (000 / timeout / refused) is a failed probe. A 404
|
||||||
|
from a wrong path is not a service fault and must not render as one.
|
||||||
|
|
||||||
|
This replaces a string comparison that was wrong in both directions
|
||||||
|
(`ep["code"] >= "400"`): it rendered 301 as red, 404 as yellow, and a real
|
||||||
|
500 as yellow. 5xx is kept as its own "server error" signal rather than
|
||||||
|
being merged with 4xx.
|
||||||
|
"""
|
||||||
|
if not code or code == "000":
|
||||||
|
return "red", "no connection"
|
||||||
|
if code.startswith("5"):
|
||||||
|
return "yellow", "server error"
|
||||||
|
if code.startswith(("2", "3", "4")):
|
||||||
|
return "green", "alive"
|
||||||
|
return "yellow", f"unexpected {code}"
|
||||||
|
|
||||||
|
|
||||||
def build_html(r):
|
def build_html(r):
|
||||||
issues = []
|
issues = []
|
||||||
|
|
||||||
@@ -590,7 +646,7 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']
|
|||||||
# ── Network Endpoints ──
|
# ── Network Endpoints ──
|
||||||
html += '<div class="card"><h2>🌐 Network Endpoints</h2><table><tr><th>Service</th><th>Status</th></tr>'
|
html += '<div class="card"><h2>🌐 Network Endpoints</h2><table><tr><th>Service</th><th>Status</th></tr>'
|
||||||
for ep in r["endpoints"]:
|
for ep in r["endpoints"]:
|
||||||
color = "green" if ep["code"] in ("200","302","401") else ("yellow" if ep["code"] >= "400" else "red")
|
color = classify_endpoint(ep["code"])[0]
|
||||||
html += f'<tr><td>{ep["name"]}</td><td class="{color}">HTTP {ep["code"]}</td></tr>'
|
html += f'<tr><td>{ep["name"]}</td><td class="{color}">HTTP {ep["code"]}</td></tr>'
|
||||||
html += '</table></div>'
|
html += '</table></div>'
|
||||||
|
|
||||||
@@ -654,60 +710,188 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']
|
|||||||
return html
|
return html
|
||||||
|
|
||||||
|
|
||||||
# ── Send Email ──
|
# ── Delivery: Zulip DM carrying the report as an HTML ATTACHMENT ──
|
||||||
|
#
|
||||||
|
# Captain's decision, clarified 2026-09-26: the report is sent as an HTML FILE,
|
||||||
|
# i.e. an attachment - NOT HTML rendered in the message body, and NOT a Markdown
|
||||||
|
# translation of it. So the styled dashboard is built exactly as before, uploaded
|
||||||
|
# through Zulip's file-upload API, and the message body stays short: subject,
|
||||||
|
# top-line status, and a pointer to the attachment.
|
||||||
|
#
|
||||||
|
# This removes the Google dependency entirely (no SMTP, no EMAIL_PASSWORD).
|
||||||
|
# The 10,000-character message cap does not apply: it bounds message TEXT only,
|
||||||
|
# and the report travels as a file.
|
||||||
|
|
||||||
def send_email(html_content, subject_prefix=""):
|
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||||
FROM = "abiba@sysloggh.com"
|
ZULIP_BOT_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||||
TO = "jerome@sysloggh.com"
|
CAPTAIN_USER_ID = 9
|
||||||
SUBJECT = f"{subject_prefix}{'🏗️ Infrastructure Report — ' + DATE_STR}"
|
ZULIP_KEY_FILE = "/root/.pi/agent/extensions/zulip/.env"
|
||||||
|
REPORT_ARTIFACT_DIR = "/var/log/daily-infra-report"
|
||||||
|
|
||||||
msg = MIMEMultipart("alternative")
|
|
||||||
msg["From"] = FROM
|
|
||||||
msg["To"] = TO
|
|
||||||
msg["Subject"] = SUBJECT
|
|
||||||
msg.attach(MIMEText("Infrastructure report in HTML format — enable images to view.", "plain"))
|
|
||||||
msg.attach(MIMEText(html_content, "html"))
|
|
||||||
|
|
||||||
|
def zulip_key():
|
||||||
|
"""abiba-bot's Zulip key, from the env or the on-host 600 file."""
|
||||||
|
key = os.environ.get("ABIBA_ZULIP_API_KEY")
|
||||||
|
if key:
|
||||||
|
return key.strip()
|
||||||
try:
|
try:
|
||||||
EMAIL_PASSWORD = "rgbuomwcydxwbszd"
|
with open(ZULIP_KEY_FILE) as fh:
|
||||||
GMAIL_EMAIL = "jtabiri@gmail.com"
|
for line in fh:
|
||||||
|
if line.startswith("ABIBA_ZULIP_API_KEY="):
|
||||||
|
return line.split("=", 1)[1].strip()
|
||||||
|
except OSError:
|
||||||
|
return None
|
||||||
|
return None
|
||||||
|
|
||||||
server = smtplib.SMTP("smtp.gmail.com", 587)
|
|
||||||
server.starttls()
|
def build_summary(r, filename, test=False):
|
||||||
server.login(GMAIL_EMAIL, EMAIL_PASSWORD)
|
"""Short Markdown body: subject, top-line status, pointer to the attachment.
|
||||||
server.sendmail(FROM, [TO], msg.as_string())
|
|
||||||
server.quit()
|
Deliberately NOT a reproduction of the report - the attachment is the report.
|
||||||
return True, "✅ Email sent to jerome@sysloggh.com"
|
"""
|
||||||
except Exception as e:
|
nodes = f"{r.get('nodes_online', 0)}/{r.get('node_count', 0)} nodes online"
|
||||||
return False, f"❌ Email failed: {e}"
|
guests = f"{r.get('running_vms', 0)}/{r.get('total_vms', 0)} guests running"
|
||||||
|
lines = [
|
||||||
|
("\U0001F9EA **TEST — **" if test else "") + "\U0001F3D7\uFE0F **Infrastructure Report — " + DATE_STR + "**",
|
||||||
|
f"**{nodes}** \u00b7 **{guests}** \u00b7 generated {TIME_STR}",
|
||||||
|
]
|
||||||
|
problems = []
|
||||||
|
if r.get("pve_probe_status") != "ok":
|
||||||
|
problems.append(f"\u274c Proxmox probe: {r.get('pve_probe_status')}")
|
||||||
|
if r.get("resources_probe_status") != "ok":
|
||||||
|
problems.append(f"\u274c Resources probe: {r.get('resources_probe_status')}")
|
||||||
|
lit = r.get("litellm", {}) or {}
|
||||||
|
checks = lit.get("checks", []) or []
|
||||||
|
if checks:
|
||||||
|
passed = sum(1 for c in checks if c.get("status") == "pass")
|
||||||
|
if passed != len(checks):
|
||||||
|
problems.append(f"\u274c LiteLLM: {passed}/{len(checks)} checks pass")
|
||||||
|
if not (r.get("zulip_ext", {}) or {}).get("connected"):
|
||||||
|
problems.append("\u274c Zulip extension: not connected")
|
||||||
|
for leg in DEGRADED_LEGS:
|
||||||
|
problems.append(f"\u26a0\uFE0F degraded: {leg}")
|
||||||
|
|
||||||
|
lines.append("\n".join(problems) if problems else "\u2705 All monitored services healthy")
|
||||||
|
lines.append(f"\U0001F4CE **Full report attached:** `{filename}`")
|
||||||
|
return "\n\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def _curl(args, timeout=60):
|
||||||
|
r = subprocess.run(["curl", "-s", "-m", str(timeout)] + args,
|
||||||
|
capture_output=True, text=True)
|
||||||
|
try:
|
||||||
|
return json.loads(r.stdout or "{}"), r.stdout
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
return {}, r.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def _curl_json(args, timeout=90):
|
||||||
|
r = subprocess.run(["curl", "-s", "-m", str(timeout)] + args,
|
||||||
|
capture_output=True, text=True)
|
||||||
|
try:
|
||||||
|
return json.loads(r.stdout or "{}"), r.stdout
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
return {}, r.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def send_zulip(html_content, report, test=False):
|
||||||
|
"""Upload the styled HTML and post a short pointer to the captain's DM.
|
||||||
|
|
||||||
|
Returns (ok, message). On ANY failure the report body is also printed to
|
||||||
|
stdout and persisted to disk, so a delivery failure can never swallow the
|
||||||
|
content - the defect this folds in.
|
||||||
|
"""
|
||||||
|
os.makedirs(REPORT_ARTIFACT_DIR, exist_ok=True)
|
||||||
|
stamp = NOW.strftime("%Y%m%d-%H%M%S")
|
||||||
|
filename = f"infra-report-{stamp}.html"
|
||||||
|
html_path = os.path.join(REPORT_ARTIFACT_DIR, filename)
|
||||||
|
try:
|
||||||
|
with open(html_path, "w") as fh:
|
||||||
|
fh.write(html_content)
|
||||||
|
except OSError as e:
|
||||||
|
print(f" \u26a0\uFE0F could not persist report artifact: {e}", file=sys.stderr)
|
||||||
|
|
||||||
|
key = zulip_key()
|
||||||
|
if not key:
|
||||||
|
print(html_content) # never swallow the content
|
||||||
|
return False, ("\u274c Delivery FAILED: no Zulip credential "
|
||||||
|
"(ABIBA_ZULIP_API_KEY unset and "
|
||||||
|
f"{ZULIP_KEY_FILE} unreadable). Report persisted to {html_path}")
|
||||||
|
|
||||||
|
auth = ["-u", f"{ZULIP_BOT_EMAIL}:{key}"]
|
||||||
|
|
||||||
|
# 1. Upload the report as a file.
|
||||||
|
up, up_raw = _curl_json(auth + [
|
||||||
|
"-X", "POST", f"{ZULIP_SITE}/api/v1/user_uploads",
|
||||||
|
"-F", f"file=@{html_path};type=text/html",
|
||||||
|
])
|
||||||
|
if up.get("result") != "success" or not up.get("uri"):
|
||||||
|
print(html_content)
|
||||||
|
return False, (f"\u274c Delivery FAILED at upload: {up.get('msg') or up_raw[:160]} "
|
||||||
|
f"(report persisted to {html_path})")
|
||||||
|
|
||||||
|
uri = up["uri"]
|
||||||
|
size = os.path.getsize(html_path)
|
||||||
|
|
||||||
|
# 2. Post a short message pointing at it.
|
||||||
|
body = build_summary(report, filename, test=test)
|
||||||
|
link = f"[{filename}]({uri})"
|
||||||
|
body = body.replace(f"`{filename}`", link)
|
||||||
|
payload, raw = _curl_json(auth + [
|
||||||
|
"-X", "POST", f"{ZULIP_SITE}/api/v1/messages",
|
||||||
|
"-d", "type=private",
|
||||||
|
"-d", f"to=[{CAPTAIN_USER_ID}]",
|
||||||
|
"--data-urlencode", f"content={body}",
|
||||||
|
])
|
||||||
|
if payload.get("result") == "success":
|
||||||
|
return True, (f"\u2705 Delivered to Zulip DM (user {CAPTAIN_USER_ID}), "
|
||||||
|
f"message id {payload.get('id')}, attachment {size} bytes at {uri}")
|
||||||
|
|
||||||
|
print(html_content)
|
||||||
|
return False, (f"\u274c Delivery FAILED at message post: {payload.get('msg') or raw[:160]} "
|
||||||
|
f"(uploaded {uri}; report persisted to {html_path})")
|
||||||
|
|
||||||
|
|
||||||
# ── Main ──
|
# ── Main ──
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
is_test = "--test-email" in sys.argv
|
is_test = ("--test-email" in sys.argv) or ("--test-zulip" in sys.argv)
|
||||||
|
|
||||||
print(f"{'🧪 TEST MODE' if is_test else '📊'} Collecting infrastructure data...")
|
print(f"{'🧪 TEST MODE' if is_test else '📊'} Collecting infrastructure data...")
|
||||||
report = collect()
|
report = collect()
|
||||||
|
|
||||||
if "--json" in sys.argv:
|
if "--json" in sys.argv:
|
||||||
print(json.dumps(report, indent=2, default=str))
|
print(json.dumps(report, indent=2, default=str))
|
||||||
|
if PROBE_FAILURES:
|
||||||
|
for leg in PROBE_FAILURES:
|
||||||
|
print(f"PROBE FAILURE: {leg}", file=sys.stderr)
|
||||||
|
sys.exit(1)
|
||||||
sys.exit(0)
|
sys.exit(0)
|
||||||
|
|
||||||
print(" Building dashboard...")
|
print(" Building dashboard...")
|
||||||
html = build_html(report)
|
html = build_html(report)
|
||||||
|
print(f" report ready: {len(html)} chars of HTML (delivered as a file attachment)")
|
||||||
|
|
||||||
if is_test:
|
if is_test:
|
||||||
prefix = "🧪 TEST — "
|
print(" Sending TEST message to the captain's Zulip DM...")
|
||||||
print(" Sending test email...")
|
|
||||||
else:
|
else:
|
||||||
prefix = ""
|
print(" Sending to the captain's Zulip DM...")
|
||||||
print(" Sending email...")
|
|
||||||
|
|
||||||
ok, msg = send_email(html, subject_prefix=prefix)
|
ok, msg = send_zulip(html, report, test=is_test)
|
||||||
print(f" {msg}")
|
print(f" {msg}")
|
||||||
|
|
||||||
# Show summary
|
# Show summary
|
||||||
|
if DEGRADED_LEGS:
|
||||||
|
print(f"\n⚠️ Degraded legs ({len(DEGRADED_LEGS)}):")
|
||||||
|
for leg in DEGRADED_LEGS:
|
||||||
|
print(f" - {leg}")
|
||||||
|
else:
|
||||||
|
print("\n✅ All legs fully credentialed")
|
||||||
|
|
||||||
|
# A failed send must exit non-zero; a degraded leg (no credential) must stay exit 0
|
||||||
|
if not ok:
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
|
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
|
||||||
print(f"\n📋 Summary:")
|
print(f"\n📋 Summary:")
|
||||||
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
|
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
|
||||||
@@ -718,3 +902,9 @@ if __name__ == "__main__":
|
|||||||
for k,v in report.get('agents',{}).items():
|
for k,v in report.get('agents',{}).items():
|
||||||
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
||||||
print(f" Agents: {', '.join(agent_parts)}")
|
print(f" Agents: {', '.join(agent_parts)}")
|
||||||
|
|
||||||
|
if PROBE_FAILURES:
|
||||||
|
print(f"\n❌ Probe failures ({len(PROBE_FAILURES)}):")
|
||||||
|
for leg in PROBE_FAILURES:
|
||||||
|
print(f" - {leg}")
|
||||||
|
sys.exit(1)
|
||||||
|
|||||||
+238
-8
@@ -101,8 +101,6 @@ GUESTS: list[Guest] = [
|
|||||||
# amdpve (192.168.68.15)
|
# amdpve (192.168.68.15)
|
||||||
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
|
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
|
||||||
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
|
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
|
||||||
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
|
|
||||||
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
|
|
||||||
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
|
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
|
||||||
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
|
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
|
||||||
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
|
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
|
||||||
@@ -110,6 +108,8 @@ GUESTS: list[Guest] = [
|
|||||||
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
|
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
|
||||||
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
|
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
|
||||||
# minipve (192.168.68.12)
|
# minipve (192.168.68.12)
|
||||||
|
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="minipve",
|
||||||
|
access_method="pct-run", probe_target="tanko (CT 112, minipve)"),
|
||||||
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
|
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
|
||||||
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
|
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
|
||||||
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
|
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
|
||||||
@@ -153,6 +153,25 @@ GPU_HOSTS = [
|
|||||||
CONNECT_TIMEOUT = 5
|
CONNECT_TIMEOUT = 5
|
||||||
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
|
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
|
||||||
|
|
||||||
|
# Host filesystem thresholds (from contract)
|
||||||
|
HOST_THRESHOLDS = {
|
||||||
|
"WARN": 85,
|
||||||
|
"AMBER": 90,
|
||||||
|
"RED": 95,
|
||||||
|
}
|
||||||
|
|
||||||
|
# State file path (absolute, so execution context doesn't matter)
|
||||||
|
STATE_FILE = pathlib.Path(__file__).resolve().parent.parent / "state" / "host-disk-bands.json"
|
||||||
|
|
||||||
|
# PVE nodes to probe for host filesystems
|
||||||
|
HOST_NODES = [
|
||||||
|
{"hostname": "acerpve", "ip": "192.168.68.9"},
|
||||||
|
{"hostname": "amdpve", "ip": "192.168.68.15"},
|
||||||
|
{"hostname": "storepve", "ip": "192.168.68.6"},
|
||||||
|
{"hostname": "minipve", "ip": "192.168.68.12"},
|
||||||
|
{"hostname": "ocupve", "ip": "192.168.68.5"},
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
|
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
|
||||||
"""Run a command and return (exit_code, stdout, stderr)."""
|
"""Run a command and return (exit_code, stdout, stderr)."""
|
||||||
@@ -298,6 +317,151 @@ def scan_fleet() -> list[dict]:
|
|||||||
return results
|
return results
|
||||||
|
|
||||||
|
|
||||||
|
def classify_band(usage_pct: float) -> str:
|
||||||
|
"""Classify a percentage into a band."""
|
||||||
|
if usage_pct >= HOST_THRESHOLDS["RED"]:
|
||||||
|
return "HOST-RED"
|
||||||
|
elif usage_pct >= HOST_THRESHOLDS["AMBER"]:
|
||||||
|
return "HOST-AMBER"
|
||||||
|
elif usage_pct >= HOST_THRESHOLDS["WARN"]:
|
||||||
|
return "HOST-WARN"
|
||||||
|
else:
|
||||||
|
return "GREEN"
|
||||||
|
|
||||||
|
|
||||||
|
def probe_host_filesystems() -> tuple[list[dict], dict[str, str]]:
|
||||||
|
"""Probe host filesystems on all PVE nodes.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
- List of host filesystem results
|
||||||
|
- Dict of volume_key -> current_band (for state file)
|
||||||
|
"""
|
||||||
|
results = []
|
||||||
|
current_bands = {}
|
||||||
|
|
||||||
|
for node in HOST_NODES:
|
||||||
|
ip = node["ip"]
|
||||||
|
hostname = node["hostname"]
|
||||||
|
|
||||||
|
# Probe df for host filesystems
|
||||||
|
probe_cmd = f'ssh {SSH_OPTS} root@{ip} "df -hP / /media/* tank 2>/dev/null | tail -n +2"'
|
||||||
|
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
|
||||||
|
|
||||||
|
if exit_code != 0:
|
||||||
|
results.append({
|
||||||
|
"target": f"{hostname} ({ip})",
|
||||||
|
"hostname": hostname,
|
||||||
|
"ip": ip,
|
||||||
|
"reachable": False,
|
||||||
|
"volumes": [],
|
||||||
|
"probe_cmd": probe_cmd,
|
||||||
|
"failure_kind": "ssh-error",
|
||||||
|
})
|
||||||
|
continue
|
||||||
|
|
||||||
|
# Parse df output and classify each volume
|
||||||
|
volumes = []
|
||||||
|
for line in stdout.splitlines():
|
||||||
|
parts = line.split()
|
||||||
|
if len(parts) < 6:
|
||||||
|
continue
|
||||||
|
|
||||||
|
dev, size, used, avail, pct_str, mount = parts[:6]
|
||||||
|
pct = float(pct_str.rstrip("%"))
|
||||||
|
band = classify_band(pct)
|
||||||
|
|
||||||
|
# Volume type classification
|
||||||
|
if mount.startswith("/media/"):
|
||||||
|
vol_type = "media"
|
||||||
|
elif mount == "/" or "pve" in dev:
|
||||||
|
vol_type = "host-root"
|
||||||
|
elif mount == "tank" or "tank" in mount:
|
||||||
|
vol_type = "pbs-datastore"
|
||||||
|
else:
|
||||||
|
vol_type = "other"
|
||||||
|
|
||||||
|
# Volume key for state file (host/volume)
|
||||||
|
volume_key = f"{hostname}/{mount}"
|
||||||
|
current_bands[volume_key] = band
|
||||||
|
|
||||||
|
volumes.append({
|
||||||
|
"mount": mount,
|
||||||
|
"device": dev,
|
||||||
|
"size": size,
|
||||||
|
"used": used,
|
||||||
|
"avail": avail,
|
||||||
|
"pct": pct,
|
||||||
|
"band": band,
|
||||||
|
"type": vol_type,
|
||||||
|
})
|
||||||
|
|
||||||
|
results.append({
|
||||||
|
"target": f"{hostname} ({ip})",
|
||||||
|
"hostname": hostname,
|
||||||
|
"ip": ip,
|
||||||
|
"reachable": True,
|
||||||
|
"volumes": volumes,
|
||||||
|
"probe_cmd": probe_cmd,
|
||||||
|
"failure_kind": None,
|
||||||
|
})
|
||||||
|
|
||||||
|
return results, current_bands
|
||||||
|
|
||||||
|
|
||||||
|
def read_state_file() -> Optional[dict[str, str]]:
|
||||||
|
"""Read the state file if it exists."""
|
||||||
|
if not STATE_FILE.exists():
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
with open(STATE_FILE) as f:
|
||||||
|
return json.load(f)
|
||||||
|
except (json.JSONDecodeError, IOError) as e:
|
||||||
|
print(f"⚠️ State file exists but unreadable: {e}", file=sys.stderr)
|
||||||
|
return {}
|
||||||
|
|
||||||
|
|
||||||
|
def write_state_file(bands: dict[str, str]) -> None:
|
||||||
|
"""Write the state file."""
|
||||||
|
STATE_FILE.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
try:
|
||||||
|
with open(STATE_FILE, "w") as f:
|
||||||
|
json.dump(bands, f, indent=2)
|
||||||
|
except IOError as e:
|
||||||
|
print(f"⚠️ State file write failed: {e}", file=sys.stderr)
|
||||||
|
|
||||||
|
|
||||||
|
def detect_transitions(current_bands: dict[str, str], prior_bands: Optional[dict[str, str]]) -> list[dict]:
|
||||||
|
"""Detect band transitions (current vs. prior)."""
|
||||||
|
if prior_bands is None:
|
||||||
|
# First run — no transitions, just establish baseline
|
||||||
|
return []
|
||||||
|
|
||||||
|
transitions = []
|
||||||
|
# Check for volumes that moved to a higher band (escalation)
|
||||||
|
for volume, current_band in current_bands.items():
|
||||||
|
prior_band = prior_bands.get(volume, "GREEN")
|
||||||
|
|
||||||
|
# Band ordering: GREEN < HOST-WARN < HOST-AMBER < HOST-RED
|
||||||
|
band_order = {"GREEN": 0, "HOST-WARN": 1, "HOST-AMBER": 2, "HOST-RED": 3}
|
||||||
|
|
||||||
|
if band_order[current_band] > band_order[prior_band]:
|
||||||
|
transitions.append({
|
||||||
|
"type": "escalation",
|
||||||
|
"volume": volume,
|
||||||
|
"from": prior_band,
|
||||||
|
"to": current_band,
|
||||||
|
})
|
||||||
|
elif band_order[current_band] < band_order[prior_band]:
|
||||||
|
transitions.append({
|
||||||
|
"type": "recovery",
|
||||||
|
"volume": volume,
|
||||||
|
"from": prior_band,
|
||||||
|
"to": current_band,
|
||||||
|
})
|
||||||
|
|
||||||
|
return transitions
|
||||||
|
|
||||||
|
|
||||||
def render_results(results: list[dict]) -> str:
|
def render_results(results: list[dict]) -> str:
|
||||||
"""Render scan results in human-readable format."""
|
"""Render scan results in human-readable format."""
|
||||||
lines = []
|
lines = []
|
||||||
@@ -316,22 +480,88 @@ def render_results(results: list[dict]) -> str:
|
|||||||
return "\n".join(lines)
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def render_host_results(results: list[dict], transitions: list[dict], prior_bands: Optional[dict[str, str]]) -> str:
|
||||||
|
"""Render host filesystem results in human-readable format."""
|
||||||
|
lines = []
|
||||||
|
lines.append("")
|
||||||
|
lines.append("=== Host Filesystem Bands ===")
|
||||||
|
lines.append("")
|
||||||
|
|
||||||
|
# Render transitions first (they're the actionable alerts)
|
||||||
|
if prior_bands is None:
|
||||||
|
lines.append(" (first run — recording baseline, no alerts)")
|
||||||
|
elif not transitions:
|
||||||
|
lines.append(" (no band changes since last scan)")
|
||||||
|
else:
|
||||||
|
for t in transitions:
|
||||||
|
volume, from_band, to_band = t["volume"], t["from"], t["to"]
|
||||||
|
if t["type"] == "escalation":
|
||||||
|
lines.append(f" ⚠️ {volume}: {from_band} → {to_band} (ESCALATION)")
|
||||||
|
else:
|
||||||
|
lines.append(f" ✅ {volume}: {from_band} → {to_band} (RECOVERY)")
|
||||||
|
|
||||||
|
# Render all volumes with their bands
|
||||||
|
lines.append("")
|
||||||
|
for node_result in results:
|
||||||
|
if not node_result["reachable"]:
|
||||||
|
lines.append(f" ❌ {node_result['target']}: UNREACHABLE ({node_result['failure_kind']})")
|
||||||
|
continue
|
||||||
|
|
||||||
|
lines.append(f" {node_result['target']}:")
|
||||||
|
for vol in node_result["volumes"]:
|
||||||
|
lines.append(f" {vol['mount']} ({vol['type']}): {vol['pct']}% ({vol['used']}/{vol['size']}, {vol['avail']} free) -> {vol['band']}")
|
||||||
|
|
||||||
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
def main() -> int:
|
||||||
import argparse
|
import argparse
|
||||||
|
|
||||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
|
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests and host filesystems.")
|
||||||
ap.add_argument("--json", action="store_true", help="machine-readable output")
|
ap.add_argument("--json", action="store_true", help="machine-readable output")
|
||||||
|
ap.add_argument("--hosts-only", action="store_true", help="scan host filesystems only")
|
||||||
|
ap.add_argument("--guests-only", action="store_true", help="scan guests only (skip host filesystems)")
|
||||||
args = ap.parse_args()
|
args = ap.parse_args()
|
||||||
|
|
||||||
results = scan_fleet()
|
# Scan guests (unless --hosts-only)
|
||||||
|
guest_results = []
|
||||||
|
if not args.hosts_only:
|
||||||
|
guest_results = scan_fleet()
|
||||||
|
|
||||||
|
# Scan host filesystems (unless --guests-only)
|
||||||
|
host_results = []
|
||||||
|
current_bands = {}
|
||||||
|
if not args.guests_only:
|
||||||
|
host_results, current_bands = probe_host_filesystems()
|
||||||
|
|
||||||
|
# Read prior state and detect transitions
|
||||||
|
prior_bands = read_state_file()
|
||||||
|
transitions = detect_transitions(current_bands, prior_bands)
|
||||||
|
|
||||||
|
# Write new state
|
||||||
|
write_state_file(current_bands)
|
||||||
|
else:
|
||||||
|
prior_bands = None
|
||||||
|
transitions = []
|
||||||
|
|
||||||
if args.json:
|
if args.json:
|
||||||
print(json.dumps(results, indent=2))
|
# JSON output
|
||||||
|
output = {
|
||||||
|
"guests": guest_results,
|
||||||
|
"hosts": host_results,
|
||||||
|
"transitions": transitions,
|
||||||
|
"prior_bands": prior_bands,
|
||||||
|
}
|
||||||
|
print(json.dumps(output, indent=2))
|
||||||
else:
|
else:
|
||||||
print(render_results(results))
|
# Human-readable output
|
||||||
|
if guest_results:
|
||||||
|
print(render_results(guest_results))
|
||||||
|
|
||||||
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
|
if host_results:
|
||||||
# (a probe error means the probe itself failed, not just that the guest was unreachable)
|
print(render_host_results(host_results, transitions, prior_bands))
|
||||||
|
|
||||||
|
# Exit 0 if all probed (reachable or not), 1 if any probe error
|
||||||
return 0
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
Executable
+227
@@ -0,0 +1,227 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Health-log freshness watchdog (dead-man's-switch).
|
||||||
|
|
||||||
|
Absence of logs raises no alarm. On 2026-09-26 the `gpu/` directory of
|
||||||
|
SyslogSolution/health-logs went silent for 12 days unnoticed, because the only
|
||||||
|
thing that would have noticed was the job that had stopped. This check lives on
|
||||||
|
a DIFFERENT host (CT 100) from the producers, so a producer host that is dead,
|
||||||
|
or a schedule that was deleted, still raises an alarm.
|
||||||
|
|
||||||
|
It asks the Gitea API for the newest commit touching each watched directory and
|
||||||
|
fails when that commit is older than the directory's threshold.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
health-log-freshness.py # check every watched directory
|
||||||
|
health-log-freshness.py --json # machine-readable output
|
||||||
|
|
||||||
|
Exit: 0 = all fresh, 1 = at least one stale or unreachable, 2 = check could not run.
|
||||||
|
|
||||||
|
Environment:
|
||||||
|
GITEA_URL default https://git.sysloggh.net
|
||||||
|
GITEA_TOKEN API token (falls back to GITEA_PAT, then to the token
|
||||||
|
embedded in ~/.git-credentials for that host)
|
||||||
|
HEALTH_LOG_MAX_AGE_H override thresholds, e.g. HEALTH_LOG_MAX_AGE_GPU=12
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import base64
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import urllib.error
|
||||||
|
import urllib.request
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
REPO = "SyslogSolution/health-logs"
|
||||||
|
GITEA_URL = os.environ.get("GITEA_URL", "https://git.sysloggh.net").rstrip("/")
|
||||||
|
|
||||||
|
# dir -> (threshold_hours, why)
|
||||||
|
# Thresholds are sized for the producer cadence plus one missed run:
|
||||||
|
# gpu/ runs every 6h -> 12h tolerates one miss, catches a second
|
||||||
|
# litellm/ runs every 6h -> 18h (it is proven healthy; a looser bound avoids
|
||||||
|
# noise while still catching a real stop)
|
||||||
|
WATCHED: dict[str, tuple[float, str]] = {
|
||||||
|
"gpu": (12.0, "gpu-self-heal.py, CT116 cron 2 */6 * * *"),
|
||||||
|
"litellm": (18.0, "litellm-health-check.sh, CT116 cron 0 */6 * * *"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _threshold(directory: str, default: float) -> float:
|
||||||
|
"""Allow HEALTH_LOG_MAX_AGE_<DIR> to override a threshold.
|
||||||
|
|
||||||
|
Documented override; used both operationally (tighten a bound while
|
||||||
|
investigating) and in tests (force a stale verdict deterministically).
|
||||||
|
"""
|
||||||
|
raw = os.environ.get(f"HEALTH_LOG_MAX_AGE_{directory.upper()}")
|
||||||
|
if raw is None:
|
||||||
|
return default
|
||||||
|
try:
|
||||||
|
return float(raw)
|
||||||
|
except ValueError:
|
||||||
|
print(f"WARN: ignoring non-numeric HEALTH_LOG_MAX_AGE_{directory.upper()}={raw!r}")
|
||||||
|
return default
|
||||||
|
|
||||||
|
# pm2/ is deliberately NOT watched: pm2-self-heal.prose.md claimed a health-logs
|
||||||
|
# posting that its executor never implemented, and the claim was retired on
|
||||||
|
# 2026-09-26 in favour of the contract-runner's per-run logs. The directory is
|
||||||
|
# left as historical evidence, not as a live obligation.
|
||||||
|
|
||||||
|
|
||||||
|
def _auth_candidates() -> list[str]:
|
||||||
|
"""Authorization header values to try, in order.
|
||||||
|
|
||||||
|
Gitea accepts either an API token (``token <tok>``) or HTTP basic auth, and
|
||||||
|
the value held under ``GITEA_TOKEN`` is not reliably an API token - on this
|
||||||
|
host it is the git account's *password*, which sent as a token returns HTTP
|
||||||
|
401. So return every candidate and let the caller use the first that works,
|
||||||
|
rather than guessing and failing.
|
||||||
|
"""
|
||||||
|
candidates: list[str] = []
|
||||||
|
for var in ("GITEA_TOKEN", "GITEA_PAT"):
|
||||||
|
if os.environ.get(var):
|
||||||
|
candidates.append(f"token {os.environ[var]}")
|
||||||
|
candidates.append(
|
||||||
|
"Basic "
|
||||||
|
+ base64.b64encode(
|
||||||
|
f"{_git_user()}:{os.environ[var]}".encode()
|
||||||
|
).decode()
|
||||||
|
)
|
||||||
|
# ~/.git-credentials may hold this Gitea under several hostnames (the public
|
||||||
|
# name and the internal IP both appear in this fleet), so accept any of them
|
||||||
|
# rather than filtering on the configured host - that filter produced zero
|
||||||
|
# candidates whenever GITEA_URL pointed at the internal address.
|
||||||
|
try:
|
||||||
|
with open(os.path.expanduser("~/.git-credentials")) as fh:
|
||||||
|
for line in fh:
|
||||||
|
m = re.match(r"https://([^:]+):([^@]+)@", line.strip())
|
||||||
|
if m:
|
||||||
|
raw = f"{m.group(1)}:{m.group(2)}".encode()
|
||||||
|
candidates.append("Basic " + base64.b64encode(raw).decode())
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
seen, out = set(), []
|
||||||
|
for c in candidates:
|
||||||
|
if c not in seen:
|
||||||
|
seen.add(c)
|
||||||
|
out.append(c)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def _git_user() -> str:
|
||||||
|
return os.environ.get("GITEA_USER", "abiba-bot")
|
||||||
|
|
||||||
|
|
||||||
|
def newest_commit_iso(directory: str, auth: list[str] | None) -> tuple[str | None, str]:
|
||||||
|
"""Return (iso_timestamp, detail) for the newest commit touching `directory`."""
|
||||||
|
url = (
|
||||||
|
f"{GITEA_URL}/api/v1/repos/{REPO}/commits"
|
||||||
|
f"?path={directory}&limit=1&stat=false"
|
||||||
|
)
|
||||||
|
req = urllib.request.Request(url, headers={"Accept": "application/json"})
|
||||||
|
headers = list(auth or [])
|
||||||
|
last = "no credential"
|
||||||
|
for i, hdr in enumerate(headers or [None]):
|
||||||
|
r = urllib.request.Request(url, headers={"Accept": "application/json"})
|
||||||
|
if hdr:
|
||||||
|
r.add_header("Authorization", hdr)
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(r, timeout=20) as resp:
|
||||||
|
data = json.loads(resp.read().decode("utf-8", "replace"))
|
||||||
|
break
|
||||||
|
except urllib.error.HTTPError as exc:
|
||||||
|
last = f"HTTP {exc.code}"
|
||||||
|
if exc.code not in (401, 403):
|
||||||
|
return None, last
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
return None, repr(exc)
|
||||||
|
else:
|
||||||
|
return None, last
|
||||||
|
|
||||||
|
if not data:
|
||||||
|
return None, "no commits"
|
||||||
|
commit = data[0].get("commit", {})
|
||||||
|
when = (
|
||||||
|
(commit.get("committer") or {}).get("date")
|
||||||
|
or (commit.get("author") or {}).get("date")
|
||||||
|
)
|
||||||
|
msg = (commit.get("message") or "").splitlines()[0][:60]
|
||||||
|
return when, msg
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
as_json = "--json" in sys.argv
|
||||||
|
auth = _auth_candidates()
|
||||||
|
now = datetime.now(timezone.utc)
|
||||||
|
failures: list[str] = []
|
||||||
|
report: dict[str, dict] = {}
|
||||||
|
|
||||||
|
if not auth:
|
||||||
|
print("FAIL: no Gitea credential available (GITEA_TOKEN/GITEA_PAT/~/.git-credentials)")
|
||||||
|
return 2
|
||||||
|
for directory, (default_age_h, why) in WATCHED.items():
|
||||||
|
max_age_h = _threshold(directory, default_age_h)
|
||||||
|
when, detail = newest_commit_iso(directory, auth)
|
||||||
|
entry: dict = {"directory": directory, "producer": why, "max_age_h": max_age_h}
|
||||||
|
|
||||||
|
if when is None:
|
||||||
|
entry.update(ok=False, reason=f"could not read newest commit: {detail}")
|
||||||
|
failures.append(
|
||||||
|
f"health-logs/{directory}/ could not be read ({detail}) — "
|
||||||
|
f"a directory that cannot be read is indistinguishable from one that stopped"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
try:
|
||||||
|
ts = datetime.fromisoformat(when.replace("Z", "+00:00"))
|
||||||
|
except ValueError:
|
||||||
|
entry.update(ok=False, reason=f"unparseable timestamp {when!r}")
|
||||||
|
failures.append(f"health-logs/{directory}/ timestamp unparseable: {when!r}")
|
||||||
|
report[directory] = entry
|
||||||
|
continue
|
||||||
|
age_h = (now - ts).total_seconds() / 3600.0
|
||||||
|
stale = age_h > max_age_h
|
||||||
|
entry.update(
|
||||||
|
ok=not stale,
|
||||||
|
newest=when,
|
||||||
|
age_h=round(age_h, 2),
|
||||||
|
newest_commit=detail,
|
||||||
|
)
|
||||||
|
if stale:
|
||||||
|
entry["reason"] = f"stale: {age_h:.1f}h > {max_age_h}h"
|
||||||
|
failures.append(
|
||||||
|
f"health-logs/{directory}/ is STALE: newest entry {when} "
|
||||||
|
f"({age_h:.1f}h old, limit {max_age_h}h) from {why}"
|
||||||
|
)
|
||||||
|
report[directory] = entry
|
||||||
|
|
||||||
|
if as_json:
|
||||||
|
print(json.dumps({"checked": report, "failures": failures}, indent=2))
|
||||||
|
return 1 if failures else 0
|
||||||
|
|
||||||
|
print("Health-log freshness — dead-man's-switch")
|
||||||
|
print("=" * 66)
|
||||||
|
for directory, entry in report.items():
|
||||||
|
mark = "✅" if entry.get("ok") else "❌"
|
||||||
|
if entry.get("newest"):
|
||||||
|
print(
|
||||||
|
f"{mark} health-logs/{directory}/ newest {entry['newest']} "
|
||||||
|
f"({entry['age_h']}h old, limit {entry['max_age_h']}h)"
|
||||||
|
)
|
||||||
|
print(f" last commit: {entry.get('newest_commit')}")
|
||||||
|
else:
|
||||||
|
print(f"{mark} health-logs/{directory}/ {entry.get('reason')}")
|
||||||
|
print(f" producer: {entry['producer']}")
|
||||||
|
|
||||||
|
print("=" * 66)
|
||||||
|
if failures:
|
||||||
|
print("VERDICT: FAIL")
|
||||||
|
for f in failures:
|
||||||
|
print(f" - {f}")
|
||||||
|
return 1
|
||||||
|
print("VERDICT: PASS — every watched health-log directory is advancing")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
Executable
+234
@@ -0,0 +1,234 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# infrastructure-monitoring.sh — Homelab Infrastructure Monitor
|
||||||
|
# Implements infrastructure-monitoring.prose.md (check-health section)
|
||||||
|
#
|
||||||
|
# Legs: Grafana, Prometheus, LiteLLM, PVE API (5 nodes), GPU exporters,
|
||||||
|
# Docker Stats, PVE Exporter
|
||||||
|
#
|
||||||
|
# Design:
|
||||||
|
# - Every target, port, path, and expected status is defined in code
|
||||||
|
# - Liveness rule: any HTTP status = ALIVE for auth-gated/redirect endpoints;
|
||||||
|
# only connection failures (000/timeout) = probe-failed
|
||||||
|
# - Bare-200 rule: expected status must match exactly (200); anything else = alert
|
||||||
|
# - PVE API uses -k flag (self-signed certs), probes /api2/json/version
|
||||||
|
# - Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116, probed via SSH
|
||||||
|
# with one retry at longer timeout (25s connect, 30s max) to distinguish
|
||||||
|
# transient timeout from host-down
|
||||||
|
# - Non-zero exit naming every failed target; no "OK" summary when any leg failed
|
||||||
|
#
|
||||||
|
# Output shape per leg:
|
||||||
|
# ✅ <name>: alive
|
||||||
|
# 🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)
|
||||||
|
#
|
||||||
|
# Failure kinds: timeout | refused | tls | unexpected:<code> (printed in the failure line)
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
# ── Configuration (documented in infrastructure-monitoring.prose.md) ────────
|
||||||
|
# Change these in ONE place; test_infra_monitoring.sh asserts against these.
|
||||||
|
|
||||||
|
GRAFANA_HOST="192.168.68.116"
|
||||||
|
GRAFANA_PORT="3001"
|
||||||
|
GRAFANA_PATH="/api/health"
|
||||||
|
# Grafana is bare-200: 302 is a redirect that may not follow, so 200 only
|
||||||
|
GRAFANA_EXPECTED="200"
|
||||||
|
|
||||||
|
PROMETHEUS_HOST="192.168.68.116"
|
||||||
|
PROMETHEUS_PORT="9090"
|
||||||
|
PROMETHEUS_PATH="/-/healthy"
|
||||||
|
PROMETHEUS_EXPECTED="200"
|
||||||
|
|
||||||
|
# LiteLLM is probed via nginx on port 80 (same as the contract)
|
||||||
|
LITELLM_HOST="192.168.68.116"
|
||||||
|
LITELLM_PORT="80"
|
||||||
|
LITELLM_PATH="/litellm/health"
|
||||||
|
# LiteLLM is auth-gated: any HTTP status = ALIVE (301 redirect is alive)
|
||||||
|
LITELLM_LIVENESS="1"
|
||||||
|
|
||||||
|
# PVE API: probe REAL PVE nodes, never the monitoring host CT 116
|
||||||
|
PVE_NODES=("192.168.68.9" "192.168.68.12" "192.168.68.6" "192.168.68.15" "192.168.68.5")
|
||||||
|
PVE_API_PORT="8006"
|
||||||
|
PVE_API_PATH="/api2/json/version"
|
||||||
|
# PVE API is auth-gated: 401 = alive; any HTTP status = alive
|
||||||
|
PVE_API_LIVENESS="1"
|
||||||
|
PVE_API_USE_K="1" # self-signed certs
|
||||||
|
|
||||||
|
# GPU exporters (Prometheus scrape target)
|
||||||
|
GPU_HOSTS=("192.168.68.8" "192.168.68.110" "192.168.68.15")
|
||||||
|
GPU_PORT="9400"
|
||||||
|
GPU_PATH="/metrics"
|
||||||
|
GPU_EXPECTED="200"
|
||||||
|
|
||||||
|
# Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116
|
||||||
|
DOCKER_STATS_PORT="9324" # harness-docker-stats (docker_container_* metrics)
|
||||||
|
PVE_EXPORTER_PORT="9221" # harness-pve-exporter (5 pve_* metrics)
|
||||||
|
CT116_SSH_HOST="192.168.68.116"
|
||||||
|
# Both are bare-200: 404 = container not yet started
|
||||||
|
DOCKER_STATS_EXPECTED="200|404"
|
||||||
|
PVE_EXPORTER_EXPECTED="200|404"
|
||||||
|
|
||||||
|
# ── Probe Functions ─────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
# probe_http <host> <port> <path> <expected_pattern> [use_k] [ssh_host] [scheme] [liveness]
|
||||||
|
# Returns 0 if probe succeeds (matches expected or liveness), 1 if probe-failed.
|
||||||
|
# Prints the result line.
|
||||||
|
#
|
||||||
|
# FIX C1: The kind value is computed and printed in the failure line.
|
||||||
|
# FIX C2: SSH retry logic is in the first attempt branch (not unreachable).
|
||||||
|
|
||||||
|
LAST_KIND=""
|
||||||
|
probe_http() {
|
||||||
|
local host="$1" port="$2" path="$3" expected="$4"
|
||||||
|
local use_k="${5:-}" ssh_host="${6:-}" scheme="${7:-http}" liveness="${8:-0}"
|
||||||
|
local url="${scheme}://${host}:${port}${path}"
|
||||||
|
local code="" kind=""
|
||||||
|
LAST_KIND=""
|
||||||
|
|
||||||
|
# Single invocation that captures both output and status
|
||||||
|
if [ -n "$ssh_host" ]; then
|
||||||
|
out=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||||
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||||
|
rc=$?
|
||||||
|
else
|
||||||
|
out=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} "$url" 2>/dev/null)
|
||||||
|
rc=$?
|
||||||
|
fi
|
||||||
|
code=$(printf '%s' "$out" | tr -d '[:space:]')
|
||||||
|
|
||||||
|
# Classify failure kind and retry if needed
|
||||||
|
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||||
|
# Distinguish timeout from TLS error from refused
|
||||||
|
case "$rc" in
|
||||||
|
35|51|58|59|60|77|83) kind="tls" ;;
|
||||||
|
*) kind="timeout" ;;
|
||||||
|
esac
|
||||||
|
# Retry once at longer timeout (25s connect, 30s max)
|
||||||
|
if [ -n "$ssh_host" ]; then
|
||||||
|
code=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||||
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||||
|
rc=$?
|
||||||
|
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||||
|
else
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} "$url" 2>/dev/null)
|
||||||
|
rc=$?
|
||||||
|
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||||
|
# On retry, classify: still 000 = keep existing kind (or timeout if empty), unexpected status = refused
|
||||||
|
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||||
|
[ -z "$kind" ] && kind="timeout"
|
||||||
|
elif ! echo "$code" | grep -qE "^(${expected})$"; then
|
||||||
|
kind="refused"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Check result
|
||||||
|
if [ -n "$code" ] && [ "$code" != "000" ]; then
|
||||||
|
if [ "$liveness" = "1" ]; then
|
||||||
|
# Any HTTP status = ALIVE for auth-gated/redirect endpoints
|
||||||
|
return 0
|
||||||
|
else
|
||||||
|
# Bare-200 or specific expected pattern
|
||||||
|
if echo "$code" | grep -qE "^(${expected})$"; then
|
||||||
|
return 0
|
||||||
|
else
|
||||||
|
kind="unexpected:$code"
|
||||||
|
LAST_KIND="$kind"
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
[ -z "$kind" ] && kind="refused"
|
||||||
|
LAST_KIND="$kind"
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
# ── Main ────────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
FAILED=()
|
||||||
|
FAILED_KIND=()
|
||||||
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
|
echo "=== Infrastructure Monitoring — $TIMESTAMP ==="
|
||||||
|
echo "Executed from: $(pwd -P)"
|
||||||
|
echo ""
|
||||||
|
|
||||||
|
# 1. Grafana (CT 116 :3001 /api/health) — bare-200
|
||||||
|
if probe_http "$GRAFANA_HOST" "$GRAFANA_PORT" "$GRAFANA_PATH" "$GRAFANA_EXPECTED"; then
|
||||||
|
echo " ✅ Grafana: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Grafana: probe-failed: ${GRAFANA_HOST}:${GRAFANA_PORT} (expected ${GRAFANA_EXPECTED}) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("grafana")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 2. Prometheus (CT 116 :9090 /-/healthy) — bare-200
|
||||||
|
if probe_http "$PROMETHEUS_HOST" "$PROMETHEUS_PORT" "$PROMETHEUS_PATH" "$PROMETHEUS_EXPECTED"; then
|
||||||
|
echo " ✅ Prometheus: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Prometheus: probe-failed: ${PROMETHEUS_HOST}:${PROMETHEUS_PORT} (expected ${PROMETHEUS_EXPECTED}) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("prometheus")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 3. LiteLLM (CT 116 :80/litellm/health via nginx) — liveness (any HTTP = alive)
|
||||||
|
if probe_http "$LITELLM_HOST" "$LITELLM_PORT" "$LITELLM_PATH" "" "" "" "http" "$LITELLM_LIVENESS"; then
|
||||||
|
echo " ✅ LiteLLM: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 LiteLLM: probe-failed: ${LITELLM_HOST}:${LITELLM_PORT}${LITELLM_PATH} (any-HTTP liveness) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("litellm")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 4. PVE API (5 real nodes :8006 /api2/json/version, -k, liveness)
|
||||||
|
PVE_FAILED=()
|
||||||
|
for node in "${PVE_NODES[@]}"; do
|
||||||
|
if probe_http "$node" "$PVE_API_PORT" "$PVE_API_PATH" "" "$PVE_API_USE_K" "" "https" "$PVE_API_LIVENESS"; then
|
||||||
|
echo " ✅ PVE API ${node}: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 PVE API ${node}: probe-failed: ${node}:${PVE_API_PORT} (any-HTTP liveness, -k for self-signed) (<${LAST_KIND}>)"
|
||||||
|
PVE_FAILED+=("$node")
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
if [ ${#PVE_FAILED[@]} -gt 0 ]; then
|
||||||
|
FAILED+=("pve-api: ${PVE_FAILED[*]}")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 5. GPU exporters (:9400/metrics) — bare-200
|
||||||
|
GPU_FAILED=()
|
||||||
|
for host in "${GPU_HOSTS[@]}"; do
|
||||||
|
if probe_http "$host" "$GPU_PORT" "$GPU_PATH" "$GPU_EXPECTED"; then
|
||||||
|
echo " ✅ GPU exporter ${host}: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 GPU exporter ${host}: probe-failed: ${host}:${GPU_PORT} (expected 200) (<${LAST_KIND}>)"
|
||||||
|
GPU_FAILED+=("$host")
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
if [ ${#GPU_FAILED[@]} -gt 0 ]; then
|
||||||
|
FAILED+=("gpu-exporters: ${GPU_FAILED[*]}")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 6. Docker Stats (CT 116 :9324, 127.0.0.1 via SSH) — 200|404
|
||||||
|
if probe_http "127.0.0.1" "$DOCKER_STATS_PORT" "/" "$DOCKER_STATS_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||||
|
echo " ✅ Docker Stats: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:${DOCKER_STATS_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("docker-stats")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 7. PVE Exporter (CT 116 :9221, 127.0.0.1 via SSH) — 200|404
|
||||||
|
if probe_http "127.0.0.1" "$PVE_EXPORTER_PORT" "/" "$PVE_EXPORTER_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||||
|
echo " ✅ PVE Exporter: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:${PVE_EXPORTER_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("pve-exporter")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||||
|
echo " ✅ All legs OK"
|
||||||
|
exit 0
|
||||||
|
else
|
||||||
|
for f in "${FAILED[@]}"; do
|
||||||
|
echo " 🔴 FAILED: $f"
|
||||||
|
done
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
@@ -55,9 +55,12 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
|||||||
try:
|
try:
|
||||||
rc, stdout, stderr = run_command(cmd, timeout)
|
rc, stdout, stderr = run_command(cmd, timeout)
|
||||||
if rc != 0:
|
if rc != 0:
|
||||||
# Determine failure kind from curl exit code
|
# Check if this is a timeout from run_command (rc=1, stderr="TIMEOUT")
|
||||||
|
if rc == 1 and stderr == "TIMEOUT":
|
||||||
|
return (000, "timeout after " + str(timeout) + "s")
|
||||||
|
# Otherwise, determine failure kind from curl exit code
|
||||||
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
||||||
if rc == 28:
|
elif rc == 28:
|
||||||
return (000, "timeout after " + str(timeout) + "s")
|
return (000, "timeout after " + str(timeout) + "s")
|
||||||
elif rc == 7:
|
elif rc == 7:
|
||||||
return (000, "connection refused")
|
return (000, "connection refused")
|
||||||
@@ -73,6 +76,20 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
|||||||
except subprocess.TimeoutExpired:
|
except subprocess.TimeoutExpired:
|
||||||
return (000, "timeout after " + str(timeout) + "s")
|
return (000, "timeout after " + str(timeout) + "s")
|
||||||
|
|
||||||
|
def check_host_health(host_ip):
|
||||||
|
"""Check if the GPU host's llama-chat-api health endpoint is reachable
|
||||||
|
|
||||||
|
Returns: (healthy: bool, detail: str)
|
||||||
|
"""
|
||||||
|
code, _ = probe_http("http://" + host_ip + ":8080/health", timeout=10)
|
||||||
|
if code == 200:
|
||||||
|
return True, "host healthy (200)"
|
||||||
|
elif code == 000:
|
||||||
|
return False, "host unreachable (timeout or refused)"
|
||||||
|
else:
|
||||||
|
return False, "host unhealthy (HTTP " + str(code) + ")"
|
||||||
|
|
||||||
|
|
||||||
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
||||||
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
||||||
cmd = "curl -s -m " + str(timeout)
|
cmd = "curl -s -m " + str(timeout)
|
||||||
@@ -115,29 +132,53 @@ def check_model_probes():
|
|||||||
|
|
||||||
results = []
|
results = []
|
||||||
|
|
||||||
|
# Host health mapping: model -> host IP
|
||||||
|
model_hosts = {
|
||||||
|
"gpu-dense": "192.168.68.8", # RTX 3090
|
||||||
|
"gpu-vision": "192.168.68.110", # RTX 5070
|
||||||
|
"strix-moe": "192.168.68.15" # Strix Halo
|
||||||
|
}
|
||||||
|
|
||||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||||
# Single-host aliases: 30s timeout each
|
host_ip = model_hosts[model]
|
||||||
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
|
# Single-host aliases: 30s initial timeout, retry once at 90s on failure
|
||||||
|
# Worst-case prefill ~76s, so 90s retry ensures we cover it
|
||||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||||
method="POST",
|
method="POST",
|
||||||
bearer_token=monitor_key,
|
bearer_token=monitor_key,
|
||||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||||
timeout=30)
|
timeout=30)
|
||||||
|
|
||||||
|
first_kind = None
|
||||||
if code == 000 and failure_kind:
|
if code == 000 and failure_kind:
|
||||||
# Report probe failure with kind, do not assert a service verdict
|
first_kind = failure_kind
|
||||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
|
time.sleep(1)
|
||||||
|
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||||
|
method="POST",
|
||||||
|
bearer_token=monitor_key,
|
||||||
|
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||||
|
timeout=90)
|
||||||
|
|
||||||
|
if code == 000 and failure_kind:
|
||||||
|
# Both attempts failed - check host health to distinguish busy from down
|
||||||
|
host_healthy, host_detail = check_host_health(host_ip)
|
||||||
|
if host_healthy:
|
||||||
|
results.append((model, False, "busy (completion timed out after retry; " + host_detail + ")"))
|
||||||
|
else:
|
||||||
|
# Host unreachable - report both kinds
|
||||||
|
if first_kind:
|
||||||
|
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
|
||||||
|
else:
|
||||||
|
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
|
||||||
elif code == 200:
|
elif code == 200:
|
||||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||||
elif code in (401, 403):
|
elif code in (401, 403):
|
||||||
# Credential fault - capture body and key alias
|
|
||||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||||
method="POST",
|
method="POST",
|
||||||
bearer_token=monitor_key,
|
bearer_token=monitor_key,
|
||||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||||
timeout=10)
|
timeout=10)
|
||||||
# Resolve key alias
|
alias = "monitor-20260813"
|
||||||
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
|
|
||||||
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||||
else:
|
else:
|
||||||
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||||
@@ -197,16 +238,16 @@ def check_admin_key_list():
|
|||||||
# Try to parse the response
|
# Try to parse the response
|
||||||
try:
|
try:
|
||||||
data = json.loads(stdout)
|
data = json.loads(stdout)
|
||||||
# Response is a dict with "keys" field
|
# Response is a dict with "keys" (paginated list) and "total_count" fields
|
||||||
if isinstance(data, dict) and "keys" in data:
|
if isinstance(data, dict) and "keys" in data:
|
||||||
key_count = len(data["keys"])
|
key_count = data.get("total_count", len(data["keys"]))
|
||||||
elif isinstance(data, list):
|
elif isinstance(data, list):
|
||||||
key_count = len(data)
|
key_count = len(data)
|
||||||
else:
|
else:
|
||||||
key_count = 0
|
key_count = 0
|
||||||
if key_count == 0:
|
if key_count == 0:
|
||||||
return "Admin Key List", False, "admin-call-failed (empty response)"
|
return "Admin Key List", False, "admin-call-failed (empty response)"
|
||||||
return "Admin Key List", True, str(key_count) + " keys"
|
return "Admin Key List", True, str(key_count) + " total (" + str(len(data.get("keys", []) if isinstance(data, dict) else data)) + " on page 1)" if isinstance(data, dict) else str(key_count) + " total"
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
|
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
|
||||||
|
|
||||||
@@ -247,6 +288,7 @@ def main():
|
|||||||
print("")
|
print("")
|
||||||
|
|
||||||
all_pass = True
|
all_pass = True
|
||||||
|
degraded = [] # Track degraded (busy) checks
|
||||||
|
|
||||||
# Run all checks
|
# Run all checks
|
||||||
checks = [
|
checks = [
|
||||||
@@ -266,9 +308,15 @@ def main():
|
|||||||
# Model probes
|
# Model probes
|
||||||
model_results = check_model_probes()
|
model_results = check_model_probes()
|
||||||
for name, passed, detail in model_results:
|
for name, passed, detail in model_results:
|
||||||
status = "✅" if passed else "❌"
|
# Check if this is a busy (degraded) verdict
|
||||||
|
if not passed and detail.startswith("busy "):
|
||||||
|
status = "⚠️"
|
||||||
|
degraded.append(name)
|
||||||
|
else:
|
||||||
|
status = "✅" if passed else "❌"
|
||||||
print(" " + status + " " + name + ": " + detail)
|
print(" " + status + " " + name + ": " + detail)
|
||||||
if not passed:
|
# Only set all_pass=False for real failures (not busy)
|
||||||
|
if not passed and not detail.startswith("busy "):
|
||||||
all_pass = False
|
all_pass = False
|
||||||
|
|
||||||
# Admin key list
|
# Admin key list
|
||||||
@@ -294,10 +342,16 @@ def main():
|
|||||||
|
|
||||||
print("")
|
print("")
|
||||||
if all_pass:
|
if all_pass:
|
||||||
print("✅ All checks passed")
|
if degraded:
|
||||||
|
print("✅ All checks passed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||||
|
else:
|
||||||
|
print("✅ All checks passed")
|
||||||
return 0
|
return 0
|
||||||
else:
|
else:
|
||||||
print("❌ Some checks failed")
|
if degraded:
|
||||||
|
print("❌ Some checks failed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||||
|
else:
|
||||||
|
print("❌ Some checks failed")
|
||||||
return 1
|
return 1
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|||||||
+1
-1
@@ -12,7 +12,6 @@ set -euo pipefail
|
|||||||
declare -A CT_NODES=(
|
declare -A CT_NODES=(
|
||||||
# amdpve (192.168.68.15)
|
# amdpve (192.168.68.15)
|
||||||
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
|
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
|
||||||
[112]=amdpve # tanko
|
|
||||||
[113]=amdpve # baggy
|
[113]=amdpve # baggy
|
||||||
[115]=amdpve # scottdenya
|
[115]=amdpve # scottdenya
|
||||||
[120]=amdpve # adguard2 (added 2026-09-12)
|
[120]=amdpve # adguard2 (added 2026-09-12)
|
||||||
@@ -21,6 +20,7 @@ declare -A CT_NODES=(
|
|||||||
[102]=minipve # adguard (was acerpve)
|
[102]=minipve # adguard (was acerpve)
|
||||||
[104]=minipve # authentik
|
[104]=minipve # authentik
|
||||||
[110]=minipve # gitea
|
[110]=minipve # gitea
|
||||||
|
[112]=minipve # tanko (was amdpve — migrated 2026-09-27; live per pvesh)
|
||||||
[116]=minipve # syslog-api
|
[116]=minipve # syslog-api
|
||||||
[119]=minipve # infisical-vault
|
[119]=minipve # infisical-vault
|
||||||
# storepve (192.168.68.6)
|
# storepve (192.168.68.6)
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
#!/bin/bash
|
#!/bin/bash
|
||||||
# pm2-self-heal — hourly PM2 process check
|
# pm2-self-heal — hourly PM2 process check
|
||||||
# Part of the pm2-self-heal prose contract
|
# Part of the pm2-self-heal prose contract
|
||||||
# Alerts via Telegram (abiba-zulip decommissioned 2026-07-04)
|
# Alerts via Telegram (primary) and Zulip DM (secondary, abiba-zulip restored 2026-08-03)
|
||||||
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||||
|
|
||||||
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
||||||
@@ -46,9 +46,75 @@ if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
|
|||||||
fi
|
fi
|
||||||
fi
|
fi
|
||||||
|
|
||||||
|
# Check abiba-zulip (live Zulip bridge, heartbeating)
|
||||||
|
ZULIP_LINE=$(echo "$STATUS" | grep "abiba-zulip")
|
||||||
|
ZULIP_STATUS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
ZULIP_RESTARTS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||||
|
|
||||||
|
if [ "$ZULIP_STATUS" != "online" ]; then
|
||||||
|
pm2 restart abiba-zulip > /dev/null 2>&1
|
||||||
|
sleep 3
|
||||||
|
ZULIP_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-zulip")
|
||||||
|
ZULIP_STATUS2=$(echo "$ZULIP_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
if [ "$ZULIP_STATUS2" = "online" ]; then
|
||||||
|
msg="⚠️ abiba-zulip was **$ZULIP_STATUS** → restarted to online"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
else
|
||||||
|
msg="🚨 abiba-zulip **failed restart** (was $ZULIP_STATUS, still $ZULIP_STATUS2)"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
fi
|
||||||
|
elif [ "$ZULIP_RESTARTS" -gt 5 ]; then
|
||||||
|
msg="⚠️ abiba-zulip has **$ZULIP_RESTARTS** restarts (high count)"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Check gitea-runner
|
||||||
|
GITEA_LINE=$(echo "$STATUS" | grep "gitea-runner")
|
||||||
|
GITEA_STATUS=$(echo "$GITEA_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
GITEA_RESTARTS=$(echo "$GITEA_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||||
|
|
||||||
|
if [ "$GITEA_STATUS" != "online" ]; then
|
||||||
|
pm2 restart gitea-runner > /dev/null 2>&1
|
||||||
|
sleep 3
|
||||||
|
GITEA_LINE2=$(pm2 status --no-color 2>/dev/null | grep "gitea-runner")
|
||||||
|
GITEA_STATUS2=$(echo "$GITEA_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
if [ "$GITEA_STATUS2" = "online" ]; then
|
||||||
|
msg="⚠️ gitea-runner was **$GITEA_STATUS** → restarted to online"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
else
|
||||||
|
msg="🚨 gitea-runner **failed restart** (was $GITEA_STATUS, still $GITEA_STATUS2)"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
fi
|
||||||
|
elif [ "$GITEA_RESTARTS" -gt 5 ]; then
|
||||||
|
msg="⚠️ gitea-runner has **$GITEA_RESTARTS** restarts (high count)"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Check zulip-watchdog
|
||||||
|
WATCHDOG_LINE=$(echo "$STATUS" | grep "zulip-watchdog")
|
||||||
|
WATCHDOG_STATUS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
WATCHDOG_RESTARTS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||||
|
|
||||||
|
if [ "$WATCHDOG_STATUS" != "online" ]; then
|
||||||
|
pm2 restart zulip-watchdog > /dev/null 2>&1
|
||||||
|
sleep 3
|
||||||
|
WATCHDOG_LINE2=$(pm2 status --no-color 2>/dev/null | grep "zulip-watchdog")
|
||||||
|
WATCHDOG_STATUS2=$(echo "$WATCHDOG_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||||
|
if [ "$WATCHDOG_STATUS2" = "online" ]; then
|
||||||
|
msg="⚠️ zulip-watchdog was **$WATCHDOG_STATUS** → restarted to online"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
else
|
||||||
|
msg="🚨 zulip-watchdog **failed restart** (was $WATCHDOG_STATUS, still $WATCHDOG_STATUS2)"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
fi
|
||||||
|
elif [ "$WATCHDOG_RESTARTS" -gt 5 ]; then
|
||||||
|
msg="⚠️ zulip-watchdog has **$WATCHDOG_RESTARTS** restarts (high count)"
|
||||||
|
ALERTS="${ALERTS}${msg}\n"
|
||||||
|
fi
|
||||||
|
|
||||||
# Log check
|
# Log check
|
||||||
{
|
{
|
||||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS alerts=${ALERTS:+yes}"
|
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS zulip=$ZULIP_STATUS gitea=$GITEA_STATUS watchdog=$WATCHDOG_STATUS alerts=${ALERTS:+yes}"
|
||||||
[ -n "$ALERTS" ] && echo "$ALERTS"
|
[ -n "$ALERTS" ] && echo "$ALERTS"
|
||||||
} >> "$LOG"
|
} >> "$LOG"
|
||||||
|
|
||||||
|
|||||||
@@ -47,8 +47,8 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
|
|||||||
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
|
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
|
||||||
|
|
||||||
**Proxmox Cluster "Tabiri" (5 nodes):**
|
**Proxmox Cluster "Tabiri" (5 nodes):**
|
||||||
- amdpve (192.168.68.15): kagentz, tanko, baggy, scottdenya, adguard2
|
- amdpve (192.168.68.15): kagentz, baggy, scottdenya, adguard2
|
||||||
- minipve (192.168.68.12): abiba, adguard, authentik, gitea, syslog-api, infisical-vault
|
- minipve (192.168.68.12): abiba, tanko, adguard, authentik, gitea, syslog-api, infisical-vault
|
||||||
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
|
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
|
||||||
- acerpve (192.168.68.9): llm-gpu
|
- acerpve (192.168.68.9): llm-gpu
|
||||||
- ocupve (192.168.68.5): ocu-llm
|
- ocupve (192.168.68.5): ocu-llm
|
||||||
|
|||||||
+16
-1
@@ -135,7 +135,22 @@ fi
|
|||||||
|
|
||||||
echo " Cross-contract: $WARNINGS total warnings across all checks"
|
echo " Cross-contract: $WARNINGS total warnings across all checks"
|
||||||
|
|
||||||
# ── 4. Summary ──
|
# ── 4. Committed-credential scan ──
|
||||||
|
# The 2026-09-17 purge removed six live credentials that had sat in .md prose
|
||||||
|
# and scripts for weeks. This step makes that class of commit FAIL the gate
|
||||||
|
# instead of printing a warning. Patterns: scripts/secret-patterns.tsv.
|
||||||
|
# Only deliberate synthetic examples may be listed in scripts/secret-allowlist.tsv,
|
||||||
|
# each with a reason. Run `bash scripts/secret-scan.sh --staged` before committing.
|
||||||
|
echo ""
|
||||||
|
echo "── 4. Secret scan (committed credentials) ──"
|
||||||
|
if bash scripts/secret-scan.sh; then
|
||||||
|
echo " ✅ No committed credentials"
|
||||||
|
else
|
||||||
|
echo " ❌ COMMITTED CREDENTIAL DETECTED"
|
||||||
|
FAILED=1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 5. Summary ──
|
||||||
echo ""
|
echo ""
|
||||||
echo "═══════════════════════════════════"
|
echo "═══════════════════════════════════"
|
||||||
if [ $FAILED -eq 1 ]; then
|
if [ $FAILED -eq 1 ]; then
|
||||||
|
|||||||
Executable
+166
@@ -0,0 +1,166 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# proxmox-monitor.sh — Proxmox Cluster + Docker Monitoring Health Check
|
||||||
|
# Implements proxmox-monitor.prose.md (check-health section)
|
||||||
|
#
|
||||||
|
# Legs: Prometheus, Grafana, Docker Stats Exporter, PVE Exporter, PBS GC
|
||||||
|
# All legs must return 200 for healthy status.
|
||||||
|
#
|
||||||
|
# Run: bash scripts/proxmox-monitor.sh
|
||||||
|
# Exits 0 if all probes pass, 1 if any fails.
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
CT116_HOST="192.168.68.116"
|
||||||
|
FAILED=()
|
||||||
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
|
|
||||||
|
echo "=== Proxmox Monitor — $TIMESTAMP ==="
|
||||||
|
echo "Executed from: $(pwd -P)"
|
||||||
|
echo ""
|
||||||
|
|
||||||
|
# 1. Prometheus health (bound to 0.0.0.0:9090 on .116)
|
||||||
|
PROM_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:9090/-/healthy 2>/dev/null)
|
||||||
|
PROM_CODE=$(printf '%s' "$PROM_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$PROM_CODE" ] || PROM_CODE="000"
|
||||||
|
|
||||||
|
if [ "$PROM_CODE" = "200" ]; then
|
||||||
|
echo " ✅ Prometheus: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Prometheus: probe-failed: ${CT116_HOST}:9090 (expected 200, got ${PROM_CODE})"
|
||||||
|
FAILED+=("prometheus")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 2. Grafana health (bound to 0.0.0.0:3001 on .116)
|
||||||
|
GRAF_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:3001/api/health 2>/dev/null)
|
||||||
|
GRAF_CODE=$(printf '%s' "$GRAF_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$GRAF_CODE" ] || GRAF_CODE="000"
|
||||||
|
|
||||||
|
if [ "$GRAF_CODE" = "200" ]; then
|
||||||
|
echo " ✅ Grafana: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Grafana: probe-failed: ${CT116_HOST}:3001 (expected 200, got ${GRAF_CODE})"
|
||||||
|
FAILED+=("grafana")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 3. Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — probe from .116 localhost)
|
||||||
|
DOCKER_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||||
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics" 2>/dev/null)
|
||||||
|
DOCKER_CODE=$(printf '%s' "$DOCKER_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$DOCKER_CODE" ] || DOCKER_CODE="000"
|
||||||
|
|
||||||
|
if [ "$DOCKER_CODE" = "200" ]; then
|
||||||
|
echo " ✅ Docker Stats: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:9324 (expected 200, got ${DOCKER_CODE})"
|
||||||
|
FAILED+=("docker-stats")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 4. PVE Exporter (bound to 127.0.0.1:9221 on .116 — probe from .116 localhost)
|
||||||
|
PVE_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||||
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics" 2>/dev/null)
|
||||||
|
PVE_CODE=$(printf '%s' "$PVE_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$PVE_CODE" ] || PVE_CODE="000"
|
||||||
|
|
||||||
|
if [ "$PVE_CODE" = "200" ]; then
|
||||||
|
echo " ✅ PVE Exporter: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:9221 (expected 200, got ${PVE_CODE})"
|
||||||
|
FAILED+=("pve-exporter")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 5. PBS GC liveness (storepve-datastore GC must have run within 48h)
|
||||||
|
# Use absolute path for pct to avoid PATH issues in non-interactive ssh
|
||||||
|
PBS_GC_OUTPUT=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192.168.68.6 \
|
||||||
|
"/sbin/pct exec 107 -- /sbin/proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null)
|
||||||
|
PBS_GC_OUTPUT=$(printf '%s' "$PBS_GC_OUTPUT" | tr -d '[:space:]')
|
||||||
|
[ -n "$PBS_GC_OUTPUT" ] || PBS_GC_OUTPUT="000"
|
||||||
|
|
||||||
|
if [ "$PBS_GC_OUTPUT" = "000" ]; then
|
||||||
|
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
else
|
||||||
|
# Parse the JSON to get storepve-datastore's state with four distinct outcomes:
|
||||||
|
# 1. probe-failed: non-zero ssh status / empty / unparseable JSON
|
||||||
|
# 2. running: collection in progress (last-run-endtime absent or 0, but upid present)
|
||||||
|
# 3. stale: no completed run within 48h
|
||||||
|
# 4. healthy: completed within 48h
|
||||||
|
PBS_GC_RESULT=$(echo "$PBS_GC_OUTPUT" | python3 -c "
|
||||||
|
import sys, json
|
||||||
|
try:
|
||||||
|
data = json.load(sys.stdin)
|
||||||
|
for store in data:
|
||||||
|
if store['store'] == 'storepve-datastore':
|
||||||
|
endtime = store.get('last-run-endtime')
|
||||||
|
upid = store.get('upid')
|
||||||
|
pending = store.get('pending-bytes', 0)
|
||||||
|
|
||||||
|
# State 2: Running (collection in progress) — last-run-endtime absent while run is in progress
|
||||||
|
if (endtime is None or endtime == 0) and upid is not None:
|
||||||
|
print(f'running|{pending}')
|
||||||
|
break
|
||||||
|
|
||||||
|
# State 3: No completed run (never-run or stale)
|
||||||
|
if endtime is None or endtime == 0:
|
||||||
|
print(f'no-completed-run|{pending}')
|
||||||
|
break
|
||||||
|
|
||||||
|
# States 3 & 4: Completed (has endtime)
|
||||||
|
print(f'completed|{endtime}|{pending}')
|
||||||
|
break
|
||||||
|
else:
|
||||||
|
print(f'absent|0')
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
print(f'unparseable|0')
|
||||||
|
" 2>/dev/null)
|
||||||
|
|
||||||
|
# Parse the state|endtime|pending format
|
||||||
|
PBS_GC_STATE=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1)
|
||||||
|
|
||||||
|
if [ "$PBS_GC_STATE" = "unparseable" ]; then
|
||||||
|
# State 1: probe-failed (unparseable JSON)
|
||||||
|
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
elif [ "$PBS_GC_STATE" = "absent" ]; then
|
||||||
|
# State 1: probe-failed (store not found)
|
||||||
|
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (storepve-datastore not found)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
elif [ "$PBS_GC_STATE" = "running" ]; then
|
||||||
|
# State 2: collection in progress — do NOT fail
|
||||||
|
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||||
|
echo " ⏳ PBS GC: running (started: in-progress, pending-bytes: ${PENDING_BYTES} B)"
|
||||||
|
elif [ "$PBS_GC_STATE" = "no-completed-run" ]; then
|
||||||
|
# State 3: no completed run within 48h
|
||||||
|
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||||
|
echo " 🔴 PBS GC: no completed run within 48h (pending-bytes: ${PENDING_BYTES} B)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
else
|
||||||
|
# States 3 & 4: completed (has endtime)
|
||||||
|
LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||||
|
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f3)
|
||||||
|
|
||||||
|
# Convert epoch to age in hours
|
||||||
|
NOW_EPOCH=$(date -u +%s)
|
||||||
|
AGE_HOURS=$(( (NOW_EPOCH - LAST_RUN_ENDTIME) / 3600 ))
|
||||||
|
|
||||||
|
if [ $AGE_HOURS -gt 48 ]; then
|
||||||
|
# State 3: stale (no completed run within 48h)
|
||||||
|
echo " 🔴 PBS GC: stale — last completed run was ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
else
|
||||||
|
# State 4: healthy (completed within 48h)
|
||||||
|
echo " ✅ PBS GC: healthy — last completed run ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||||
|
echo ""
|
||||||
|
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||||
|
echo " ✅ All legs OK"
|
||||||
|
exit 0
|
||||||
|
else
|
||||||
|
for f in "${FAILED[@]}"; do
|
||||||
|
echo " 🔴 FAILED: $f"
|
||||||
|
done
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
Executable
+209
@@ -0,0 +1,209 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# revision-preflight.sh — prove the copy a contract is about to execute is the
|
||||||
|
# copy that is merged.
|
||||||
|
#
|
||||||
|
# Usage:
|
||||||
|
# revision-preflight.sh [options] <script-path> <clone-path>
|
||||||
|
#
|
||||||
|
# Options:
|
||||||
|
# --ref <ref> Ref to compare against (default: origin/master)
|
||||||
|
# --no-fetch Do not refresh the ref first (see FRESHNESS)
|
||||||
|
# --fetch-timeout <secs> Bound the default fetch (default: 20; 0 = no bound)
|
||||||
|
# --quiet Print nothing on success
|
||||||
|
# -h, --help Show this help
|
||||||
|
#
|
||||||
|
# Exit codes:
|
||||||
|
# 0 the executing script byte-matches <ref>:<repo-relative-path>
|
||||||
|
# 1 the copy is NOT the merged one -> REASON=mismatch:<class>
|
||||||
|
# 2 the check could not be performed -> REASON=cannot-verify:<class>
|
||||||
|
#
|
||||||
|
# Every non-zero exit prints one machine-readable line
|
||||||
|
# REASON=<class>
|
||||||
|
# followed by the human explanation. The two top-level classes are deliberately
|
||||||
|
# distinct: "I could not check" is a different situation from "this copy is
|
||||||
|
# wrong", and an operator must never have to guess which they are looking at.
|
||||||
|
#
|
||||||
|
# cannot-verify:fetch-failed the remote could not be reached (or timed out)
|
||||||
|
# cannot-verify:ref-unresolvable <ref> does not exist in the clone
|
||||||
|
# mismatch:path-absent the script does not exist in <ref>
|
||||||
|
# mismatch:content the script differs from <ref>
|
||||||
|
# mismatch:detached-head the clone is on a detached HEAD
|
||||||
|
# mismatch:clone-ahead local HEAD is ahead of <ref> (mid-review?)
|
||||||
|
#
|
||||||
|
# FRESHNESS
|
||||||
|
# A guard is only as good as the ref it compares against. On 2026-09-25 a
|
||||||
|
# stale local origin/master made an ancestry check on this fleet report
|
||||||
|
# "unlanded work" for a branch that had in fact merged, and it would equally
|
||||||
|
# have passed a stale script as current. So by default this guard FETCHES the
|
||||||
|
# remote before comparing, bounded by --fetch-timeout so a hung remote cannot
|
||||||
|
# block a scheduled contract. With --no-fetch it compares against whatever the
|
||||||
|
# local ref points at and says so out loud; it never silently assumes
|
||||||
|
# freshness.
|
||||||
|
#
|
||||||
|
# FAIL CLOSED
|
||||||
|
# An unresolvable path or ref is a FAILURE, never a warning. "Cannot verify"
|
||||||
|
# is precisely the state a stale or hand-edited copy produces, so treating it
|
||||||
|
# as success would defeat the guard. The original draft did exactly that: it
|
||||||
|
# resolved the master revision with
|
||||||
|
# `git show origin/master:$(basename "$SCRIPT")`, which drops the scripts/
|
||||||
|
# prefix, queries the repo root, fails, and exited 0 — passing a script that
|
||||||
|
# exists in no revision at all.
|
||||||
|
#
|
||||||
|
# WHICH CLONE
|
||||||
|
# Pass the clone the contract is actually executing from. See
|
||||||
|
# docs/contract-execution-pinning.md for which clone each contract pins.
|
||||||
|
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
REF="origin/master"
|
||||||
|
FETCH=1
|
||||||
|
QUIET=0
|
||||||
|
FETCH_TIMEOUT="${REVISION_PREFLIGHT_FETCH_TIMEOUT:-20}"
|
||||||
|
|
||||||
|
usage() {
|
||||||
|
sed -n '2,55p' "$0" | sed 's/^# \{0,1\}//'
|
||||||
|
}
|
||||||
|
|
||||||
|
while [[ $# -gt 0 ]]; do
|
||||||
|
case "$1" in
|
||||||
|
--ref)
|
||||||
|
[[ $# -ge 2 ]] || { echo "revision-preflight: --ref needs a value" >&2; exit 2; }
|
||||||
|
REF="$2"; shift 2 ;;
|
||||||
|
--no-fetch) FETCH=0; shift ;;
|
||||||
|
--fetch-timeout)
|
||||||
|
[[ $# -ge 2 ]] || { echo "revision-preflight: --fetch-timeout needs a value" >&2; exit 2; }
|
||||||
|
FETCH_TIMEOUT="$2"; shift 2 ;;
|
||||||
|
--quiet) QUIET=1; shift ;;
|
||||||
|
-h|--help) usage; exit 0 ;;
|
||||||
|
--) shift; break ;;
|
||||||
|
-*) echo "revision-preflight: unknown option: $1" >&2; exit 2 ;;
|
||||||
|
*) break ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
|
||||||
|
if [[ $# -lt 2 ]]; then
|
||||||
|
usage >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
SCRIPT="$1"
|
||||||
|
CLONE="$2"
|
||||||
|
|
||||||
|
say() { [[ $QUIET -eq 1 ]] || echo "$@" >&2; }
|
||||||
|
|
||||||
|
# refuse <class> <explanation...> -> the copy is not the merged one
|
||||||
|
refuse() {
|
||||||
|
local class="$1"; shift
|
||||||
|
echo "REASON=mismatch:${class}" >&2
|
||||||
|
echo "❌ revision-preflight: MISMATCH (${class}) — refusing to report from this copy" >&2
|
||||||
|
for line in "$@"; do echo " $line" >&2; done
|
||||||
|
exit 1
|
||||||
|
}
|
||||||
|
|
||||||
|
# unverifiable <class> <explanation...> -> the check could not be performed
|
||||||
|
unverifiable() {
|
||||||
|
local class="$1"; shift
|
||||||
|
echo "REASON=cannot-verify:${class}" >&2
|
||||||
|
echo "❌ revision-preflight: CANNOT VERIFY (${class}) — refusing to report unverified" >&2
|
||||||
|
for line in "$@"; do echo " $line" >&2; done
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
|
||||||
|
# ── 1. inputs must exist ──────────────────────────────────────────────────────
|
||||||
|
if [[ ! -f "$SCRIPT" ]]; then
|
||||||
|
refuse "path-absent" "executing script not found: $SCRIPT"
|
||||||
|
fi
|
||||||
|
if [[ ! -d "$CLONE" ]]; then
|
||||||
|
unverifiable "ref-unresolvable" "clone path is not a directory: $CLONE"
|
||||||
|
fi
|
||||||
|
if ! git -C "$CLONE" rev-parse --git-dir >/dev/null 2>&1; then
|
||||||
|
unverifiable "ref-unresolvable" "not a git clone: $CLONE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 2. resolve the repo-relative path (the original defect) ───────────────────
|
||||||
|
CLONE_ABS=$(cd "$CLONE" && pwd)
|
||||||
|
SCRIPT_ABS=$(cd "$(dirname "$SCRIPT")" && pwd)/$(basename "$SCRIPT")
|
||||||
|
case "$SCRIPT_ABS" in
|
||||||
|
"$CLONE_ABS"/*) REL="${SCRIPT_ABS#"$CLONE_ABS"/}" ;;
|
||||||
|
*) refuse "content" "script is outside the clone: $SCRIPT_ABS is not under $CLONE_ABS" ;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
# ── 3. refresh the ref, bounded, so a hung remote cannot block a contract ────
|
||||||
|
if [[ $FETCH -eq 1 ]]; then
|
||||||
|
REMOTE="${REF%%/*}"
|
||||||
|
[[ "$REMOTE" == "$REF" ]] && REMOTE="origin"
|
||||||
|
FETCH_CMD=(git -C "$CLONE" fetch --quiet "$REMOTE")
|
||||||
|
if [[ "$FETCH_TIMEOUT" != "0" ]]; then
|
||||||
|
if ! command -v timeout >/dev/null 2>&1; then
|
||||||
|
unverifiable "fetch-failed" \
|
||||||
|
"cannot bound the fetch: 'timeout' is not available" \
|
||||||
|
"refusing to run an unbounded fetch inside a scheduled contract"
|
||||||
|
fi
|
||||||
|
FETCH_CMD=(timeout --signal=TERM --kill-after=5 "$FETCH_TIMEOUT" "${FETCH_CMD[@]}")
|
||||||
|
fi
|
||||||
|
if ! "${FETCH_CMD[@]}" 2>/dev/null; then
|
||||||
|
unverifiable "fetch-failed" \
|
||||||
|
"could not fetch '$REMOTE' in $CLONE_ABS (bound: ${FETCH_TIMEOUT}s)" \
|
||||||
|
"cannot compare against a possibly stale '$REF'" \
|
||||||
|
"re-run with network access, raise --fetch-timeout, or pass --no-fetch deliberately"
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
say "⚠️ revision-preflight: --no-fetch — comparing against the LOCAL '$REF'; freshness is assumed, not verified"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 4. resolve the merged revision; unresolvable is a failure ────────────────
|
||||||
|
if ! git -C "$CLONE" rev-parse --verify --quiet "$REF" >/dev/null; then
|
||||||
|
unverifiable "ref-unresolvable" \
|
||||||
|
"ref '$REF' does not resolve in $CLONE_ABS" \
|
||||||
|
"the clone may never have fetched, or the ref name may be wrong"
|
||||||
|
fi
|
||||||
|
REF_COMMIT=$(git -C "$CLONE" rev-parse --short "$REF")
|
||||||
|
|
||||||
|
TMPFILE=$(mktemp)
|
||||||
|
trap 'rm -f "$TMPFILE"' EXIT
|
||||||
|
|
||||||
|
if ! git -C "$CLONE" show "$REF:$REL" > "$TMPFILE" 2>/dev/null; then
|
||||||
|
refuse "path-absent" \
|
||||||
|
"'$REL' does not exist in $REF ($REF_COMMIT)" \
|
||||||
|
"a path absent from $REF can never be a merged copy" \
|
||||||
|
"script: $SCRIPT_ABS"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 5. compare ───────────────────────────────────────────────────────────────
|
||||||
|
EXEC_SHA=$(sha256sum "$SCRIPT" | cut -d' ' -f1)
|
||||||
|
MERGED_SHA=$(sha256sum "$TMPFILE" | cut -d' ' -f1)
|
||||||
|
|
||||||
|
if [[ "$EXEC_SHA" != "$MERGED_SHA" ]]; then
|
||||||
|
DETAIL=("script: $SCRIPT_ABS"
|
||||||
|
"clone: $CLONE_ABS"
|
||||||
|
"executed: $EXEC_SHA"
|
||||||
|
"merged: $MERGED_SHA ($REF:$REL @ $REF_COMMIT)")
|
||||||
|
|
||||||
|
# Name WHY it differs: a detached HEAD or a branch legitimately ahead of the
|
||||||
|
# ref is a much more benign situation than a hand-edited file, and the
|
||||||
|
# operator must be able to tell them apart.
|
||||||
|
if ! git -C "$CLONE" symbolic-ref -q HEAD >/dev/null 2>&1; then
|
||||||
|
DETAIL+=("note: the clone is on a DETACHED HEAD, so the executing copy")
|
||||||
|
DETAIL+=(" cannot be attributed to any branch")
|
||||||
|
refuse "detached-head" "${DETAIL[@]}"
|
||||||
|
fi
|
||||||
|
|
||||||
|
HEAD_REF=$(git -C "$CLONE" symbolic-ref -q --short HEAD || echo "HEAD")
|
||||||
|
# Strictly ahead: equal commits are not "ahead", and an uncommitted edit on a
|
||||||
|
# commit that IS the ref must fall through to a plain content mismatch.
|
||||||
|
REF_OID=$(git -C "$CLONE" rev-parse "$REF" 2>/dev/null || echo "")
|
||||||
|
HEAD_OID=$(git -C "$CLONE" rev-parse HEAD 2>/dev/null || echo "")
|
||||||
|
if [[ -n "$REF_OID" && "$REF_OID" != "$HEAD_OID" ]] \
|
||||||
|
&& git -C "$CLONE" merge-base --is-ancestor "$REF" HEAD 2>/dev/null; then
|
||||||
|
AHEAD=$(git -C "$CLONE" rev-list --count "$REF..HEAD" 2>/dev/null || echo "?")
|
||||||
|
DETAIL+=("note: '$HEAD_REF' is AHEAD of $REF by $AHEAD commit(s)")
|
||||||
|
DETAIL+=(" (a legitimate mid-review state, not a hand-edited file)")
|
||||||
|
refuse "clone-ahead" "${DETAIL[@]}"
|
||||||
|
fi
|
||||||
|
|
||||||
|
DETAIL+=("branch: $HEAD_REF")
|
||||||
|
refuse "content" "${DETAIL[@]}"
|
||||||
|
fi
|
||||||
|
|
||||||
|
say "✅ revision-preflight: $REL matches $REF @ $REF_COMMIT ($EXEC_SHA)"
|
||||||
|
exit 0
|
||||||
Executable
+395
@@ -0,0 +1,395 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Agent-consumption layer in front of SearXNG + Firecrawl.
|
||||||
|
|
||||||
|
Multi-engine aggregation returns results with no dedupe, no filtering and no
|
||||||
|
reranking. Measured 2026-09-26 that put bestbuy.com and merriam-webster.com into
|
||||||
|
"best practices agent context management", and put four SEO blogs ABOVE the real
|
||||||
|
Proxmox forum threads on a precise technical query. Identical queries also ranked
|
||||||
|
differently between runs, so the fix has to be deterministic rather than
|
||||||
|
dependent on engine mood.
|
||||||
|
|
||||||
|
This module turns the raw result list into something an agent can actually use:
|
||||||
|
|
||||||
|
1. DEDUPE the same page arriving from several engines
|
||||||
|
2. DROP clear non-answers (homepages, shopping, dictionaries, logins)
|
||||||
|
3. DEMOTE config-listed low-authority hosts; PROMOTE primary sources
|
||||||
|
4. STABLE SORT so ordering is reproducible run to run
|
||||||
|
5. EXTRACT page text for the top N under an explicit character budget,
|
||||||
|
so one call returns usable material instead of a snippet
|
||||||
|
6. EMIT stable JSON with engine provenance
|
||||||
|
|
||||||
|
Policy lives in config/search-ranking.yaml, not in this file.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
search-agent-consume.py "query text" # JSON to stdout
|
||||||
|
search-agent-consume.py --no-extract "query" # ranking only, no Firecrawl
|
||||||
|
search-agent-consume.py --explain "query" # include drop/demote reasons
|
||||||
|
|
||||||
|
Exit: 0 ok, 1 no results survived filtering, 2 the layer could not run.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
SEARXNG_URL = os.environ.get("SEARXNG_URL", "http://192.168.68.7:8888").rstrip("/")
|
||||||
|
FIRECRAWL_URL = os.environ.get("FIRECRAWL_URL", "http://192.168.68.7:3002").rstrip("/")
|
||||||
|
CONFIG_PATH = os.environ.get(
|
||||||
|
"SEARCH_RANKING_CONFIG",
|
||||||
|
str(Path(__file__).resolve().parent.parent / "config" / "search-ranking.yaml"),
|
||||||
|
)
|
||||||
|
HTTP_TIMEOUT = float(os.environ.get("SEARCH_CONSUME_TIMEOUT", "25"))
|
||||||
|
|
||||||
|
|
||||||
|
def _load_config() -> dict:
|
||||||
|
"""Load the ranking policy.
|
||||||
|
|
||||||
|
PyYAML is used when present; otherwise a tiny built-in parser handles the
|
||||||
|
flat lists in this specific file, so the layer never hard-fails on a host
|
||||||
|
without PyYAML.
|
||||||
|
"""
|
||||||
|
text = Path(CONFIG_PATH).read_text()
|
||||||
|
try:
|
||||||
|
import yaml # type: ignore
|
||||||
|
|
||||||
|
return yaml.safe_load(text)
|
||||||
|
except ImportError:
|
||||||
|
return _parse_flat_yaml(text)
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_flat_yaml(text: str) -> dict:
|
||||||
|
"""Minimal fallback parser: top-level keys, nested one level, flat lists."""
|
||||||
|
import re
|
||||||
|
|
||||||
|
out: dict = {}
|
||||||
|
stack: list[tuple[int, dict]] = [(-1, out)]
|
||||||
|
section: dict | None = None
|
||||||
|
for raw in text.splitlines():
|
||||||
|
line = raw.split("#", 1)[0].rstrip()
|
||||||
|
if not line.strip():
|
||||||
|
continue
|
||||||
|
indent = len(line) - len(line.lstrip())
|
||||||
|
body = line.strip()
|
||||||
|
if body.startswith("- "):
|
||||||
|
if section is not None:
|
||||||
|
section.setdefault("_list", []).append(
|
||||||
|
body[2:].strip().strip("'\"")
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
if ":" in body:
|
||||||
|
key, _, val = body.partition(":")
|
||||||
|
key, val = key.strip(), val.strip()
|
||||||
|
if val:
|
||||||
|
# write to the INNERMOST open section, not the document root
|
||||||
|
stack[-1][1][key] = _scalar(val)
|
||||||
|
section = None
|
||||||
|
else:
|
||||||
|
while stack and indent <= stack[-1][0]:
|
||||||
|
stack.pop()
|
||||||
|
parent = stack[-1][1]
|
||||||
|
new: dict = {}
|
||||||
|
parent[key] = new
|
||||||
|
stack.append((indent, new))
|
||||||
|
section = new
|
||||||
|
# flatten "_list" holders back into their parent as plain lists
|
||||||
|
def fix(node):
|
||||||
|
if isinstance(node, dict):
|
||||||
|
if set(node.keys()) == {"_list"}:
|
||||||
|
return node["_list"]
|
||||||
|
return {k: fix(v) for k, v in node.items()}
|
||||||
|
return node
|
||||||
|
|
||||||
|
return fix(out)
|
||||||
|
|
||||||
|
|
||||||
|
def _scalar(v: str):
|
||||||
|
if v.lower() in ("true", "false"):
|
||||||
|
return v.lower() == "true"
|
||||||
|
try:
|
||||||
|
return int(v)
|
||||||
|
except ValueError:
|
||||||
|
pass
|
||||||
|
try:
|
||||||
|
return float(v)
|
||||||
|
except ValueError:
|
||||||
|
pass
|
||||||
|
return v.strip("'\"")
|
||||||
|
|
||||||
|
|
||||||
|
# ── filtering ────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def _host(url: str) -> str:
|
||||||
|
return (urllib.parse.urlparse(url).netloc or "").lower().split(":")[0]
|
||||||
|
|
||||||
|
|
||||||
|
def _registrable(host: str) -> str:
|
||||||
|
"""Best-effort registrable domain so sub.forum.proxmox.com matches proxmox.com."""
|
||||||
|
parts = host.split(".")
|
||||||
|
if len(parts) <= 2:
|
||||||
|
return host
|
||||||
|
# handle common two-label public suffixes
|
||||||
|
two = ".".join(parts[-2:])
|
||||||
|
if parts[-2] in ("co", "com", "org", "net", "ac", "gov") and len(parts) >= 3:
|
||||||
|
return ".".join(parts[-3:])
|
||||||
|
return two
|
||||||
|
|
||||||
|
|
||||||
|
def _host_in(host: str, domains) -> bool:
|
||||||
|
if not domains:
|
||||||
|
return False
|
||||||
|
reg = _registrable(host)
|
||||||
|
for d in domains:
|
||||||
|
d = str(d).lower()
|
||||||
|
if host == d or host.endswith("." + d) or reg == d:
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _normalise_url(url: str) -> str:
|
||||||
|
"""Strip tracking params and fragments so the same page dedupes."""
|
||||||
|
p = urllib.parse.urlparse(url)
|
||||||
|
q = [
|
||||||
|
(k, v)
|
||||||
|
for k, v in urllib.parse.parse_qsl(p.query, keep_blank_values=True)
|
||||||
|
if not k.lower().startswith(("utm_", "fbclid", "gclid", "mc_", "ref"))
|
||||||
|
]
|
||||||
|
path = p.path.rstrip("/") or "/"
|
||||||
|
return urllib.parse.urlunparse(
|
||||||
|
(p.scheme.lower(), p.netloc.lower(), path, "", urllib.parse.urlencode(q), "")
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def non_answer_reason(result: dict, cfg: dict) -> str | None:
|
||||||
|
"""Return why this result is a non-answer, or None if it may be returned."""
|
||||||
|
na = cfg.get("non_answer", {}) or {}
|
||||||
|
url = result.get("url", "")
|
||||||
|
p = urllib.parse.urlparse(url)
|
||||||
|
host = _host(url)
|
||||||
|
path = p.path or ""
|
||||||
|
|
||||||
|
if _host_in(host, na.get("hosts")):
|
||||||
|
return "shopping_or_dictionary_host"
|
||||||
|
|
||||||
|
if na.get("host_root", True) and path in ("", "/"):
|
||||||
|
# A preferred host's front door may legitimately be the answer
|
||||||
|
# (a repo, a docs site). Everything else is navigational.
|
||||||
|
if not _host_in(host, cfg.get("prefer_domains")):
|
||||||
|
return "navigational_host_root"
|
||||||
|
|
||||||
|
low = url.lower()
|
||||||
|
for pat in na.get("path_patterns", []) or []:
|
||||||
|
if pat.lower() in low:
|
||||||
|
return f"path_pattern:{pat}"
|
||||||
|
|
||||||
|
qkeys = {k.lower() for k in (na.get("query_keys") or [])}
|
||||||
|
if qkeys & {k.lower() for k, _ in urllib.parse.parse_qsl(p.query)}:
|
||||||
|
return "search_or_shopping_query"
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def source_type(url: str, cfg: dict) -> str:
|
||||||
|
host = _host(url)
|
||||||
|
if _host_in(host, ["github.com", "gitlab.com", "codeberg.org", "sourceforge.net"]):
|
||||||
|
return "code"
|
||||||
|
if _host_in(host, ["stackoverflow.com", "stackexchange.com", "superuser.com",
|
||||||
|
"serverfault.com", "askubuntu.com"]):
|
||||||
|
return "qa"
|
||||||
|
if _host_in(host, ["forum.proxmox.com", "forum.", "discourse"]) or "forum." in host:
|
||||||
|
return "forum"
|
||||||
|
if _host_in(host, ["news.ycombinator.com", "lobste.rs", "reddit.com"]):
|
||||||
|
return "discussion"
|
||||||
|
if _host_in(host, cfg.get("prefer_domains")):
|
||||||
|
return "official"
|
||||||
|
if _host_in(host, cfg.get("demote_domains")):
|
||||||
|
return "content-farm"
|
||||||
|
return "web"
|
||||||
|
|
||||||
|
|
||||||
|
def rank(results: list[dict], cfg: dict) -> tuple[list[dict], list[dict]]:
|
||||||
|
"""Dedupe, drop non-answers, demote/ promote, stable sort.
|
||||||
|
|
||||||
|
Returns (kept, dropped) where dropped carries the reason, because a filter
|
||||||
|
nobody can audit is a filter nobody should trust.
|
||||||
|
"""
|
||||||
|
rank_cfg = cfg.get("ranking", {}) or {}
|
||||||
|
demote_pen = float(rank_cfg.get("demote_penalty", 1000))
|
||||||
|
prefer_bonus = float(rank_cfg.get("prefer_bonus", 100))
|
||||||
|
multi_bonus = float(rank_cfg.get("multi_engine_bonus", 25))
|
||||||
|
|
||||||
|
seen: dict[str, dict] = {}
|
||||||
|
dropped: list[dict] = []
|
||||||
|
|
||||||
|
for pos, r in enumerate(results):
|
||||||
|
url = r.get("url")
|
||||||
|
if not url:
|
||||||
|
continue
|
||||||
|
key = _normalise_url(url)
|
||||||
|
engine = r.get("engine", "?")
|
||||||
|
|
||||||
|
# 1. dedupe: same normalised URL from several engines
|
||||||
|
if key in seen:
|
||||||
|
seen[key].setdefault("engines", []).append(engine)
|
||||||
|
seen[key]["duplicate_of"] = True
|
||||||
|
continue
|
||||||
|
|
||||||
|
reason = non_answer_reason(r, cfg)
|
||||||
|
if reason:
|
||||||
|
dropped.append({"url": url, "reason": reason, "position": pos + 1})
|
||||||
|
continue
|
||||||
|
|
||||||
|
seen[key] = {
|
||||||
|
"title": (r.get("title") or "").strip(),
|
||||||
|
"url": url,
|
||||||
|
"engines": [engine],
|
||||||
|
"position": pos,
|
||||||
|
"score": 0.0,
|
||||||
|
}
|
||||||
|
|
||||||
|
kept = []
|
||||||
|
for item in seen.values():
|
||||||
|
host = _host(item["url"])
|
||||||
|
score = -float(item["position"]) # original order is the base signal
|
||||||
|
if _host_in(host, cfg.get("demote_domains")):
|
||||||
|
score -= demote_pen
|
||||||
|
if _host_in(host, cfg.get("prefer_domains")):
|
||||||
|
score += prefer_bonus
|
||||||
|
if len(item["engines"]) > 1:
|
||||||
|
score += multi_bonus * (len(item["engines"]) - 1)
|
||||||
|
item["score"] = round(score, 2)
|
||||||
|
item["host"] = host
|
||||||
|
item["source_type"] = source_type(item["url"], cfg)
|
||||||
|
kept.append(item)
|
||||||
|
|
||||||
|
# stable: score desc, then original position asc => reproducible run to run
|
||||||
|
kept.sort(key=lambda i: (-i["score"], i["position"]))
|
||||||
|
return kept, dropped
|
||||||
|
|
||||||
|
|
||||||
|
# ── extraction ───────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def _post_json(url: str, payload: dict, timeout: float) -> dict:
|
||||||
|
req = urllib.request.Request(
|
||||||
|
url,
|
||||||
|
data=json.dumps(payload).encode(),
|
||||||
|
headers={"Content-Type": "application/json"},
|
||||||
|
)
|
||||||
|
with urllib.request.urlopen(req, timeout=timeout) as resp:
|
||||||
|
return json.loads(resp.read().decode("utf-8", "replace"))
|
||||||
|
|
||||||
|
|
||||||
|
def extract(items: list[dict], cfg: dict) -> dict:
|
||||||
|
"""Fetch page text for the top N under a global character budget."""
|
||||||
|
ex = cfg.get("extraction", {}) or {}
|
||||||
|
top_n = int(ex.get("top_n", 5))
|
||||||
|
total_budget = int(ex.get("total_chars", 12000))
|
||||||
|
per_item = int(ex.get("per_item_chars", 4000))
|
||||||
|
timeout = float(ex.get("timeout_seconds", 45))
|
||||||
|
|
||||||
|
used = 0
|
||||||
|
failures = 0
|
||||||
|
t0 = time.time()
|
||||||
|
for item in items[:top_n]:
|
||||||
|
remaining = total_budget - used
|
||||||
|
if remaining <= 200:
|
||||||
|
item["excerpt"] = ""
|
||||||
|
item["extraction"] = "skipped_budget_exhausted"
|
||||||
|
continue
|
||||||
|
cap = min(per_item, remaining)
|
||||||
|
try:
|
||||||
|
data = _post_json(
|
||||||
|
f"{FIRECRAWL_URL}/v1/scrape",
|
||||||
|
{"url": item["url"], "formats": ["markdown"]},
|
||||||
|
timeout,
|
||||||
|
)
|
||||||
|
md = ((data.get("data") or {}).get("markdown") or "").strip()
|
||||||
|
if not md:
|
||||||
|
item["excerpt"] = ""
|
||||||
|
item["extraction"] = "empty"
|
||||||
|
failures += 1
|
||||||
|
continue
|
||||||
|
item["excerpt"] = md[:cap]
|
||||||
|
item["extraction"] = "ok" if len(md) <= cap else "truncated"
|
||||||
|
used += len(item["excerpt"])
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
item["excerpt"] = ""
|
||||||
|
item["extraction"] = f"failed:{type(exc).__name__}"
|
||||||
|
failures += 1
|
||||||
|
return {
|
||||||
|
"extracted": min(top_n, len(items)),
|
||||||
|
"chars_used": used,
|
||||||
|
"budget": total_budget,
|
||||||
|
"failures": failures,
|
||||||
|
"seconds": round(time.time() - t0, 2),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# ── entry point ──────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def consume(query: str, do_extract: bool = True, explain: bool = False) -> dict:
|
||||||
|
cfg = _load_config()
|
||||||
|
url = f"{SEARXNG_URL}/search?" + urllib.parse.urlencode(
|
||||||
|
{"q": query, "format": "json"}
|
||||||
|
)
|
||||||
|
with urllib.request.urlopen(url, timeout=HTTP_TIMEOUT) as resp:
|
||||||
|
raw = json.loads(resp.read().decode("utf-8", "replace"))
|
||||||
|
|
||||||
|
results = raw.get("results", [])
|
||||||
|
kept, dropped = rank(results, cfg)
|
||||||
|
extraction = extract(kept, cfg) if do_extract else None
|
||||||
|
|
||||||
|
out = {
|
||||||
|
"query": query,
|
||||||
|
"raw_result_count": len(results),
|
||||||
|
"returned_count": len(kept),
|
||||||
|
"dropped_count": len(dropped),
|
||||||
|
"engines": sorted({r.get("engine", "?") for r in results}),
|
||||||
|
"results": [
|
||||||
|
{
|
||||||
|
"rank": i + 1,
|
||||||
|
"title": it["title"],
|
||||||
|
"url": it["url"],
|
||||||
|
"host": it["host"],
|
||||||
|
"source_type": it["source_type"],
|
||||||
|
"engines": sorted(set(it["engines"])),
|
||||||
|
"score": it["score"],
|
||||||
|
"excerpt": it.get("excerpt", ""),
|
||||||
|
"extraction": it.get("extraction", "not_attempted"),
|
||||||
|
}
|
||||||
|
for i, it in enumerate(kept)
|
||||||
|
],
|
||||||
|
"extraction": extraction,
|
||||||
|
}
|
||||||
|
if explain:
|
||||||
|
out["dropped"] = dropped
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
args = [a for a in sys.argv[1:] if not a.startswith("--")]
|
||||||
|
do_extract = "--no-extract" not in sys.argv
|
||||||
|
explain = "--explain" in sys.argv
|
||||||
|
if not args:
|
||||||
|
print(__doc__)
|
||||||
|
return 2
|
||||||
|
query = " ".join(args)
|
||||||
|
try:
|
||||||
|
out = consume(query, do_extract=do_extract, explain=explain)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
print(f"LAYER FAILED: {type(exc).__name__}: {exc}", file=sys.stderr)
|
||||||
|
return 2
|
||||||
|
print(json.dumps(out, indent=2))
|
||||||
|
return 0 if out["returned_count"] else 1
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
Executable
+313
@@ -0,0 +1,313 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Search-stack visibility check.
|
||||||
|
|
||||||
|
The fleet shares one SearXNG instance (search) plus one extraction service
|
||||||
|
(Firecrawl). A broken search stack used to fail silently: one engine answered
|
||||||
|
and nobody could tell that the other engines had stopped contributing, or that
|
||||||
|
an enabled engine was returning nothing at all without reporting an error.
|
||||||
|
|
||||||
|
This check makes those failures visible and non-zero:
|
||||||
|
|
||||||
|
* runs two fixed queries against SearXNG; FAILS when fewer than two engines
|
||||||
|
contribute to a query, printing the contributing engines and every
|
||||||
|
``unresponsive_engines`` entry;
|
||||||
|
* FAILS when a known page cannot be extracted to non-empty markdown through
|
||||||
|
Firecrawl;
|
||||||
|
* reports every *silent zero* engine explicitly -- an engine that is enabled,
|
||||||
|
is eligible for the query category, is not listed in
|
||||||
|
``unresponsive_engines``, and still contributed no results.
|
||||||
|
|
||||||
|
Exit code 0 = healthy, 1 = degraded, 2 = the check could not run at all.
|
||||||
|
|
||||||
|
Environment overrides (all optional):
|
||||||
|
SEARXNG_URL default http://192.168.68.7:8888
|
||||||
|
FIRECRAWL_URL default http://192.168.68.7:3002
|
||||||
|
SEARCH_CHECK_QUERIES comma-separated fixed queries
|
||||||
|
SEARCH_CHECK_MIN_ENGINES default 2
|
||||||
|
SEARCH_CHECK_TIMEOUT per-request timeout in seconds, default 25
|
||||||
|
SEARCH_CHECK_EXTRACT_URL page used for the extraction leg
|
||||||
|
SEARCH_CHECK_ENGINES comma-separated engine names the stack is expected to
|
||||||
|
run; a silent zero is reported for any of them that is
|
||||||
|
enabled but contributes nothing with no error
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
SEARXNG_URL = os.environ.get("SEARXNG_URL", "http://192.168.68.7:8888").rstrip("/")
|
||||||
|
FIRECRAWL_URL = os.environ.get("FIRECRAWL_URL", "http://192.168.68.7:3002").rstrip("/")
|
||||||
|
QUERIES = [
|
||||||
|
q.strip()
|
||||||
|
for q in os.environ.get(
|
||||||
|
"SEARCH_CHECK_QUERIES", "proxmox backup server,python asyncio tutorial"
|
||||||
|
).split(",")
|
||||||
|
if q.strip()
|
||||||
|
]
|
||||||
|
MIN_ENGINES = int(os.environ.get("SEARCH_CHECK_MIN_ENGINES", "2"))
|
||||||
|
TIMEOUT = float(os.environ.get("SEARCH_CHECK_TIMEOUT", "25"))
|
||||||
|
EXTRACT_URL = os.environ.get(
|
||||||
|
"SEARCH_CHECK_EXTRACT_URL", "https://en.wikipedia.org/wiki/Proxmox_Virtual_Environment"
|
||||||
|
)
|
||||||
|
|
||||||
|
# The general web-search engines this stack intentionally runs. A general query
|
||||||
|
# is expected to draw on these; an enabled one that returns nothing without an
|
||||||
|
# error is the silent-zero failure this check exists to expose. Specialised
|
||||||
|
# engines (images, videos, translate, currency, arxiv, npm, ...) are excluded on
|
||||||
|
# purpose -- contributing nothing to a general query is correct for them.
|
||||||
|
DEFAULT_EXPECTED_ENGINES = [
|
||||||
|
"bing",
|
||||||
|
"brave",
|
||||||
|
"google cse",
|
||||||
|
"yandex",
|
||||||
|
"duckduckgo",
|
||||||
|
]
|
||||||
|
EXPECTED_ENGINES = [
|
||||||
|
e.strip()
|
||||||
|
for e in os.environ.get(
|
||||||
|
"SEARCH_CHECK_ENGINES", ",".join(DEFAULT_EXPECTED_ENGINES)
|
||||||
|
).split(",")
|
||||||
|
if e.strip()
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def _get_json(url: str) -> dict:
|
||||||
|
req = urllib.request.Request(url, headers={"User-Agent": "search-stack-check/1.0"})
|
||||||
|
with urllib.request.urlopen(req, timeout=TIMEOUT) as resp:
|
||||||
|
return json.loads(resp.read().decode("utf-8", "replace"))
|
||||||
|
|
||||||
|
|
||||||
|
def _post_json(url: str, payload: dict) -> dict:
|
||||||
|
data = json.dumps(payload).encode("utf-8")
|
||||||
|
req = urllib.request.Request(
|
||||||
|
url,
|
||||||
|
data=data,
|
||||||
|
headers={
|
||||||
|
"Content-Type": "application/json",
|
||||||
|
"User-Agent": "search-stack-check/1.0",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
with urllib.request.urlopen(req, timeout=TIMEOUT) as resp:
|
||||||
|
return json.loads(resp.read().decode("utf-8", "replace"))
|
||||||
|
|
||||||
|
|
||||||
|
def enabled_expected_engines() -> set[str]:
|
||||||
|
"""Expected engines that SearXNG reports as actually enabled."""
|
||||||
|
cfg = _get_json(f"{SEARXNG_URL}/config")
|
||||||
|
enabled = {e["name"] for e in cfg.get("engines", []) if e.get("enabled")}
|
||||||
|
return {name for name in EXPECTED_ENGINES if name in enabled}
|
||||||
|
|
||||||
|
|
||||||
|
def unresponsive_names(pairs: list) -> dict[str, str]:
|
||||||
|
"""``unresponsive_engines`` is a list of [name, reason] pairs (or strings)."""
|
||||||
|
out: dict[str, str] = {}
|
||||||
|
for item in pairs or []:
|
||||||
|
if isinstance(item, (list, tuple)) and len(item) >= 2:
|
||||||
|
out[str(item[0])] = str(item[1])
|
||||||
|
elif isinstance(item, str):
|
||||||
|
out[item] = "unresponsive"
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
# ── QUALITY GUARD (search-agent-consumption) ─────────────────────────────────
|
||||||
|
# The agent-consumption layer applies a deterministic demote/drop policy. Without
|
||||||
|
# an assertion here it could silently rot back to raw engine ordering - the same
|
||||||
|
# way the endpoint colours silently rotted before 2026-09-26.
|
||||||
|
QUALITY_QUERIES = [
|
||||||
|
"best practices agent context management",
|
||||||
|
"proxmox thin pool metadata exhaustion recovery",
|
||||||
|
]
|
||||||
|
# A demoted (content-farm) host must never occupy the top 3 for these queries.
|
||||||
|
QUALITY_TOP_N = 3
|
||||||
|
# Non-answers that must never be returned for these queries at all.
|
||||||
|
QUALITY_BANNED_HOSTS = ["bestbuy.com", "merriam-webster.com"]
|
||||||
|
|
||||||
|
|
||||||
|
def _consumption_layer_path():
|
||||||
|
here = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
return os.path.join(here, "search-agent-consume.py")
|
||||||
|
|
||||||
|
|
||||||
|
def check_ranking_quality() -> list[str]:
|
||||||
|
"""Return a list of quality failures; empty means healthy."""
|
||||||
|
import subprocess as _sp
|
||||||
|
|
||||||
|
layer = _consumption_layer_path()
|
||||||
|
if not os.path.exists(layer):
|
||||||
|
return [f"agent-consumption layer missing: {layer}"]
|
||||||
|
|
||||||
|
failures: list[str] = []
|
||||||
|
for query in QUALITY_QUERIES:
|
||||||
|
r = _sp.run([sys.executable, layer, "--no-extract", "--explain", query],
|
||||||
|
capture_output=True, text=True, timeout=120)
|
||||||
|
if r.returncode != 0:
|
||||||
|
failures.append(f"{query!r}: layer exited {r.returncode} ({r.stderr[:120]})")
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
data = json.loads(r.stdout)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
failures.append(f"{query!r}: layer returned unparseable JSON")
|
||||||
|
continue
|
||||||
|
|
||||||
|
results = data.get("results", [])
|
||||||
|
if len(results) < QUALITY_TOP_N:
|
||||||
|
failures.append(f"{query!r}: only {len(results)} results returned")
|
||||||
|
continue
|
||||||
|
|
||||||
|
# load the demote list from the SAME config the layer uses
|
||||||
|
cfg_path = os.path.join(os.path.dirname(layer), "..", "config", "search-ranking.yaml")
|
||||||
|
demoted: set[str] = set()
|
||||||
|
try:
|
||||||
|
sys.path.insert(0, os.path.dirname(layer))
|
||||||
|
import importlib.util as _iu
|
||||||
|
spec = _iu.spec_from_file_location("_sac_cfg", layer)
|
||||||
|
mod = _iu.module_from_spec(spec)
|
||||||
|
spec.loader.exec_module(mod)
|
||||||
|
demoted = set(mod._load_config().get("demote_domains", []) or [])
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
failures.append(f"{query!r}: could not load demote_domains from config")
|
||||||
|
|
||||||
|
for item in results[:QUALITY_TOP_N]:
|
||||||
|
host = (item.get("host") or "")
|
||||||
|
for d in demoted:
|
||||||
|
if host == d or host.endswith("." + d):
|
||||||
|
failures.append(
|
||||||
|
f"{query!r}: demoted host {host} in top {QUALITY_TOP_N}"
|
||||||
|
)
|
||||||
|
for item in results:
|
||||||
|
host = (item.get("host") or "")
|
||||||
|
for b in QUALITY_BANNED_HOSTS:
|
||||||
|
if host == b or host.endswith("." + b):
|
||||||
|
failures.append(f"{query!r}: non-answer host {host} returned")
|
||||||
|
return failures
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
failures: list[str] = []
|
||||||
|
print(f"Search stack check -- {SEARXNG_URL}")
|
||||||
|
print(f"Queries: {QUERIES!r} min contributing engines: {MIN_ENGINES}")
|
||||||
|
print("=" * 72)
|
||||||
|
|
||||||
|
try:
|
||||||
|
eligible = enabled_expected_engines()
|
||||||
|
except Exception as exc: # noqa: BLE001 - report, do not traceback
|
||||||
|
print(f"FAIL: could not read /config from SearXNG: {exc!r}")
|
||||||
|
return 2
|
||||||
|
print(f"Expected engines, enabled ({len(eligible)}): {sorted(eligible)}")
|
||||||
|
missing = sorted(set(EXPECTED_ENGINES) - eligible)
|
||||||
|
if missing:
|
||||||
|
print(f"Expected engines NOT enabled: {missing}")
|
||||||
|
failures.append(f"expected engines not enabled in SearXNG: {missing}")
|
||||||
|
|
||||||
|
contributed: dict[str, int] = {name: 0 for name in eligible}
|
||||||
|
silent_zero_all: dict[str, list[str]] = {}
|
||||||
|
|
||||||
|
for query in QUERIES:
|
||||||
|
url = f"{SEARXNG_URL}/search?" + urllib.parse.urlencode(
|
||||||
|
{"q": query, "format": "json"}
|
||||||
|
)
|
||||||
|
print("-" * 72)
|
||||||
|
print(f"QUERY: {query!r}")
|
||||||
|
try:
|
||||||
|
data = _get_json(url)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
print(f" FAIL: query request failed: {exc!r}")
|
||||||
|
failures.append(f"query {query!r} request failed: {exc!r}")
|
||||||
|
continue
|
||||||
|
|
||||||
|
results = data.get("results", [])
|
||||||
|
engines: dict[str, int] = {}
|
||||||
|
for r in results:
|
||||||
|
name = r.get("engine", "?")
|
||||||
|
engines[name] = engines.get(name, 0) + 1
|
||||||
|
unresponsive = unresponsive_names(data.get("unresponsive_engines", []))
|
||||||
|
|
||||||
|
print(f" results: {len(results)}")
|
||||||
|
print(f" contributing engines: {engines or '(none)'}")
|
||||||
|
print(f" unresponsive_engines: {unresponsive or '(none)'}")
|
||||||
|
|
||||||
|
for name in engines:
|
||||||
|
contributed[name] = contributed.get(name, 0) + engines[name]
|
||||||
|
|
||||||
|
if len(engines) < MIN_ENGINES:
|
||||||
|
msg = (
|
||||||
|
f"query {query!r} had only {len(engines)} contributing engine(s) "
|
||||||
|
f"({sorted(engines)}); need >= {MIN_ENGINES}"
|
||||||
|
)
|
||||||
|
print(f" FAIL: {msg}")
|
||||||
|
failures.append(msg)
|
||||||
|
|
||||||
|
silent = sorted(
|
||||||
|
n for n in eligible if n not in engines and n not in unresponsive
|
||||||
|
)
|
||||||
|
if silent:
|
||||||
|
silent_zero_all[query] = silent
|
||||||
|
print(
|
||||||
|
" SILENT ZERO (enabled, no error, no results -- reported, "
|
||||||
|
f"not fatal): {silent}"
|
||||||
|
)
|
||||||
|
|
||||||
|
print("=" * 72)
|
||||||
|
print("Engine contribution across all queries:")
|
||||||
|
for name in sorted(contributed):
|
||||||
|
status = "ZERO" if contributed[name] == 0 else "ok"
|
||||||
|
print(f" {name:<24} {contributed[name]:>4} {status}")
|
||||||
|
|
||||||
|
if silent_zero_all:
|
||||||
|
print("-" * 72)
|
||||||
|
print("SILENT-ZERO ENGINES REPORTED (no error raised, no results returned):")
|
||||||
|
for query, names in silent_zero_all.items():
|
||||||
|
print(f" {query!r}: {names}")
|
||||||
|
print(" NOTE: a silent zero is REPORTED, not counted as a failure. These")
|
||||||
|
print(" engines are expected to answer a general query, but contributing")
|
||||||
|
print(" nothing to one query can be legitimate (result de-duplication, or")
|
||||||
|
print(" an engine that only fires on certain query shapes). Only the")
|
||||||
|
print(f" <{MIN_ENGINES}-contributing-engine floor and the extraction leg fail the run.")
|
||||||
|
|
||||||
|
print("-" * 72)
|
||||||
|
print(f"EXTRACTION: scraping {EXTRACT_URL} via {FIRECRAWL_URL}/v1/scrape")
|
||||||
|
try:
|
||||||
|
payload = _post_json(
|
||||||
|
f"{FIRECRAWL_URL}/v1/scrape",
|
||||||
|
{"url": EXTRACT_URL, "formats": ["markdown"]},
|
||||||
|
)
|
||||||
|
markdown = ((payload.get("data") or {}).get("markdown") or "").strip()
|
||||||
|
if not markdown:
|
||||||
|
msg = "extraction returned empty markdown"
|
||||||
|
print(f" FAIL: {msg}")
|
||||||
|
failures.append(msg)
|
||||||
|
else:
|
||||||
|
print(f" ok: {len(markdown)} chars of markdown returned")
|
||||||
|
print(f" first line: {markdown.splitlines()[0][:120]!r}")
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
msg = f"extraction request failed: {exc!r}"
|
||||||
|
print(f" FAIL: {msg}")
|
||||||
|
failures.append(msg)
|
||||||
|
|
||||||
|
print("-" * 72)
|
||||||
|
print("RANKING QUALITY (agent-consumption layer)")
|
||||||
|
quality = check_ranking_quality()
|
||||||
|
if quality:
|
||||||
|
for q in quality:
|
||||||
|
print(f" FAIL: {q}")
|
||||||
|
failures.extend(quality)
|
||||||
|
else:
|
||||||
|
print(" ok: no demoted host in the top 3; no banned non-answer returned")
|
||||||
|
|
||||||
|
print("=" * 72)
|
||||||
|
if failures:
|
||||||
|
print("VERDICT: FAIL")
|
||||||
|
for f in failures:
|
||||||
|
print(f" - {f}")
|
||||||
|
return 1
|
||||||
|
print("VERDICT: PASS -- multiple engines contributing, extraction healthy")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
@@ -0,0 +1,55 @@
|
|||||||
|
# secret-allowlist.tsv — exceptions for scripts/secret-scan.sh, every entry with a reason.
|
||||||
|
#
|
||||||
|
# Format: <rule-id|*><TAB><path-glob><TAB><literal-substring><TAB><reason>
|
||||||
|
# Blank lines and lines whose first field starts with '#' are ignored.
|
||||||
|
# A finding is suppressed only when ALL THREE of rule, path and literal match:
|
||||||
|
# * the rule id equals the finding's rule id, or is '*'
|
||||||
|
# * the finding's repo-relative path matches <path-glob> (bash glob)
|
||||||
|
# * the finding's line contains <literal-substring> verbatim
|
||||||
|
# An entry whose reason is empty is a hard error (exit 2) — no silent exceptions.
|
||||||
|
#
|
||||||
|
# RULE: never allowlist a live credential, and never broaden an entry (rule '*',
|
||||||
|
# a wide path glob, or a short generic literal) just to silence a finding.
|
||||||
|
# If the finding is real, remove the credential from the file.
|
||||||
|
#
|
||||||
|
# Entries are one per deliberate synthetic example, so the file reads as an
|
||||||
|
# audit trail of reviewed exceptions rather than a list of things to ignore.
|
||||||
|
# Rule '*' is used only where the same literal is matched by more than one rule.
|
||||||
|
#
|
||||||
|
# ── The 2026-09-17 purge placeholders ─────────────────────────────────────
|
||||||
|
# PR #112 replaced six live credentials with `«vault: <project>/<env> <SECRET>»`
|
||||||
|
# references. Those references are safe by construction (they name where the
|
||||||
|
# secret is read from), but they are listed here explicitly rather than being
|
||||||
|
# filtered by a general "vault" rule, so a new occurrence still needs a
|
||||||
|
# deliberate, reasoned entry.
|
||||||
|
secret-assign litellm-api-keys.prose.md MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY» 2026-09-17 purge: replaced the live Mumuni LiteLLM key with its vault reference; no literal credential.
|
||||||
|
secret-assign litellm-api-keys.prose.md MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY» 2026-09-17 purge: replaced the live Mumuni Zulip key with its vault reference; no literal credential.
|
||||||
|
* infrastructure-control.prose.md PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
|
||||||
|
cred-prose infrastructure-control.prose.md Admin credentials: 2026-09-17 purge: the Stirling admin user/password are two `«vault: ...»` references; no literal credential.
|
||||||
|
* scripts/daily-infra-report.py PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
|
||||||
|
secret-assign stirling-pdf-agent-access.prose.md «vault: infrastructure/production STIRLING_API_KEY» 2026-09-17 purge: Stirling PDF API key is read from the vault; the curl example only names the vault path.
|
||||||
|
bearer-token agent-zero-fix-summary.md «vault: agents/production OPENROUTER_API_KEY» 2026-09-17 purge: OpenRouter key is read from the vault; the example curl only names the vault path.
|
||||||
|
# ── Deliberate synthetic examples in contracts (not from the purge) ───────
|
||||||
|
# These exist to teach the rule they illustrate. They are listed here so the
|
||||||
|
# guard is never taught to skip the words "synthetic"/"example" — a fabricated
|
||||||
|
# example is always an explicit exception, never a pattern-level exemption.
|
||||||
|
* hermes-key-enforcement.prose.md sk-synthetic-external-example Rule 15 illustration of a hardcoded external key that is tolerated; fabricated, never a live key.
|
||||||
|
* hermes-key-enforcement.prose.md sk-synthetic-example-12345 Rule 15 illustration of a forbidden hardcoded key; fabricated, never a live key.
|
||||||
|
openai-key hermes-key-enforcement.prose.md sk-synthetic-litellm- Fabricated key name inside a `grep 'LITELLM_API_KEY=...'` example; not a live key.
|
||||||
|
secret-assign hermes-key-enforcement.prose.md sk-NEW_KEY Placeholder standing for the rotated key in an `infisical secrets set` command; not a literal key.
|
||||||
|
openrouter-key agent-zero-openrouter-key.prose.md sk-or-v1-synthetic Synthetic key prefix in the contract's example response; the real key is read from the vault.
|
||||||
|
openai-key litellm-api-keys.prose.md sk-synthetic-tanko-example Fabricated key name in migration history prose; not a live key.
|
||||||
|
openai-key litellm-self-heal.prose.md sk-syslog-local-master-key Deprecated local LiteLLM master key name documented as no-live-usage; kept for history, not a usable credential.
|
||||||
|
# ── Redacted evidence, not a credential ──────────────────────────────────
|
||||||
|
secret-assign docs/probe-drift-round2-evidence.md =sk-... Probe evidence records redacted key trailers (`sk-...x6uw`); the usable part of the key is not present.
|
||||||
|
# ── tests/test_secret_scan.sh fixtures ───────────────────────────────────
|
||||||
|
# The self-test plants these fabricated values into a TEMP tree, whose path no
|
||||||
|
# entry here covers, so each still fails the guard when planted (see the test's
|
||||||
|
# "... fails the guard" cases). They are listed only so the repo-wide scan of
|
||||||
|
# the test file itself stays quiet.
|
||||||
|
* tests/test_secret_scan.sh sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef Self-test fixture: fabricated OpenRouter-shaped key written to a temp tree; the guard must fail on it there.
|
||||||
|
bearer-token tests/test_secret_scan.sh Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb Self-test fixture: fabricated Bearer token written to a temp tree; the guard must fail on it there.
|
||||||
|
proxmox-token tests/test_secret_scan.sh PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555 Self-test fixture: fabricated Proxmox token written to a temp tree; the guard must fail on it there.
|
||||||
|
private-key tests/test_secret_scan.sh -----BEGIN OPENSSH PRIVATE KEY----- Self-test fixture: fabricated PEM banner written to a temp tree; the guard must fail on it there.
|
||||||
|
cred-prose tests/test_secret_scan.sh Admin credentials: Self-test fixture: fabricated prose credential line written to a temp tree; the guard must fail on it there.
|
||||||
|
secret-assign tests/test_secret_scan.sh DB_PASSWORD=correct-horse-battery-staple Self-test fixture: fabricated password assignment written to a temp tree; the guard must fail on it there.
|
||||||
|
Can't render this file because it contains an unexpected character in line 23 and column 25.
|
@@ -0,0 +1,22 @@
|
|||||||
|
# secret-patterns.tsv — checked-in pattern list for scripts/secret-scan.sh
|
||||||
|
#
|
||||||
|
# Format: <rule-id><TAB><POSIX ERE><TAB><description><TAB><check>
|
||||||
|
# Blank lines and lines whose first field starts with '#' are ignored.
|
||||||
|
# <check> is optional; the only value today is "value", which tells the scanner
|
||||||
|
# to run the matched value through its inert-value classifier (see
|
||||||
|
# value_is_inert in secret-scan.sh) so bare identifiers, env refs and dotted
|
||||||
|
# code access are not reported as credentials. Omit the column to report every
|
||||||
|
# regex hit.
|
||||||
|
# Matching is case-insensitive, so `API_KEY` and `api_key` both count.
|
||||||
|
#
|
||||||
|
# Add a rule here, never inline in secret-scan.sh: this file is the single
|
||||||
|
# auditable list of what the guard considers credential-shaped.
|
||||||
|
openai-key \bsk-[A-Za-z0-9_-]{16,} OpenAI/LiteLLM-style "sk-" secret key (also hyphenated sk-proj- keys)
|
||||||
|
openrouter-key \bsk-or-v1-[A-Za-z0-9_-]{8,} OpenRouter API key
|
||||||
|
stripe-live-key \bsk_live_[A-Za-z0-9]{8,} Stripe live secret key
|
||||||
|
proxmox-token PVEAPIToken=[^[:space:]"']+ Proxmox API token literal
|
||||||
|
bearer-token bearer[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) literal Bearer token (http header or prose)
|
||||||
|
auth-header authorization:[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) Authorization header carrying a raw literal value
|
||||||
|
private-key -----BEGIN [A-Z ]*PRIVATE KEY----- PEM private key block
|
||||||
|
cred-prose credentials?[[:space:]]*[:=][[:space:]]*[^[:space:]] prose credential line carrying a value
|
||||||
|
secret-assign (api[_-]?key|apikey|passwd|password|secret|token)s?["']?[[:space:]]*[:=][[:space:]]*["']?(«.{3,}»|[A-Za-z0-9_./+=-]{8,}) credential assignment carrying a literal value value
|
||||||
|
Can't render this file because it contains an unexpected character in line 5 and column 48.
|
Executable
+269
@@ -0,0 +1,269 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# secret-scan.sh — commit-time secret guard. FAILS (exit 1) on a credential-shaped
|
||||||
|
# string, so a build cannot go green with a credential committed to it.
|
||||||
|
#
|
||||||
|
# Usage:
|
||||||
|
# scripts/secret-scan.sh # scan the whole git-tracked tree (default)
|
||||||
|
# scripts/secret-scan.sh --tree
|
||||||
|
# scripts/secret-scan.sh --path DIR # scan an arbitrary directory (git not required)
|
||||||
|
# scripts/secret-scan.sh --staged # scan added lines in the index (pre-commit)
|
||||||
|
# scripts/secret-scan.sh --diff REF # scan added lines since REF (e.g. origin/master)
|
||||||
|
# --quiet only print the verdict and findings, no per-mode banner
|
||||||
|
#
|
||||||
|
# Exit codes: 0 clean, 1 credential found, 2 usage/config error.
|
||||||
|
#
|
||||||
|
# Patterns live in scripts/secret-patterns.tsv
|
||||||
|
# Exceptions live in scripts/secret-allowlist.tsv (every entry carries a reason;
|
||||||
|
# a missing reason is a hard error, so the guard fails closed).
|
||||||
|
#
|
||||||
|
# Dependencies are deliberately bash + coreutils + grep + sed/awk + git. The
|
||||||
|
# Gitea Actions runner executes job steps INSIDE the runner container, which
|
||||||
|
# has no node and no python by default: keep this script free of both.
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
SELF_DIR=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||||
|
ROOT=$(cd -- "$SELF_DIR/.." && pwd)
|
||||||
|
PATTERNS_FILE="$SELF_DIR/secret-patterns.tsv"
|
||||||
|
ALLOWLIST_FILE="$SELF_DIR/secret-allowlist.tsv"
|
||||||
|
|
||||||
|
# The guard's own definition files are not scannable content: the pattern list
|
||||||
|
# necessarily contains the pattern text, and the allowlist necessarily contains
|
||||||
|
# the allowed literals. Narrow, exact-path exclusion — not a wildcard.
|
||||||
|
SELF_FILES=(
|
||||||
|
"scripts/secret-scan.sh"
|
||||||
|
"scripts/secret-patterns.tsv"
|
||||||
|
"scripts/secret-allowlist.tsv"
|
||||||
|
)
|
||||||
|
|
||||||
|
MODE="tree"
|
||||||
|
PATH_DIR=""
|
||||||
|
DIFF_REF=""
|
||||||
|
QUIET=0
|
||||||
|
|
||||||
|
usage() {
|
||||||
|
sed -n '2,20p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
|
||||||
|
while [ $# -gt 0 ]; do
|
||||||
|
case "$1" in
|
||||||
|
--tree) MODE="tree" ;;
|
||||||
|
--path) MODE="path"; PATH_DIR="${2:-}"; shift ;;
|
||||||
|
--staged) MODE="staged" ;;
|
||||||
|
--diff) MODE="diff"; DIFF_REF="${2:-}"; shift ;;
|
||||||
|
--quiet) QUIET=1 ;;
|
||||||
|
-h|--help) usage ;;
|
||||||
|
*) echo "secret-scan: unknown argument '$1'" >&2; usage ;;
|
||||||
|
esac
|
||||||
|
shift
|
||||||
|
done
|
||||||
|
|
||||||
|
[ -f "$PATTERNS_FILE" ] || { echo "secret-scan: missing $PATTERNS_FILE" >&2; exit 2; }
|
||||||
|
[ -f "$ALLOWLIST_FILE" ] || { echo "secret-scan: missing $ALLOWLIST_FILE" >&2; exit 2; }
|
||||||
|
if [ "$MODE" = "path" ] && [ -z "$PATH_DIR" ]; then
|
||||||
|
echo "secret-scan: --path needs a directory" >&2; exit 2
|
||||||
|
fi
|
||||||
|
if [ "$MODE" = "diff" ] && [ -z "$DIFF_REF" ]; then
|
||||||
|
echo "secret-scan: --diff needs a base ref" >&2; exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Load patterns ──────────────────────────────────────────────────────────
|
||||||
|
RULE_IDS=()
|
||||||
|
RULE_RES=()
|
||||||
|
RULE_DESCS=()
|
||||||
|
RULE_CHECKS=()
|
||||||
|
COMBINED=""
|
||||||
|
while IFS=$'\t' read -r id re desc check; do
|
||||||
|
case "$id" in ''|'#'*) continue ;; esac
|
||||||
|
[ -n "$re" ] || continue
|
||||||
|
RULE_IDS+=("$id"); RULE_RES+=("$re"); RULE_DESCS+=("$desc"); RULE_CHECKS+=("${check:-}")
|
||||||
|
if [ -z "$COMBINED" ]; then COMBINED="($re)"; else COMBINED="$COMBINED|($re)"; fi
|
||||||
|
done < "$PATTERNS_FILE"
|
||||||
|
if [ "${#RULE_IDS[@]}" -eq 0 ]; then
|
||||||
|
echo "secret-scan: no patterns loaded from $PATTERNS_FILE" >&2; exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Load allowlist (fails closed on a missing reason) ──────────────────────
|
||||||
|
AL_RULES=()
|
||||||
|
AL_GLOBS=()
|
||||||
|
AL_LITS=()
|
||||||
|
AL_REASONS=()
|
||||||
|
AL_LINENO=0
|
||||||
|
while IFS=$'\t' read -r rule glob lit reason; do
|
||||||
|
AL_LINENO=$((AL_LINENO + 1))
|
||||||
|
case "$rule" in ''|'#'*) continue ;; esac
|
||||||
|
if [ -z "$glob" ] || [ -z "$lit" ] || [ -z "$reason" ]; then
|
||||||
|
echo "secret-scan: ❌ $ALLOWLIST_FILE:$AL_LINENO — allowlist entry needs <rule> <path-glob> <literal> <reason>; reason-based exceptions only, refusing to run" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
AL_RULES+=("$rule"); AL_GLOBS+=("$glob"); AL_LITS+=("$lit"); AL_REASONS+=("$reason")
|
||||||
|
done < "$ALLOWLIST_FILE"
|
||||||
|
|
||||||
|
# nocasematch is toggled only around the regex test; path globs must stay
|
||||||
|
# case-sensitive, so it is never left on.
|
||||||
|
MATCH=""
|
||||||
|
regex_match() { # regex_match <regex> <text> -> MATCH holds the matched text
|
||||||
|
local re="$1" text="$2"
|
||||||
|
shopt -s nocasematch
|
||||||
|
if [[ $text =~ $re ]]; then
|
||||||
|
MATCH="${BASH_REMATCH[0]}"
|
||||||
|
shopt -u nocasematch
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
shopt -u nocasematch
|
||||||
|
MATCH=""
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
allowlisted() { # allowlisted <rule> <path> <text>
|
||||||
|
local rule="$1" path="$2" text="$3" i
|
||||||
|
for i in "${!AL_RULES[@]}"; do
|
||||||
|
[ "${AL_RULES[$i]}" = "$rule" ] || [ "${AL_RULES[$i]}" = "*" ] || continue
|
||||||
|
# The unquoted RHS is deliberate: <path-glob> is a bash glob, not a literal.
|
||||||
|
# shellcheck disable=SC2053
|
||||||
|
[[ $path == ${AL_GLOBS[$i]} ]] || continue
|
||||||
|
[[ $text == *"${AL_LITS[$i]}"* ]] || continue
|
||||||
|
return 0
|
||||||
|
done
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
mask_value() { # mask_value <text> <match> — never echo a credential to logs.
|
||||||
|
# Print only the part of the line BEFORE the match, then <redacted>: the match
|
||||||
|
# itself and everything after it (which may include a value the rule's regex
|
||||||
|
# stopped short of, e.g. `credentials:` followed by a backticked password) is
|
||||||
|
# never written to stdout.
|
||||||
|
local text="$1" m="$2"
|
||||||
|
if [ -n "$m" ] && [[ $text == *"$m"* ]]; then
|
||||||
|
printf '%s<redacted>' "${text%%"$m"*}"
|
||||||
|
else
|
||||||
|
printf '%s' "$text"
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
FINDINGS=0
|
||||||
|
SUPPRESSED=0
|
||||||
|
INERT=0
|
||||||
|
SCANNED=0
|
||||||
|
|
||||||
|
# value_is_inert <value> <text-after-match> — true when a matched assignment value
|
||||||
|
# is plainly not a credential: empty, an env/command reference, a path, dotted
|
||||||
|
# code access, a short or single-class identifier (a variable or key NAME, not a
|
||||||
|
# value), a well-known placeholder word, or a value the file deliberately
|
||||||
|
# truncates with '…' / '...' (a redacted prefix is not a usable credential).
|
||||||
|
# Deliberately does NOT know the words "synthetic" or "example": a fabricated
|
||||||
|
# example must be an explicit allowlist entry.
|
||||||
|
value_is_inert() {
|
||||||
|
local v="$1" rest="$2"
|
||||||
|
case "$rest" in '…'*|'...'*) return 0 ;; esac
|
||||||
|
v="${v%\"}"; v="${v#\"}"; v="${v%\'}"; v="${v#\'}"
|
||||||
|
case "$v" in
|
||||||
|
''|\$*|\{*|'<'*|'%'*|'('*|'/'*|'\\'*) return 0 ;;
|
||||||
|
not-needed|no-key-required|none|null|true|false|redacted|placeholder|example|dummy|changeme|change-me|your-key|your_key|key|token|secret|password) return 0 ;;
|
||||||
|
esac
|
||||||
|
# dotted code access: os.environ.get / process.env.ZULIP_API_KEY / cfg.a
|
||||||
|
if [[ $v =~ ^[a-z_][a-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)+$ ]]; then return 0; fi
|
||||||
|
# bare identifier (no punctuation beyond _): a NAME, not a value. A real
|
||||||
|
# secret in this shape is long and mixes letters with digits.
|
||||||
|
if [[ $v =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
|
||||||
|
[ "${#v}" -lt 20 ] && return 0
|
||||||
|
[[ $v =~ [0-9] ]] || return 0
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
report_finding() { # report_finding <path> <line> <text>
|
||||||
|
local path="$1" line="$2" text="$3" i val
|
||||||
|
for i in "${!RULE_IDS[@]}"; do
|
||||||
|
regex_match "${RULE_RES[$i]}" "$text" || continue
|
||||||
|
SCANNED=$((SCANNED + 1))
|
||||||
|
if [ "${RULE_CHECKS[$i]}" = "value" ]; then
|
||||||
|
val="${MATCH#*[:=]}"
|
||||||
|
val="${val# }"
|
||||||
|
if value_is_inert "$val" "${text#*"$MATCH"}"; then
|
||||||
|
INERT=$((INERT + 1))
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
if allowlisted "${RULE_IDS[$i]}" "$path" "$text"; then
|
||||||
|
SUPPRESSED=$((SUPPRESSED + 1))
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
FINDINGS=$((FINDINGS + 1))
|
||||||
|
printf ' ❌ %s:%s [%s] %s\n' "$path" "$line" "${RULE_IDS[$i]}" "${RULE_DESCS[$i]}"
|
||||||
|
printf ' | %s\n' "$(mask_value "$text" "$MATCH")"
|
||||||
|
done
|
||||||
|
}
|
||||||
|
|
||||||
|
self_excluded() { # self_excluded <repo-relative-path>
|
||||||
|
local p="$1" s
|
||||||
|
for s in "${SELF_FILES[@]}"; do
|
||||||
|
[ "$p" = "$s" ] && return 0
|
||||||
|
done
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
# ── Collect candidate lines and scan them ─────────────────────────────────
|
||||||
|
if [ "$MODE" = "tree" ] || [ "$MODE" = "path" ]; then
|
||||||
|
if [ "$MODE" = "tree" ]; then
|
||||||
|
BASE="$ROOT"
|
||||||
|
git -C "$BASE" rev-parse --git-dir >/dev/null 2>&1 || { echo "secret-scan: --tree needs a git checkout (use --path DIR)" >&2; exit 2; }
|
||||||
|
mapfile -d '' candidate < <(git -C "$BASE" ls-files -z 2>/dev/null)
|
||||||
|
if [ "${#candidate[@]}" -eq 0 ]; then
|
||||||
|
echo "secret-scan: ❌ no tracked files — refusing to report clean" >&2; exit 2
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
BASE=$(cd -- "$PATH_DIR" 2>/dev/null && pwd) || { echo "secret-scan: --path '$PATH_DIR' is not a directory" >&2; exit 2; }
|
||||||
|
mapfile -t candidate < <(cd -- "$BASE" && find . -type f -not -path './.git/*' | sed 's|^\./||')
|
||||||
|
if [ "${#candidate[@]}" -eq 0 ]; then
|
||||||
|
echo "secret-scan: ❌ no files under $BASE — refusing to report clean" >&2; exit 2
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
[ "$QUIET" -eq 1 ] || echo "── secret scan ($MODE): ${#candidate[@]} files under $BASE ──"
|
||||||
|
for rel in "${candidate[@]}"; do
|
||||||
|
[ -f "$BASE/$rel" ] || continue
|
||||||
|
self_excluded "$rel" && continue
|
||||||
|
while IFS= read -r hit; do
|
||||||
|
[ -n "$hit" ] || continue
|
||||||
|
report_finding "$rel" "${hit%%:*}" "${hit#*:}"
|
||||||
|
done < <(grep -nEIi -e "$COMBINED" "$BASE/$rel" 2>/dev/null || true)
|
||||||
|
done
|
||||||
|
else
|
||||||
|
# --staged / --diff: only ADDED lines, with the post-change line number.
|
||||||
|
if [ "$MODE" = "staged" ]; then
|
||||||
|
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines in the index ──"
|
||||||
|
DIFF_TEXT=$(git -C "$ROOT" diff --cached --unified=0 --no-color -- . 2>/dev/null)
|
||||||
|
else
|
||||||
|
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines since $DIFF_REF ──"
|
||||||
|
DIFF_TEXT=$(git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"...HEAD 2>/dev/null \
|
||||||
|
|| git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"..HEAD 2>/dev/null)
|
||||||
|
fi
|
||||||
|
if [ -z "$DIFF_TEXT" ]; then
|
||||||
|
[ "$QUIET" -eq 1 ] || echo " (no added lines)"
|
||||||
|
fi
|
||||||
|
while IFS=$'\t' read -r rel line text; do
|
||||||
|
[ -n "$rel" ] || continue
|
||||||
|
self_excluded "$rel" && continue
|
||||||
|
report_finding "$rel" "$line" "$text"
|
||||||
|
done < <(printf '%s\n' "$DIFF_TEXT" | awk '
|
||||||
|
/^\+\+\+ / { f=$2; sub(/^b\//,"",f); next }
|
||||||
|
/^@@ / { if (match($0, /\+[0-9]+/)) ln=substr($0, RSTART+1, RLENGTH-1)+0; next }
|
||||||
|
(/^\+/ && !/^\+\+\+/) { print f "\t" ln "\t" substr($0,2); ln++; next }
|
||||||
|
')
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Verdict ────────────────────────────────────────────────────────────────
|
||||||
|
if [ "$FINDINGS" -gt 0 ]; then
|
||||||
|
echo ""
|
||||||
|
echo "❌ SECRET SCAN FAILED — $FINDINGS credential-shaped string(s) in ${MODE} content."
|
||||||
|
echo " Fix: remove the credential and read it from the vault/env."
|
||||||
|
echo " Only a deliberate synthetic example may be added to scripts/secret-allowlist.tsv,"
|
||||||
|
echo " one entry per file/rule/literal, with a reason. Never allowlist a live credential."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "✅ secret scan clean (${MODE}; ${SUPPRESSED} allowlisted exception(s), ${INERT} inert value(s) ignored)"
|
||||||
|
exit 0
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
#!/bin/bash
|
#!/bin/bash
|
||||||
# swap-gpu-dense-model.sh — Swap RTX 3090 from qwen3.6-27B-code to SmartCode-Fable-5
|
# swap-gpu-dense-model.sh — Swap RTX 3090 from qwen3.6-27B-code to SmartCode-Fable-5
|
||||||
# Run when download completes: ssh root@192.168.68.8 'bash -s' < this script
|
# Run when download completes: ssh llmuser@192.168.68.8 'sudo bash -s' < this script
|
||||||
#
|
#
|
||||||
# Usage: bash swap-gpu-dense-model.sh
|
# Usage: bash swap-gpu-dense-model.sh
|
||||||
# Requires: new model at /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf
|
# Requires: new model at /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf
|
||||||
|
|||||||
Executable
+228
@@ -0,0 +1,228 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# test_infra_monitoring.sh — Asserts that probe calls use the documented targets.
|
||||||
|
#
|
||||||
|
# Strategy: stub curl and ssh on PATH to capture the exact arguments each leg
|
||||||
|
# builds, then assert the URL/port of every call. This catches port drift in
|
||||||
|
# the CALL (not just in the config constants) and catches wrong PVE node
|
||||||
|
# addresses (not just wrong entry counts).
|
||||||
|
#
|
||||||
|
# Run: bash scripts/test_infra_monitoring.sh
|
||||||
|
# Exits 0 if all assertions pass, 1 otherwise.
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
SCRIPT="${SCRIPT_DIR}/infra-monitoring.sh"
|
||||||
|
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
|
||||||
|
assert() {
|
||||||
|
local desc="$1" condition="$2"
|
||||||
|
if eval "$condition"; then
|
||||||
|
echo " ✅ $desc"
|
||||||
|
PASS=$((PASS+1))
|
||||||
|
else
|
||||||
|
echo " 🔴 $desc"
|
||||||
|
FAIL=$((FAIL+1))
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
echo "=== test_infra_monitoring.sh ==="
|
||||||
|
echo ""
|
||||||
|
|
||||||
|
# ── Stub curl: capture argv to a file, return 200 ──────────────────────────
|
||||||
|
STUB_DIR=$(mktemp -d)
|
||||||
|
trap 'rm -rf "$STUB_DIR"' EXIT
|
||||||
|
|
||||||
|
# Stub curl: first arg after flags is the URL; capture all args
|
||||||
|
cat > "$STUB_DIR/curl" << 'STUBEOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo "$@" >> "${CURL_STUB_LOG:-/dev/null}"
|
||||||
|
# Print 200 for %{http_code}
|
||||||
|
printf '%s\n' "200"
|
||||||
|
exit 0
|
||||||
|
STUBEOF
|
||||||
|
chmod +x "$STUB_DIR/curl"
|
||||||
|
|
||||||
|
# Stub ssh: first arg after options is the remote command; capture it
|
||||||
|
cat > "$STUB_DIR/ssh" << 'SSTUBEOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo "SSH $@" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||||
|
# The last arg is the remote command — extract and log curl args
|
||||||
|
for arg in "$@"; do
|
||||||
|
if [[ "$arg" == curl* ]]; then
|
||||||
|
echo "$arg" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
printf '%s\n' "200"
|
||||||
|
exit 0
|
||||||
|
SSTUBEOF
|
||||||
|
chmod +x "$STUB_DIR/ssh"
|
||||||
|
|
||||||
|
# ── Run the monitor with stubs ────────────────────────────────────────────
|
||||||
|
CURL_LOG="$STUB_DIR/curl_calls.log"
|
||||||
|
SSH_LOG="$STUB_DIR/ssh_calls.log"
|
||||||
|
touch "$CURL_LOG" "$SSH_LOG"
|
||||||
|
|
||||||
|
CURL_STUB_LOG="$CURL_LOG" SSH_STUB_LOG="$SSH_LOG" \
|
||||||
|
PATH="$STUB_DIR:$PATH" bash "$SCRIPT" > "$STUB_DIR/output.txt" 2>&1
|
||||||
|
|
||||||
|
# ── 1. Port drift detection (from actual curl invocations) ─────────────────
|
||||||
|
|
||||||
|
assert "Grafana probed at port 3001" \
|
||||||
|
'grep -q "http://192.168.68.116:3001/api/health" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "Prometheus probed at port 9090" \
|
||||||
|
'grep -q "http://192.168.68.116:9090/-/healthy" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "LiteLLM probed via nginx at port 80" \
|
||||||
|
'grep -q "http://192.168.68.116:80/litellm/health" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE API probed at port 8006" \
|
||||||
|
'grep -q ":8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "GPU exporter probed at port 9400" \
|
||||||
|
'grep -q ":9400/metrics" "$CURL_LOG"'
|
||||||
|
|
||||||
|
# ── 2. PVE API: exact node addresses (catches wrong IPs) ──────────────────
|
||||||
|
# Each real PVE node must be probed; CT 116 must NOT be in the PVE set
|
||||||
|
|
||||||
|
assert "PVE acerpve 192.168.68.9 probed" \
|
||||||
|
'grep -q "https://192.168.68.9:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE minipve 192.168.68.12 probed" \
|
||||||
|
'grep -q "https://192.168.68.12:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE storepve 192.168.68.6 probed" \
|
||||||
|
'grep -q "https://192.168.68.6:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE amdpve 192.168.68.15 probed" \
|
||||||
|
'grep -q "https://192.168.68.15:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE ocupve 192.168.68.5 probed" \
|
||||||
|
'grep -q "https://192.168.68.5:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
# CT 116 (.116) must NOT appear as a PVE API target
|
||||||
|
assert "CT 116 (.116) NOT probed as PVE API node" \
|
||||||
|
'! grep -q "https://192.168.68.116:8006" "$CURL_LOG"'
|
||||||
|
|
||||||
|
# ── 3. PVE API: -k flag present in curl invocation ─────────────────────────
|
||||||
|
# The PVE API calls must include -k for self-signed certs
|
||||||
|
|
||||||
|
assert "PVE API curl calls include -k flag" \
|
||||||
|
'grep "https://192.168.68.9:8006" "$CURL_LOG" | grep -q -- "-k"'
|
||||||
|
|
||||||
|
# ── 4. Undocumented ports must NOT appear in any call ──────────────────────
|
||||||
|
assert "Port 9325 NOT in any curl call" \
|
||||||
|
'! grep -q ":9325" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "Port 9405 NOT in any curl call" \
|
||||||
|
'! grep -q ":9405" "$CURL_LOG"'
|
||||||
|
|
||||||
|
# ── 5. Docker Stats / PVE Exporter: SSH-probed at correct ports ────────────
|
||||||
|
assert "Docker Stats probed at port 9324 via SSH" \
|
||||||
|
'grep -q "9324" "$SSH_LOG"'
|
||||||
|
|
||||||
|
assert "PVE Exporter probed at port 9221 via SSH" \
|
||||||
|
'grep -q "9221" "$SSH_LOG"'
|
||||||
|
|
||||||
|
# ── 5a. Per-leg assertions (proves which leg owns which port) ──────────────
|
||||||
|
# The script source must show Docker Stats using $DOCKER_STATS_PORT and
|
||||||
|
# PVE Exporter using $PVE_EXPORTER_PORT in the correct leg sections
|
||||||
|
assert "Docker Stats leg uses DOCKER_STATS_PORT constant" \
|
||||||
|
'grep -A 3 "# 6. Docker Stats" "$SCRIPT" | grep -q "\$DOCKER_STATS_PORT"'
|
||||||
|
|
||||||
|
assert "PVE Exporter leg uses PVE_EXPORTER_PORT constant" \
|
||||||
|
'grep -A 3 "# 7. PVE Exporter" "$SCRIPT" | grep -q "\$PVE_EXPORTER_PORT"'
|
||||||
|
|
||||||
|
# Verify the constants themselves are set to the correct values
|
||||||
|
assert "DOCKER_STATS_PORT constant set to 9324" \
|
||||||
|
'grep -q "^DOCKER_STATS_PORT=\"9324\"" "$SCRIPT"'
|
||||||
|
|
||||||
|
assert "PVE_EXPORTER_PORT constant set to 9221" \
|
||||||
|
'grep -q "^PVE_EXPORTER_PORT=\"9221\"" "$SCRIPT"'
|
||||||
|
|
||||||
|
# ── 5b. Stale port 9323 (dockerd) NOT probed ──────────────────────────────
|
||||||
|
assert "Port 9323 (dockerd) NOT in SSH log" \
|
||||||
|
'! grep -q "9323" "$SSH_LOG"'
|
||||||
|
|
||||||
|
# ── 6. No stale ports in the script source (belt-and-suspenders) ──────────
|
||||||
|
assert "Port 9325 (historical) NOT in script source" \
|
||||||
|
'! grep -q "9325" "$SCRIPT"'
|
||||||
|
|
||||||
|
assert "Port 9405 (historical) NOT in script source" \
|
||||||
|
'! grep -q "9405" "$SCRIPT"'
|
||||||
|
|
||||||
|
# ── 7. Failure-line content includes non-empty kind ────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||||
|
|
||||||
|
# Test 7a: Unexpected status (500) → kind should be unexpected:500
|
||||||
|
cat > "$TMP_DIR/curl" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
# Stub: return 500 for Grafana port, 200 otherwise
|
||||||
|
for arg in "$@"; do
|
||||||
|
if [[ "$arg" == *":3001"* ]]; then
|
||||||
|
echo "500"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
echo "200"
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/curl"
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||||
|
assert "Grafana failure line exists (unexpected status)" \
|
||||||
|
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||||
|
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||||
|
assert "Grafana failure kind is non-empty (unexpected status)" \
|
||||||
|
'[[ -n "$KIND" ]]'
|
||||||
|
|
||||||
|
# Test 7b: TLS error (000 + exit 60) → kind should be tls
|
||||||
|
cat > "$TMP_DIR/curl" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
# Stub: return 000 and exit 60 for Grafana port (TLS error) on ALL invocations
|
||||||
|
for arg in "$@"; do
|
||||||
|
if [[ "$arg" == *":3001"* ]]; then
|
||||||
|
echo "000"
|
||||||
|
exit 60
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
echo "200"
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/curl"
|
||||||
|
# Also stub ssh to return 000 + exit 60 for the retry
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
for arg in "$@"; do
|
||||||
|
if [[ "$arg" == curl* ]]; then
|
||||||
|
echo "000"
|
||||||
|
exit 60
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
echo "200"
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
export PATH="$TMP_DIR:$PATH"
|
||||||
|
OUT=$(bash "$SCRIPT" 2>&1)
|
||||||
|
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||||
|
assert "Grafana failure line exists (TLS error)" \
|
||||||
|
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||||
|
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||||
|
assert "Grafana failure kind is tls" \
|
||||||
|
'[[ "$KIND" == "tls" ]]'
|
||||||
|
|
||||||
|
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||||
|
echo ""
|
||||||
|
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||||
|
if [ $FAIL -gt 0 ]; then
|
||||||
|
echo " 🔴 TESTS FAILED"
|
||||||
|
exit 1
|
||||||
|
else
|
||||||
|
echo " ✅ ALL TESTS PASSED"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
Executable
+125
@@ -0,0 +1,125 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# test_proxmox_monitor.sh — Stub-driven tests for PBS GC liveness leg
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
SCRIPT="$(cd "$(dirname "$0")" && pwd)/proxmox-monitor.sh"
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
|
||||||
|
# ── Helpers ────────────────────────────────────────────────────────────────
|
||||||
|
assert() {
|
||||||
|
local desc="$1" cond="$2"
|
||||||
|
if eval "$cond" 2>/dev/null; then
|
||||||
|
echo " ✅ $desc"
|
||||||
|
PASS=$((PASS+1))
|
||||||
|
else
|
||||||
|
echo " 🔴 $desc"
|
||||||
|
FAIL=$((FAIL+1))
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
# ── 1. Healthy: fresh GC, 0 B pending ─────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
FRESH_ENDTIME=$(( $(date -u +%s) - (1 * 3600) )) # 1 hour ago
|
||||||
|
cat > "$TMP_DIR/ssh" << EOF
|
||||||
|
#!/bin/bash
|
||||||
|
# Stub: return valid JSON with fresh endtime
|
||||||
|
echo '[{"store":"storepve-datastore","last-run-endtime":$FRESH_ENDTIME,"pending-bytes":0}]'
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Healthy: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Healthy: shows healthy verdict" '[[ "$PBS_LINE" == *"healthy"* ]]'
|
||||||
|
assert "Healthy: shows pending-bytes 0 B" '[[ "$PBS_LINE" == *"pending-bytes: 0 B"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 2. Stale: GC >48h old ─────────────────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
OLD_ENDTIME=$(( $(date -u +%s) - (49 * 3600) )) # 49 hours ago
|
||||||
|
cat > "$TMP_DIR/ssh" << EOF
|
||||||
|
#!/bin/bash
|
||||||
|
echo '[{"store":"storepve-datastore","last-run-endtime":$OLD_ENDTIME,"pending-bytes":1048576}]'
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Stale: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Stale: shows stale verdict" '[[ "$PBS_LINE" == *"stale"* ]]'
|
||||||
|
assert "Stale: shows pending-bytes 1048576 B" '[[ "$PBS_LINE" == *"pending-bytes: 1048576 B"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 3. Probe-failed: empty output ─────────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Probe-failed (empty): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Probe-failed (empty): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 4. Probe-failed: unparseable output ───────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo "proxmox-backup-manager: command not found"
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Probe-failed (unparseable): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Probe-failed (unparseable): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 5. Null endtime: never-run ────────────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo '[{"store":"storepve-datastore","last-run-endtime":null,"pending-bytes":0}]'
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Null endtime: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Null endtime: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 6. Datastore absent ───────────────────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo '[{"store":"s3-archive","last-run-endtime":null,"pending-bytes":0}]'
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Datastore absent: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Datastore absent: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── Summary ────────────────────────────────────────────────────────────────
|
||||||
|
echo ""
|
||||||
|
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||||
|
if [ $FAIL -gt 0 ]; then
|
||||||
|
echo " 🔴 TESTS FAILED"
|
||||||
|
exit 1
|
||||||
|
else
|
||||||
|
echo " ✅ ALL TESTS PASSED"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
+79
-27
@@ -7,11 +7,22 @@
|
|||||||
# agent leg is retired — see the note after the Tanko leg.
|
# agent leg is retired — see the note after the Tanko leg.
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
|
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
|
||||||
|
# Never fall back to a literal key.
|
||||||
|
# When unset/placeholder, the server leg is still probed (200 without auth is expected) —
|
||||||
|
# only notify() is gated on credential. The pi/Tanko/kagentz
|
||||||
|
# legs do not need the Zulip API key. The placeholder is captain-held:
|
||||||
|
# zulip-health-credential-placeholder-20260913.
|
||||||
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||||
ZULIP_SITE="https://chat.sysloggh.net"
|
ZULIP_SITE="https://chat.sysloggh.net"
|
||||||
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||||
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
|
||||||
OWNER_ZULIP_ID="9"
|
OWNER_ZULIP_ID="9"
|
||||||
|
|
||||||
|
# Track whether the Zulip API credential is usable
|
||||||
|
ZULIP_CRED_OK=1
|
||||||
|
if [ -z "$ZULIP_API_KEY" ] || [[ "$ZULIP_API_KEY" == *"placeholder"* ]] || [[ "$ZULIP_API_KEY" == *"REDACTED"* ]]; then
|
||||||
|
ZULIP_CRED_OK=0
|
||||||
|
fi
|
||||||
|
|
||||||
LOG="/root/zulip-health-monitor.log"
|
LOG="/root/zulip-health-monitor.log"
|
||||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
@@ -22,26 +33,32 @@ notify() {
|
|||||||
local severity="$1" msg="$2"
|
local severity="$1" msg="$2"
|
||||||
echo "[$severity] $msg"
|
echo "[$severity] $msg"
|
||||||
|
|
||||||
# Zulip DM to owner
|
# Zulip DM to owner (skip if no credential)
|
||||||
local content="${severity} Zulip Monitor: ${msg}"
|
if [ "$ZULIP_CRED_OK" -eq 1 ]; then
|
||||||
local form
|
local content="${severity} Zulip Monitor: ${msg}"
|
||||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
local form
|
||||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
-d "${form}" > /dev/null 2>&1 || true
|
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
-d "${form}" > /dev/null 2>&1 || true
|
||||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||||
> /dev/null 2>&1 \
|
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
|
> /dev/null 2>&1 \
|
||||||
|
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)">> "$LOG"
|
||||||
|
else
|
||||||
|
echo " ALERT SUPPRESSED (no credential): ${severity} ${msg}" >> "$LOG"
|
||||||
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
# ── Global: Zulip Server ──
|
# ── Global: Zulip Server ──
|
||||||
|
# F3: Always probe server regardless of credential — 200 without auth is expected
|
||||||
|
# (verified live: server_settings returns 200 with no credential or wrong key).
|
||||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||||
https://chat.sysloggh.net/api/v1/server_settings \
|
https://chat.sysloggh.net/api/v1/server_settings \
|
||||||
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null) || SERVER_CODE="000"
|
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
|
||||||
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
|
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
|
||||||
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
|
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
|
||||||
if [ "$SERVER_CODE" != "200" ]; then
|
if [ "$SERVER_CODE" != "200" ]; then
|
||||||
@@ -118,16 +135,16 @@ case "$PI_VERDICT" in
|
|||||||
esac
|
esac
|
||||||
# -- abiba-leg-end
|
# -- abiba-leg-end
|
||||||
|
|
||||||
# ── Platform B: Tanko (DSH dsh-web on amdpve CT 112) ──
|
# ── Platform B: Tanko (DSH dsh-web on minipve CT 112) ──
|
||||||
# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker
|
# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker
|
||||||
# key availability varies — so probes run from the amdpve vantage via `pct exec`.
|
# key availability varies — so probes run from the minipve vantage via `pct exec`.
|
||||||
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on amdpve
|
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on minipve
|
||||||
# (192.168.68.15). The gateway binds 127.0.0.1:3080 loopback-only by design — a
|
# (192.168.68.12). The gateway binds 127.0.0.1:3080 loopback-only by design — a
|
||||||
# remote :3080 probe is refused and is NOT a fault.
|
# remote :3080 probe is refused and is NOT a fault.
|
||||||
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.12 \
|
||||||
"pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true)
|
"pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true)
|
||||||
[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown"
|
[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown"
|
||||||
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.12 \
|
||||||
"pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true)
|
"pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true)
|
||||||
[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000"
|
[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000"
|
||||||
|
|
||||||
@@ -159,6 +176,11 @@ fi
|
|||||||
# never contact her former host.
|
# never contact her former host.
|
||||||
|
|
||||||
# ── Platform C: Agent Zero (kagentz) ──
|
# ── Platform C: Agent Zero (kagentz) ──
|
||||||
|
# C1: A2A liveness (no credential needed) — probes the container's internal :80/a2a/
|
||||||
|
# C2: A2A response verification (needs LITELLM_KEY) — probes POST /a2a with auth
|
||||||
|
# C3: Public access path (no credential needed) — probes https://kagentz.sysloggh.net/
|
||||||
|
|
||||||
|
# C1: A2A liveness (container-internal probe)
|
||||||
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||||
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
|
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
|
||||||
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||||
@@ -167,23 +189,53 @@ AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
|||||||
if [ "$AZ_A2A_CODE" = "000" ]; then
|
if [ "$AZ_A2A_CODE" = "000" ]; then
|
||||||
notify "🔴" "kagentz A2A server DOWN (connection failed)"
|
notify "🔴" "kagentz A2A server DOWN (connection failed)"
|
||||||
ISSUES=$((ISSUES + 1))
|
ISSUES=$((ISSUES + 1))
|
||||||
echo " kagentz: ❌ A2A down (HTTP 000)" >> "$LOG"
|
echo " kagentz C1: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||||
else
|
else
|
||||||
case "$AZ_A2A_CODE" in
|
case "$AZ_A2A_CODE" in
|
||||||
200|401)
|
200|401)
|
||||||
echo " kagentz: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
echo " kagentz C1: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||||
*)
|
*)
|
||||||
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
|
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
|
||||||
ISSUES=$((ISSUES + 1))
|
ISSUES=$((ISSUES + 1))
|
||||||
echo " kagentz: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
echo " kagentz C1: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||||
|
esac
|
||||||
|
fi
|
||||||
|
|
||||||
|
# C3: Public access path (the captain's point of view)
|
||||||
|
# Probes the public URL that NetBird proxies to the container. 200/302/401 = alive,
|
||||||
|
# 502 = proxy's "upstream refused" page (incident), connection failed = incident.
|
||||||
|
# Never restarts anything — the contract forbids restarting the platform.
|
||||||
|
KAGENTZ_PUBLIC_CODE=$(curl -s -o /dev/null --connect-timeout 10 --max-time 15 \
|
||||||
|
-w '%{http_code}' https://kagentz.sysloggh.net/ 2>/dev/null) || KAGENTZ_PUBLIC_CODE="000"
|
||||||
|
KAGENTZ_PUBLIC_CODE=$(printf '%s' "$KAGENTZ_PUBLIC_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$KAGENTZ_PUBLIC_CODE" ] || KAGENTZ_PUBLIC_CODE="000"
|
||||||
|
|
||||||
|
if [ "$KAGENTZ_PUBLIC_CODE" = "000" ]; then
|
||||||
|
notify "🔴" "kagentz public URL DOWN (connection failed)"
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " kagentz C3: ❌ public URL down (HTTP 000)" >> "$LOG"
|
||||||
|
elif [ "$KAGENTZ_PUBLIC_CODE" = "502" ]; then
|
||||||
|
notify "🔴" "kagentz public URL 502 (upstream refused)"
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " kagentz C3: ❌ public URL 502 (upstream refused)" >> "$LOG"
|
||||||
|
else
|
||||||
|
case "$KAGENTZ_PUBLIC_CODE" in
|
||||||
|
200|302|401)
|
||||||
|
echo " kagentz C3: ✅ public URL alive (HTTP $KAGENTZ_PUBLIC_CODE)" >> "$LOG" ;;
|
||||||
|
*)
|
||||||
|
notify "🟡" "kagentz public URL answered HTTP $KAGENTZ_PUBLIC_CODE — running, unexpected status"
|
||||||
|
ISSUES=$((ISSUES + 1))
|
||||||
|
echo " kagentz C3: 🟡 public URL unexpected http=$KAGENTZ_PUBLIC_CODE (running, warning)" >> "$LOG" ;;
|
||||||
esac
|
esac
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# ── Summary ──
|
# ── Summary ──
|
||||||
|
# The run verdict is non-optimistic: when ISSUES > 0, the run is an INCIDENT.
|
||||||
|
# The lane must quote this Result line verbatim in its status report.
|
||||||
if [ "$ISSUES" -eq 0 ]; then
|
if [ "$ISSUES" -eq 0 ]; then
|
||||||
echo " Result: ✅ All healthy" >> "$LOG"
|
echo " Result: ✅ 0 issues (all healthy)" >> "$LOG"
|
||||||
else
|
else
|
||||||
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
echo " Result: 🔴 INCIDENT — $ISSUES issue(s) found" >> "$LOG"
|
||||||
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
||||||
fi
|
fi
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,137 @@
|
|||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: search-agent-consumption
|
||||||
|
description: >
|
||||||
|
Agent-consumption layer in front of SearXNG + Firecrawl. Raw multi-engine
|
||||||
|
aggregation returns results with no dedupe, no filtering and no reranking;
|
||||||
|
measured 2026-09-26 that put bestbuy.com and merriam-webster.com into "best
|
||||||
|
practices agent context management", and put four SEO blogs above the real
|
||||||
|
Proxmox forum threads on a precise technical query. Identical queries also
|
||||||
|
ranked DIFFERENTLY between runs, which is why the layer is deterministic
|
||||||
|
rather than dependent on engine behaviour.
|
||||||
|
|
||||||
|
Pipeline: dedupe -> drop non-answers -> demote content farms / promote primary
|
||||||
|
sources -> stable sort -> extract page text for the top N under an explicit
|
||||||
|
character budget -> stable JSON. Policy lives in config, not code.
|
||||||
|
|
||||||
|
Call it when an agent needs search RESULTS rather than links: it returns usable
|
||||||
|
page text in one call instead of a snippet plus a second fetch.
|
||||||
|
|
||||||
|
version: 1.0.0
|
||||||
|
---
|
||||||
|
|
||||||
|
## Where the policy lives
|
||||||
|
|
||||||
|
`config/search-ranking.yaml` — reviewable, no code change needed to adjust:
|
||||||
|
|
||||||
|
| key | effect |
|
||||||
|
| --- | --- |
|
||||||
|
| `non_answer.hosts` / `path_patterns` / `query_keys` / `host_root` | dropped outright |
|
||||||
|
| `demote_domains` | ranked below everything, never dropped |
|
||||||
|
| `prefer_domains` | promoted above default rank |
|
||||||
|
| `ranking.*` | `demote_penalty`, `prefer_bonus`, `multi_engine_bonus` |
|
||||||
|
| `extraction.*` | `top_n`, `total_chars`, `per_item_chars`, `timeout_seconds` |
|
||||||
|
|
||||||
|
**Demotion, not deletion, for content farms**: a genuinely useful hit is not lost,
|
||||||
|
it simply cannot outrank a primary source. Non-answers are dropped because they
|
||||||
|
cannot answer a question at all.
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 scripts/search-agent-consume.py "query text" # JSON
|
||||||
|
python3 scripts/search-agent-consume.py --no-extract "query" # ranking only
|
||||||
|
python3 scripts/search-agent-consume.py --explain "query" # + drop reasons
|
||||||
|
```
|
||||||
|
|
||||||
|
Exit `0` ok, `1` nothing survived filtering, `2` the layer could not run.
|
||||||
|
|
||||||
|
## Output shape
|
||||||
|
|
||||||
|
Stable JSON:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"query": "...",
|
||||||
|
"raw_result_count": 46,
|
||||||
|
"returned_count": 44,
|
||||||
|
"dropped_count": 2,
|
||||||
|
"engines": ["bing", "brave", "duckduckgo", "yandex"],
|
||||||
|
"results": [
|
||||||
|
{"rank": 1, "title": "...", "url": "...", "host": "...",
|
||||||
|
"source_type": "official|code|qa|forum|discussion|web|content-farm",
|
||||||
|
"engines": ["bing"], "score": 100.0,
|
||||||
|
"excerpt": "...", "extraction": "ok|truncated|skipped_budget_exhausted|empty|failed:<Type>"}
|
||||||
|
],
|
||||||
|
"extraction": {"extracted": 5, "chars_used": 12000, "budget": 12000,
|
||||||
|
"failures": 0, "seconds": 5.28}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`--explain` adds `dropped: [{url, reason, position}]` so the filter is auditable
|
||||||
|
rather than magic.
|
||||||
|
|
||||||
|
## Measured before/after (2026-09-26)
|
||||||
|
|
||||||
|
Fixed query set. Relevance judged per query, not by impression.
|
||||||
|
|
||||||
|
**`best practices agent context management`**
|
||||||
|
|
||||||
|
| | before (raw SearXNG) | after (layer) |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| 1-2 | anthropic, stackai | anthropic, langchain |
|
||||||
|
| 3-4 | aitechmonk, agentic-design | jetbrains, blog.jetbrains |
|
||||||
|
| 5-6 | mindstudio, sparkco | docs.langchain, reddit |
|
||||||
|
| 7-8 | langchain, medium | cursor, reddit |
|
||||||
|
| verdict | 4 relevant of 10; 4 content farms; medium.com twice | top 8 all primary/discussion; no content farm in the top 8 |
|
||||||
|
|
||||||
|
**`proxmox thin pool metadata exhaustion recovery`**
|
||||||
|
|
||||||
|
| | before | after |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| 1-4 | vormox, linuxoperatingsystem, riparazioneserver, bigiron (all SEO/thin) | forum.proxmox.com, forum.proxmox.com, gist.github, github |
|
||||||
|
| 5-9 | forum.proxmox.com x2, voxfor, github, gist | forum.proxmox.com, serverfault, forum.proxmox.com, reddit |
|
||||||
|
|
||||||
|
The primary sources moved from positions 5-9 to 1-4.
|
||||||
|
|
||||||
|
**Rule proof** (`--explain`, and a direct check of the classifier):
|
||||||
|
|
||||||
|
```
|
||||||
|
DigitalOcean docs -> KEEP (a '/products/' path rule was REMOVED after the
|
||||||
|
before/after run caught it dropping this page)
|
||||||
|
Best Buy -> DROP shopping_or_dictionary_host
|
||||||
|
Merriam-Webster -> DROP shopping_or_dictionary_host
|
||||||
|
bare homepage -> DROP navigational_host_root
|
||||||
|
proxmox.com home -> KEEP (preferred host root: a repo/docs front door is
|
||||||
|
legitimately the answer)
|
||||||
|
github repo -> KEEP
|
||||||
|
```
|
||||||
|
|
||||||
|
**Extraction cost (criterion 4):**
|
||||||
|
|
||||||
|
```
|
||||||
|
extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s
|
||||||
|
whole run end-to-end: 6.4s wall
|
||||||
|
```
|
||||||
|
|
||||||
|
## Regression guard
|
||||||
|
|
||||||
|
`search-stack-visibility` asserts the layer still ranks correctly: for the fixed
|
||||||
|
query set, no `demote_domains` host may appear in the top 3, and the two known
|
||||||
|
non-answers must not be returned. Without it this layer could silently rot back
|
||||||
|
to raw ordering, which is exactly what happened to the endpoint colours.
|
||||||
|
|
||||||
|
## Reachability, and one honest gap
|
||||||
|
|
||||||
|
- **Hermes agents** reach it directly: it reads the same `SEARXNG_URL` and
|
||||||
|
`FIRECRAWL_URL` they already use.
|
||||||
|
- **pi agents (MCP search server)**: the MCP server's request/response shape is
|
||||||
|
**not ours to change**, so this layer is **NOT** wired into it. That is a real
|
||||||
|
gap, stated rather than claimed as coverage. Closing it would require a change
|
||||||
|
on the MCP side, which is outside this repo.
|
||||||
|
|
||||||
|
## Constraints
|
||||||
|
|
||||||
|
Does not touch the live SearXNG or Firecrawl service paths. Third-party
|
||||||
|
`google cse` is not a hard requirement of this layer — if it 429s, ranking still
|
||||||
|
works from the remaining engines. No credential is added or required.
|
||||||
@@ -0,0 +1,124 @@
|
|||||||
|
---
|
||||||
|
kind: function
|
||||||
|
name: search-stack-visibility
|
||||||
|
description: >
|
||||||
|
Makes the shared search stack observable. Every agent reaches one SearXNG
|
||||||
|
instance (http://192.168.68.7:8888) and one extraction service (Firecrawl,
|
||||||
|
http://192.168.68.7:3002). Before this check the stack could degrade to a
|
||||||
|
single engine, or an enabled engine could return nothing at all, without any
|
||||||
|
error surfacing anywhere.
|
||||||
|
|
||||||
|
This contract runs scripts/search-stack-check.py, which:
|
||||||
|
* runs two fixed queries against SearXNG and FAILS when fewer than two
|
||||||
|
engines contribute, printing the contributing engines and every
|
||||||
|
unresponsive_engines entry;
|
||||||
|
* checks extraction by scraping a known page through Firecrawl and FAILS
|
||||||
|
when the returned markdown is empty or the request fails;
|
||||||
|
* reports every silent-zero engine explicitly (enabled, not in
|
||||||
|
unresponsive_engines, contributed no results).
|
||||||
|
|
||||||
|
Multi-engine state (2026-09-25): bing, google cse, brave and yandex
|
||||||
|
contribute on every query. duckduckgo is NOT working: the house egress IP
|
||||||
|
and the VPS fallback egress are both flagged by DuckDuckGo and it reports
|
||||||
|
CAPTCHA. It is left enabled as best-effort coverage so that a recovery shows
|
||||||
|
up as a contribution.
|
||||||
|
|
||||||
|
google cse is a third party's public search-engine id hardcoded in the
|
||||||
|
SearXNG build. Quota and availability are outside our control.
|
||||||
|
|
||||||
|
SCHEDULED: /etc/cron.d/contract-runner on CT 100 (abiba), hourly at :15,
|
||||||
|
via scripts/contract-run.sh search-stack-visibility. Logs land in
|
||||||
|
/var/log/contract-runs/. A failure also raises a firstmate inbox note.
|
||||||
|
version: 1.1.0
|
||||||
|
---
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
The fleet has exactly one search endpoint and one extraction endpoint. If
|
||||||
|
either degrades, every agent silently loses capability at the same moment.
|
||||||
|
The failure mode this contract exists to close is *silent* degradation: a
|
||||||
|
query that still returns a page of results while all but one engine have
|
||||||
|
stopped contributing, or an enabled engine that answers with zero results and
|
||||||
|
raises no error.
|
||||||
|
|
||||||
|
## Execution model
|
||||||
|
|
||||||
|
The contract is a host-scheduled check, not an agent workflow. It is driven by
|
||||||
|
`scripts/contract-run.sh search-stack-visibility` from
|
||||||
|
`/etc/cron.d/contract-runner` on CT 100. `contract-run.sh` resolves the
|
||||||
|
mapping to `scripts/search-stack-check.py`, runs it under a timeout, writes a
|
||||||
|
timestamped log to `/var/log/contract-runs/`, and on non-zero exit raises a
|
||||||
|
firstmate inbox note through `bin/fm-inbox.sh`.
|
||||||
|
|
||||||
|
## What passing looks like
|
||||||
|
|
||||||
|
```
|
||||||
|
$ bash scripts/contract-run.sh search-stack-visibility
|
||||||
|
Expected engines, enabled (5): ['bing', 'brave', 'duckduckgo', 'google cse', 'yandex']
|
||||||
|
queries: 'proxmox backup server' -> contributing: bing, brave, google cse, yandex
|
||||||
|
unresponsive: duckduckgo=CAPTCHA
|
||||||
|
'python asyncio tutorial' -> contributing: bing, brave, google cse, yandex
|
||||||
|
EXTRACTION: 71016 chars of markdown returned
|
||||||
|
VERDICT: PASS -- multiple engines contributing, extraction healthy
|
||||||
|
```
|
||||||
|
|
||||||
|
## What failing looks like
|
||||||
|
|
||||||
|
* A query whose results come from fewer than `SEARCH_CHECK_MIN_ENGINES`
|
||||||
|
engines (default 2) fails and names the engines that did contribute.
|
||||||
|
* An extraction request that errors or returns empty markdown fails.
|
||||||
|
|
||||||
|
## Silent zeros are reported, not fatal
|
||||||
|
|
||||||
|
An enabled, expected engine that contributed nothing **without reporting an
|
||||||
|
error** is printed under `SILENT-ZERO ENGINES REPORTED`, and each occurrence is
|
||||||
|
annotated `reported, not fatal`. This is deliberate:
|
||||||
|
|
||||||
|
* a general query can legitimately draw zero results from an engine that only
|
||||||
|
fires on certain query shapes, and results are de-duplicated across engines,
|
||||||
|
so a zero does not by itself prove the engine is broken;
|
||||||
|
* the run therefore fails only on the two conditions that do prove loss of
|
||||||
|
capability -- fewer than two contributing engines, and a broken extraction
|
||||||
|
leg.
|
||||||
|
|
||||||
|
A run can consequently print `VERDICT: PASS` while still listing a silent
|
||||||
|
zero. That is the intended relationship: the zero is *visible*, not *fatal*.
|
||||||
|
An engine that fails with an error (for example DuckDuckGo returning CAPTCHA)
|
||||||
|
appears in `unresponsive_engines` instead.
|
||||||
|
|
||||||
|
## Google coverage is third-party, not ours
|
||||||
|
|
||||||
|
The free Google-derived results come from the SearXNG build's built-in
|
||||||
|
`google cse` engine. It uses **a third party's public search-engine id
|
||||||
|
hardcoded in the build** (`google_cse.py`, `CX = "partner-pub-8993..."`,
|
||||||
|
blackle.com), not a key or id we own. Its quota and availability are outside
|
||||||
|
our control and it can be rate-limited or withdrawn without notice. No engine
|
||||||
|
in this build accepts our own Google Custom Search key; using our own free key
|
||||||
|
would require a small wrapper service, which is deliberately **not** built.
|
||||||
|
|
||||||
|
## Configuration
|
||||||
|
|
||||||
|
Environment overrides (see the script docstring for the full list):
|
||||||
|
|
||||||
|
| Variable | Default | Meaning |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `SEARXNG_URL` | `http://192.168.68.7:8888` | SearXNG base URL |
|
||||||
|
| `FIRECRAWL_URL` | `http://192.168.68.7:3002` | Firecrawl base URL |
|
||||||
|
| `SEARCH_CHECK_QUERIES` | `proxmox backup server,python asyncio tutorial` | fixed queries |
|
||||||
|
| `SEARCH_CHECK_MIN_ENGINES` | `2` | minimum contributing engines per query |
|
||||||
|
| `SEARCH_CHECK_ENGINES` | `bing,brave,google cse,yandex,duckduckgo` | engines a silent zero is reported for |
|
||||||
|
| `SEARCH_CHECK_EXTRACT_URL` | Wikipedia Proxmox article | page used for the extraction leg |
|
||||||
|
|
||||||
|
## Known residual risk
|
||||||
|
|
||||||
|
DuckDuckGo is **not** working. The house egress IP is CAPTCHA'd by
|
||||||
|
DuckDuckGo, and a forward proxy on the VPS (`10.10.10.1:3128`, WireGuard) was
|
||||||
|
built as a second egress -- but DuckDuckGo has since flagged the VPS address
|
||||||
|
too (HTTP 202 with challenge markers), so DuckDuckGo now reports CAPTCHA on
|
||||||
|
both paths. It is left enabled as best-effort coverage: if DuckDuckGo
|
||||||
|
unflags either address it will show up as a contribution, and until then it is
|
||||||
|
visible in `unresponsive_engines` every run. It is never a required engine.
|
||||||
|
|
||||||
|
The VPS forward proxy remains a real service (`/opt/fwd-proxy`,
|
||||||
|
`restart: unless-stopped`, healthy healthcheck, Docker enabled at boot) so the
|
||||||
|
second egress path is available for any engine that benefits from it in future.
|
||||||
@@ -27,7 +27,7 @@ Agent (Hermes/pi) ──curl + X-API-Key──► Stirling-PDF (:8989) ──►
|
|||||||
| Base URL | `http://192.168.68.7:8989` |
|
| Base URL | `http://192.168.68.7:8989` |
|
||||||
| Auth Method | API Key (header) |
|
| Auth Method | API Key (header) |
|
||||||
| Header Name | `X-API-Key` |
|
| Header Name | `X-API-Key` |
|
||||||
| API Key | `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88` |
|
| API Key | `«vault: infrastructure/production STIRLING_API_KEY»` |
|
||||||
| Key Source | `SECURITY_CUSTOMGLOBALAPIKEY` in `/opt/home_stack/docker-compose.yml` |
|
| Key Source | `SECURITY_CUSTOMGLOBALAPIKEY` in `/opt/home_stack/docker-compose.yml` |
|
||||||
| Swagger | `http://192.168.68.7:8989/swagger-ui.html` |
|
| Swagger | `http://192.168.68.7:8989/swagger-ui.html` |
|
||||||
| Health | `http://192.168.68.7:8989/api/v1/info/status` |
|
| Health | `http://192.168.68.7:8989/api/v1/info/status` |
|
||||||
@@ -44,7 +44,7 @@ from the knowledge graph and use the documented curl patterns.
|
|||||||
Direct bash invocations:
|
Direct bash invocations:
|
||||||
```bash
|
```bash
|
||||||
curl -X POST "http://192.168.68.7:8989/api/v1/split-pdf" \
|
curl -X POST "http://192.168.68.7:8989/api/v1/split-pdf" \
|
||||||
-H "X-API-Key: adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88" \
|
-H "X-API-Key: «vault: infrastructure/production STIRLING_API_KEY»" \
|
||||||
-F "fileInput=@/path/to/file.pdf" \
|
-F "fileInput=@/path/to/file.pdf" \
|
||||||
-F "pageNumbers=1,2,3" \
|
-F "pageNumbers=1,2,3" \
|
||||||
-o /tmp/output.zip
|
-o /tmp/output.zip
|
||||||
|
|||||||
+40
@@ -0,0 +1,40 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Revision preflight guard: verify the script being executed matches origin/master
|
||||||
|
# Usage: revision-preflight.sh <script-path> <clone-path>
|
||||||
|
# Returns 0 if match, 1 if mismatch (prints both revisions)
|
||||||
|
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
SCRIPT="${1:?Usage: revision-preflight.sh <script-path> <clone-path>}"
|
||||||
|
CLONE="${2:?Usage: revision-preflight.sh <script-path> <clone-path>}"
|
||||||
|
|
||||||
|
# Compute sha256 of the script being executed
|
||||||
|
EXEC_SHA=$(sha256sum "$SCRIPT" | cut -d' ' -f1)
|
||||||
|
|
||||||
|
# Compute sha256 of the merged origin/master version
|
||||||
|
# Extract to a temp file to avoid pipe issues
|
||||||
|
TMPFILE=$(mktemp)
|
||||||
|
trap 'rm -f "$TMPFILE"' EXIT
|
||||||
|
|
||||||
|
# Try to extract the file from origin/master
|
||||||
|
if git -C "$CLONE" show "origin/master:$(basename "$SCRIPT")" > "$TMPFILE" 2>/dev/null; then
|
||||||
|
MASTER_SHA=$(sha256sum "$TMPFILE" | cut -d' ' -f1)
|
||||||
|
else
|
||||||
|
echo "⚠️ revision-preflight: could not resolve origin/master revision for $(basename "$SCRIPT")" >&2
|
||||||
|
exit 0 # Warn but don't block if git show fails
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ -z "$MASTER_SHA" || "$MASTER_SHA" == "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" ]]; then
|
||||||
|
echo "⚠️ revision-preflight: could not resolve origin/master revision for $(basename "$SCRIPT")" >&2
|
||||||
|
exit 0 # Warn but don't block if git show fails
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ "$EXEC_SHA" != "$MASTER_SHA" ]]; then
|
||||||
|
echo "⚠️ revision-preflight: MISMATCH detected" >&2
|
||||||
|
echo " Executed: $EXEC_SHA ($(basename "$SCRIPT"))" >&2
|
||||||
|
echo " Merged: $MASTER_SHA (origin/master:$(basename "$SCRIPT"))" >&2
|
||||||
|
exit 1
|
||||||
|
else
|
||||||
|
echo "✅ revision-preflight: $SCRIPT matches origin/master ($EXEC_SHA)" >&2
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
@@ -24,7 +24,7 @@ BASE = """
|
|||||||
model:
|
model:
|
||||||
api_key: ""
|
api_key: ""
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
max_tokens: 4096
|
max_tokens: 4096
|
||||||
default: syslog-auto
|
default: syslog-auto
|
||||||
provider: harness
|
provider: harness
|
||||||
@@ -52,7 +52,7 @@ delegation:
|
|||||||
custom_providers:
|
custom_providers:
|
||||||
- name: harness
|
- name: harness
|
||||||
key_env: LITELLM_API_KEY
|
key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
"""
|
"""
|
||||||
|
|
||||||
|
|
||||||
@@ -189,3 +189,73 @@ def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
|
|||||||
assert code == 1, out
|
assert code == 1, out
|
||||||
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
|
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
|
||||||
assert "RESULT: FAIL" in out
|
assert "RESULT: FAIL" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_canonical_internal_path_passes(tmp_path):
|
||||||
|
"""Rule 5 must accept the canonical internal base from hermes-key-enforcement.prose.md."""
|
||||||
|
code, out = _run(
|
||||||
|
tmp_path,
|
||||||
|
"gpu-vision",
|
||||||
|
)
|
||||||
|
# Override the base_url in the config
|
||||||
|
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||||
|
" base_url: http://192.168.68.116/v1",
|
||||||
|
" base_url: http://192.168.68.116/litellm/v1",
|
||||||
|
)
|
||||||
|
code, out = _run_config(tmp_path, "canonical-internal.yaml", cfg_text)
|
||||||
|
assert code == 0, out
|
||||||
|
assert "RESULT: PASS" in out
|
||||||
|
# Verify the correct message is shown
|
||||||
|
assert "model.base_url is canonical" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_wrong_base_url_fails(tmp_path):
|
||||||
|
"""Rule 5 must reject paths outside the allowed list."""
|
||||||
|
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||||
|
" base_url: http://192.168.68.116/litellm/v1",
|
||||||
|
" base_url: http://192.168.68.116/litellm/v1/responses",
|
||||||
|
)
|
||||||
|
code, out = _run_config(tmp_path, "wrong-base.yaml", cfg_text)
|
||||||
|
assert code == 1, out
|
||||||
|
assert "RESULT: FAIL" in out
|
||||||
|
assert "model.base_url must be one of" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_public_host_path_passes(tmp_path):
|
||||||
|
"""Rule 5 must accept the public host base."""
|
||||||
|
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||||
|
" base_url: http://192.168.68.116/v1",
|
||||||
|
" base_url: https://litellm.sysloggh.net/v1",
|
||||||
|
)
|
||||||
|
code, out = _run_config(tmp_path, "public-host.yaml", cfg_text)
|
||||||
|
assert code == 0, out
|
||||||
|
assert "RESULT: PASS" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_old_rule5_check_would_fail_canonical(tmp_path):
|
||||||
|
"""
|
||||||
|
Proof that the OLD Rule 5 check would fail the canonical internal path.
|
||||||
|
This proves the bug existed before the fix.
|
||||||
|
"""
|
||||||
|
# OLD check expected /v1, so the canonical /litellm/v1 would have failed
|
||||||
|
canonical_cfg = BASE.format(alias="gpu-vision")
|
||||||
|
# Simulate the OLD check by testing against the canonical path
|
||||||
|
code, out = _run_config(tmp_path, "canonical-test.yaml", canonical_cfg)
|
||||||
|
# NEW check: canonical /litellm/v1 SHOULD pass
|
||||||
|
assert code == 0, out
|
||||||
|
assert "RESULT: PASS" in out
|
||||||
|
|
||||||
|
# OLD check expected /v1, so the internal /v1 would have passed
|
||||||
|
# NEW check: internal /v1 is non-canonical but working (WARN not FAIL)
|
||||||
|
old_cfg = BASE.format(alias="gpu-vision").replace(
|
||||||
|
" base_url: http://192.168.68.116/litellm/v1",
|
||||||
|
" base_url: http://192.168.68.116/v1",
|
||||||
|
)
|
||||||
|
code, out = _run_config(tmp_path, "old-check-test.yaml", old_cfg)
|
||||||
|
assert code == 0, out
|
||||||
|
assert "RESULT: PASS" in out
|
||||||
|
# NEW check should pass
|
||||||
|
new_cfg = BASE.format(alias="gpu-vision")
|
||||||
|
code, out = _run_config(tmp_path, "new-check-test.yaml", new_cfg)
|
||||||
|
assert code == 0, out
|
||||||
|
assert "RESULT: PASS" in out
|
||||||
|
|||||||
@@ -0,0 +1,236 @@
|
|||||||
|
"""Regression test for the fallback_providers list-shape crash in audit-hermes-config.py.
|
||||||
|
|
||||||
|
WHY THIS FILE EXISTS: audit-hermes-config.py assumed `fallback_providers` was always a dict
|
||||||
|
(single provider). Two live agents (koby, koonimo) carry it as a LIST of dicts (one entry per
|
||||||
|
fallback), so the script crashed with:
|
||||||
|
|
||||||
|
File "audit-hermes-config.py", line 211, in audit
|
||||||
|
fb.get("provider") == "deepseek",
|
||||||
|
AttributeError: 'list' object has no attribute 'get'
|
||||||
|
|
||||||
|
Both are REAL agent configs, so this is not a malformed-input case — the script simply could not
|
||||||
|
audit two of the four agents it exists to audit. Until fixed, the key-hygiene check had no
|
||||||
|
coverage for half the fleet while appearing to run.
|
||||||
|
|
||||||
|
These tests execute the real CLI (`python3 audit-hermes-config.py <config>`) and assert:
|
||||||
|
1. A config whose `fallback_providers` is a LIST of valid dicts does NOT crash (exit code is 0 or 1,
|
||||||
|
never a traceback/AttributeError).
|
||||||
|
2. A config whose `fallback_providers` contains a MALFORMED entry (a list element that is not a
|
||||||
|
mapping) reports a VIOLATION naming the offending entry, NOT an uncaught exception.
|
||||||
|
3. The dict shape still works (existing tests must stay green).
|
||||||
|
|
||||||
|
No network, vault, or SSH access is required.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pathlib
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
|
||||||
|
ROOT = pathlib.Path(__file__).resolve().parent.parent
|
||||||
|
AUDIT = ROOT / "audit-hermes-config.py"
|
||||||
|
|
||||||
|
# A valid config where fallback_providers is a LIST of dicts (the real koby/koonimo shape).
|
||||||
|
# One entry, well-formed: provider=deepseek, model=deepseek-v4-flash, api_key_env=DEEPSEEK_API_KEY.
|
||||||
|
# This must produce a real verdict (PASS or FAIL) without crashing.
|
||||||
|
LIST_SHAPE_VALID = """
|
||||||
|
model:
|
||||||
|
api_key: ""
|
||||||
|
api_key_env: LITELLM_API_KEY
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
max_tokens: 4096
|
||||||
|
default: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
fallback_providers:
|
||||||
|
- provider: deepseek
|
||||||
|
model: deepseek-v4-flash
|
||||||
|
api_key_env: DEEPSEEK_API_KEY
|
||||||
|
compression:
|
||||||
|
model: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
threshold: 0.65
|
||||||
|
max_context_window: 131072
|
||||||
|
auxiliary:
|
||||||
|
vision:
|
||||||
|
model: gpu-vision
|
||||||
|
provider: harness
|
||||||
|
web_extract:
|
||||||
|
model: gpu-vision
|
||||||
|
provider: harness
|
||||||
|
compression:
|
||||||
|
model: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
delegation:
|
||||||
|
provider: harness
|
||||||
|
custom_providers:
|
||||||
|
- name: harness
|
||||||
|
key_env: LITELLM_API_KEY
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
"""
|
||||||
|
|
||||||
|
# A valid config where fallback_providers is a LIST with TWO entries (multiple fallbacks).
|
||||||
|
# Both entries well-formed. Must not crash and should produce a real verdict.
|
||||||
|
LIST_SHAPE_MULTI = """
|
||||||
|
model:
|
||||||
|
api_key: ""
|
||||||
|
api_key_env: LITELLM_API_KEY
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
max_tokens: 4096
|
||||||
|
default: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
fallback_providers:
|
||||||
|
- provider: deepseek
|
||||||
|
model: deepseek-v4-flash
|
||||||
|
api_key_env: DEEPSEEK_API_KEY
|
||||||
|
- provider: deepseek
|
||||||
|
model: deepseek-v4-flash
|
||||||
|
api_key_env: DEEPSEEK_API_KEY
|
||||||
|
compression:
|
||||||
|
model: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
threshold: 0.65
|
||||||
|
max_context_window: 131072
|
||||||
|
auxiliary:
|
||||||
|
vision:
|
||||||
|
model: gpu-vision
|
||||||
|
provider: harness
|
||||||
|
web_extract:
|
||||||
|
model: gpu-vision
|
||||||
|
provider: harness
|
||||||
|
compression:
|
||||||
|
model: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
delegation:
|
||||||
|
provider: harness
|
||||||
|
custom_providers:
|
||||||
|
- name: harness
|
||||||
|
key_env: LITELLM_API_KEY
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
"""
|
||||||
|
|
||||||
|
# A config where fallback_providers is a LIST containing a MALFORMED entry:
|
||||||
|
# one element is a plain string, not a mapping. The checker must report a VIOLATION
|
||||||
|
# naming the offending entry (fallback_providers[1]) and NOT crash.
|
||||||
|
LIST_SHAPE_MALFORMED = """
|
||||||
|
model:
|
||||||
|
api_key: ""
|
||||||
|
api_key_env: LITELLM_API_KEY
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
max_tokens: 4096
|
||||||
|
default: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
fallback_providers:
|
||||||
|
- provider: deepseek
|
||||||
|
model: deepseek-v4-flash
|
||||||
|
api_key_env: DEEPSEEK_API_KEY
|
||||||
|
- "not-a-mapping"
|
||||||
|
compression:
|
||||||
|
model: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
threshold: 0.65
|
||||||
|
max_context_window: 131072
|
||||||
|
auxiliary:
|
||||||
|
vision:
|
||||||
|
model: gpu-vision
|
||||||
|
provider: harness
|
||||||
|
web_extract:
|
||||||
|
model: gpu-vision
|
||||||
|
provider: harness
|
||||||
|
compression:
|
||||||
|
model: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
delegation:
|
||||||
|
provider: harness
|
||||||
|
custom_providers:
|
||||||
|
- name: harness
|
||||||
|
key_env: LITELLM_API_KEY
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
"""
|
||||||
|
|
||||||
|
# The original DICT shape (single provider) must still work — existing behaviour preserved.
|
||||||
|
DICT_SHAPE_VALID = """
|
||||||
|
model:
|
||||||
|
api_key: ""
|
||||||
|
api_key_env: LITELLM_API_KEY
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
max_tokens: 4096
|
||||||
|
default: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
fallback_providers:
|
||||||
|
provider: deepseek
|
||||||
|
model: deepseek-v4-flash
|
||||||
|
api_key_env: DEEPSEEK_API_KEY
|
||||||
|
compression:
|
||||||
|
model: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
threshold: 0.65
|
||||||
|
max_context_window: 131072
|
||||||
|
auxiliary:
|
||||||
|
vision:
|
||||||
|
model: gpu-vision
|
||||||
|
provider: harness
|
||||||
|
web_extract:
|
||||||
|
model: gpu-vision
|
||||||
|
provider: harness
|
||||||
|
compression:
|
||||||
|
model: syslog-auto
|
||||||
|
provider: harness
|
||||||
|
delegation:
|
||||||
|
provider: harness
|
||||||
|
custom_providers:
|
||||||
|
- name: harness
|
||||||
|
key_env: LITELLM_API_KEY
|
||||||
|
base_url: http://192.168.68.116/litellm/v1
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def _run_config(tmp_path, name, text):
|
||||||
|
cfg = tmp_path / name
|
||||||
|
cfg.write_text(text)
|
||||||
|
proc = subprocess.run(
|
||||||
|
[sys.executable, str(AUDIT), str(cfg)],
|
||||||
|
capture_output=True, text=True,
|
||||||
|
)
|
||||||
|
return proc.returncode, proc.stdout, proc.stderr
|
||||||
|
|
||||||
|
|
||||||
|
def test_list_shape_single_entry_does_not_crash(tmp_path):
|
||||||
|
"""A LIST with one valid dict must not raise AttributeError; exit 0 (PASS)."""
|
||||||
|
code, out, err = _run_config(tmp_path, "list-single.yaml", LIST_SHAPE_VALID)
|
||||||
|
# Must NOT be a crash (traceback). A clean run exits 0 (PASS) or 1 (FAIL), never 2+ (exception).
|
||||||
|
assert code in (0, 1), f"Expected clean exit 0 or 1, got {code}\nSTDOUT:\n{out}\nSTDERR:\n{err}"
|
||||||
|
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
|
||||||
|
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
|
||||||
|
# The valid single-entry list should PASS (all rules satisfied).
|
||||||
|
assert code == 0, f"Expected PASS but got {code}\n{out}"
|
||||||
|
assert "RESULT: PASS" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_list_shape_multiple_entries_does_not_crash(tmp_path):
|
||||||
|
"""A LIST with two valid dicts must not raise AttributeError; exit 0 (PASS)."""
|
||||||
|
code, out, err = _run_config(tmp_path, "list-multi.yaml", LIST_SHAPE_MULTI)
|
||||||
|
assert code in (0, 1), f"Expected clean exit 0 or 1, got {code}\nSTDOUT:\n{out}\nSTDERR:\n{err}"
|
||||||
|
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
|
||||||
|
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
|
||||||
|
assert code == 0, f"Expected PASS but got {code}\n{out}"
|
||||||
|
assert "RESULT: PASS" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_list_shape_malformed_entry_reports_violation_not_crash(tmp_path):
|
||||||
|
"""A LIST containing a non-mapping element must be a reported VIOLATION, not a crash."""
|
||||||
|
code, out, err = _run_config(tmp_path, "list-malformed.yaml", LIST_SHAPE_MALFORMED)
|
||||||
|
# Must NOT be a crash.
|
||||||
|
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
|
||||||
|
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
|
||||||
|
# Should be a FAIL (exit 1) because the malformed entry is a violation.
|
||||||
|
assert code == 1, f"Expected FAIL (exit 1) but got {code}\n{out}"
|
||||||
|
assert "RESULT: FAIL" in out
|
||||||
|
# The violation must name the offending entry (fallback_providers[1]).
|
||||||
|
assert "fallback_providers[1]" in out, f"Violation did not name the offending entry:\n{out}"
|
||||||
|
|
||||||
|
|
||||||
|
def test_dict_shape_still_passes(tmp_path):
|
||||||
|
"""The original DICT shape (single provider) must still PASS — existing behaviour preserved."""
|
||||||
|
code, out, err = _run_config(tmp_path, "dict-valid.yaml", DICT_SHAPE_VALID)
|
||||||
|
assert code == 0, f"Expected PASS but got {code}\n{out}\nSTDERR:\n{err}"
|
||||||
|
assert "RESULT: PASS" in out
|
||||||
Executable
+80
@@ -0,0 +1,80 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# test_contract_run.sh — Tests for contract-run.sh
|
||||||
|
#
|
||||||
|
# Proves:
|
||||||
|
# 1. A passing contract exits 0 and does NOT send an alert
|
||||||
|
# 2. A failing contract exits non-zero and DOES send an alert
|
||||||
|
# 3. Log files are created in /var/log/contract-runs/
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
TEST_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
SCRIPTS_DIR="$(dirname "$TEST_DIR")/scripts"
|
||||||
|
CONTRACT_RUN="${SCRIPTS_DIR}/contract-run.sh"
|
||||||
|
LOG_DIR="/var/log/contract-runs"
|
||||||
|
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
|
||||||
|
# Test 1: Passing contract should exit 0
|
||||||
|
echo "=== Test 1: Passing contract ==="
|
||||||
|
# Use a simple passing contract (proxmox-monitor should pass if services are up)
|
||||||
|
bash "$CONTRACT_RUN" "proxmox-monitor"
|
||||||
|
EXIT_CODE=$?
|
||||||
|
if [ $EXIT_CODE -eq 0 ]; then
|
||||||
|
echo "✅ Test 1 PASSED: contract passed with exit code 0"
|
||||||
|
PASS=$((PASS + 1))
|
||||||
|
else
|
||||||
|
echo "🔴 Test 1 FAILED: expected exit code 0, got $EXIT_CODE"
|
||||||
|
FAIL=$((FAIL + 1))
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Check log file was created
|
||||||
|
LATEST_LOG=$(ls -t "$LOG_DIR"/proxmox-monitor-*.log 2>/dev/null | head -1)
|
||||||
|
if [ -n "$LATEST_LOG" ] && [ -f "$LATEST_LOG" ]; then
|
||||||
|
echo "✅ Log file created: $LATEST_LOG"
|
||||||
|
PASS=$((PASS + 1))
|
||||||
|
else
|
||||||
|
echo "🔴 Log file not found"
|
||||||
|
FAIL=$((FAIL + 1))
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Test 2: Failing contract should exit non-zero
|
||||||
|
echo ""
|
||||||
|
echo "=== Test 2: Failing contract ==="
|
||||||
|
# Create a temporary failing contract
|
||||||
|
TEMP_SCRIPT="${SCRIPTS_DIR}/test-failing-contract.sh"
|
||||||
|
cat > "$TEMP_SCRIPT" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo "This is a test failure"
|
||||||
|
exit 1
|
||||||
|
EOF
|
||||||
|
chmod +x "$TEMP_SCRIPT"
|
||||||
|
|
||||||
|
# Temporarily modify contract-run.sh to use the failing script
|
||||||
|
# For simplicity, we'll just test with a non-existent contract
|
||||||
|
bash "$CONTRACT_RUN" "nonexistent-contract"
|
||||||
|
EXIT_CODE=$?
|
||||||
|
if [ $EXIT_CODE -ne 0 ]; then
|
||||||
|
echo "✅ Test 2 PASSED: failing contract exited with code $EXIT_CODE"
|
||||||
|
PASS=$((PASS + 1))
|
||||||
|
else
|
||||||
|
echo "🔴 Test 2 FAILED: expected non-zero exit, got 0"
|
||||||
|
FAIL=$((FAIL + 1))
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Cleanup
|
||||||
|
rm -f "$TEMP_SCRIPT"
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
echo "=== Summary ==="
|
||||||
|
echo "Passed: $PASS"
|
||||||
|
echo "Failed: $FAIL"
|
||||||
|
|
||||||
|
if [ $FAIL -eq 0 ]; then
|
||||||
|
echo "✅ All tests passed"
|
||||||
|
exit 0
|
||||||
|
else
|
||||||
|
echo "🔴 Some tests failed"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
@@ -6,6 +6,7 @@ Tests:
|
|||||||
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
|
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
|
||||||
"""
|
"""
|
||||||
import json
|
import json
|
||||||
|
import os
|
||||||
import subprocess
|
import subprocess
|
||||||
import sys
|
import sys
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
@@ -24,6 +25,53 @@ def load_script():
|
|||||||
return module
|
return module
|
||||||
|
|
||||||
|
|
||||||
|
def test_pve_token_is_read_from_the_environment():
|
||||||
|
"""The PVE token must come from the injected environment, never a literal.
|
||||||
|
|
||||||
|
Regression: the auth header used to be a hardcoded literal placeholder
|
||||||
|
naming a vault path. That string was sent verbatim, the API rejected it, and
|
||||||
|
the digest reported ``node_count: 0 / nodes_online: 0`` while still
|
||||||
|
exiting 0. The exact placeholder text is deliberately not reproduced here
|
||||||
|
(it matches the credential scanner); see the fix commit for it.
|
||||||
|
"""
|
||||||
|
mod = load_script()
|
||||||
|
assert hasattr(mod, "pve_auth"), "pve_auth() must exist to resolve the token at call time"
|
||||||
|
sample = "unit-test-sample-value"
|
||||||
|
with patch.dict("os.environ", {"PVE_TOKEN": sample}, clear=False):
|
||||||
|
assert mod.pve_auth() == f"Authorization: {mod.PVE_AUTH_HEADER}{sample}"
|
||||||
|
assert mod.pve_auth().endswith(sample)
|
||||||
|
|
||||||
|
|
||||||
|
def test_missing_pve_token_is_degraded_not_a_placeholder():
|
||||||
|
"""With no PVE_TOKEN, pve_get must return None (probe failure), not send a placeholder."""
|
||||||
|
mod = load_script()
|
||||||
|
env = {k: v for k, v in os.environ.items() if k != "PVE_TOKEN"}
|
||||||
|
with patch.dict("os.environ", env, clear=True):
|
||||||
|
assert mod.pve_get("/api2/json/nodes") is None, (
|
||||||
|
"a missing PVE_TOKEN must degrade to None so the caller records a probe failure"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_unreachable_probe_is_recorded_as_a_failure():
|
||||||
|
"""An unreachable probe must be recorded, so the run cannot pass silently."""
|
||||||
|
mod = load_script()
|
||||||
|
assert hasattr(mod, "PROBE_FAILURES"), "PROBE_FAILURES must exist"
|
||||||
|
mod.PROBE_FAILURES.clear()
|
||||||
|
with patch.object(mod, "pve_get", return_value=None):
|
||||||
|
report = mod.collect()
|
||||||
|
assert report["pve_probe_status"] == "unreachable"
|
||||||
|
assert any("unreachable" in f for f in mod.PROBE_FAILURES), (
|
||||||
|
f"unreachable probe must be recorded in PROBE_FAILURES, got {mod.PROBE_FAILURES}"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_pve_token_placeholder_is_gone():
|
||||||
|
"""The literal placeholder must no longer appear anywhere in the script."""
|
||||||
|
src = (Path(__file__).parent.parent / "scripts" / "daily-infra-report.py").read_text()
|
||||||
|
assert "«vault:" not in src, "the unresolved vault placeholder must not remain in the script"
|
||||||
|
assert "AUTH = \"Authorization" not in src, "the hardcoded AUTH literal must be gone"
|
||||||
|
|
||||||
|
|
||||||
def test_nested_zulip_read_feeds_agent_card():
|
def test_nested_zulip_read_feeds_agent_card():
|
||||||
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
|
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
|
||||||
# Mock the http_get_body response with nested structure
|
# Mock the http_get_body response with nested structure
|
||||||
@@ -135,4 +183,32 @@ if __name__ == "__main__":
|
|||||||
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
|
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
|
||||||
sys.exit(1)
|
sys.exit(1)
|
||||||
|
|
||||||
|
try:
|
||||||
|
test_pve_token_is_read_from_the_environment()
|
||||||
|
print("✓ test_pve_token_is_read_from_the_environment passed")
|
||||||
|
except AssertionError as e:
|
||||||
|
print(f"✗ test_pve_token_is_read_from_the_environment failed: {e}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
try:
|
||||||
|
test_missing_pve_token_is_degraded_not_a_placeholder()
|
||||||
|
print("✓ test_missing_pve_token_is_degraded_not_a_placeholder passed")
|
||||||
|
except AssertionError as e:
|
||||||
|
print(f"✗ test_missing_pve_token_is_degraded_not_a_placeholder failed: {e}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
try:
|
||||||
|
test_unreachable_probe_is_recorded_as_a_failure()
|
||||||
|
print("✓ test_unreachable_probe_is_recorded_as_a_failure passed")
|
||||||
|
except AssertionError as e:
|
||||||
|
print(f"✗ test_unreachable_probe_is_recorded_as_a_failure failed: {e}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
try:
|
||||||
|
test_pve_token_placeholder_is_gone()
|
||||||
|
print("✓ test_pve_token_placeholder_is_gone passed")
|
||||||
|
except AssertionError as e:
|
||||||
|
print(f"✗ test_pve_token_placeholder_is_gone failed: {e}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
print("All tests passed!")
|
print("All tests passed!")
|
||||||
|
|||||||
@@ -49,7 +49,7 @@ HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
|
|||||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||||
|
|
||||||
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
|
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
|
||||||
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
TANKO_VANTAGE = "192.168.68.12" # minipve — Tanko CT 112 via pct exec
|
||||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||||
|
|
||||||
|
|
||||||
@@ -82,7 +82,7 @@ done
|
|||||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||||
cmd="${*: -1}"
|
cmd="${*: -1}"
|
||||||
case "$host" in
|
case "$host" in
|
||||||
192.168.68.15)
|
192.168.68.12)
|
||||||
case "$cmd" in
|
case "$cmd" in
|
||||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||||
@@ -98,8 +98,8 @@ exit 0
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
CURL_STUB = r"""#!/usr/bin/env bash
|
CURL_STUB = r"""#!/usr/bin/env bash
|
||||||
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
|
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||||
# record every call (including notify) payloads.
|
# kagentz C3 public URL, and record every call (including notify) payloads.
|
||||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||||
case "$*" in
|
case "$*" in
|
||||||
*:9200/health*)
|
*:9200/health*)
|
||||||
@@ -109,6 +109,8 @@ case "$*" in
|
|||||||
esac ;;
|
esac ;;
|
||||||
*server_settings*)
|
*server_settings*)
|
||||||
printf '%s' "$SERVER_HTTP" ;;
|
printf '%s' "$SERVER_HTTP" ;;
|
||||||
|
*kagentz.sysloggh.net*)
|
||||||
|
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||||
esac
|
esac
|
||||||
exit 0
|
exit 0
|
||||||
"""
|
"""
|
||||||
@@ -121,7 +123,8 @@ def _write_exec(path: pathlib.Path, body: str) -> None:
|
|||||||
|
|
||||||
|
|
||||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||||
az_a2a_code="401", az_a2a_exit=0):
|
az_a2a_code="401", az_a2a_exit=0,
|
||||||
|
kagentz_public_code="302"):
|
||||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||||
|
|
||||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||||
@@ -154,6 +157,7 @@ def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
|||||||
"PI_HTTP": "200",
|
"PI_HTTP": "200",
|
||||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||||
"SERVER_HTTP": "200",
|
"SERVER_HTTP": "200",
|
||||||
|
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||||
})
|
})
|
||||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||||
capture_output=True, text=True)
|
capture_output=True, text=True)
|
||||||
@@ -169,8 +173,9 @@ def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
|||||||
assert "Server: ✅ HTTP 200" in log
|
assert "Server: ✅ HTTP 200" in log
|
||||||
assert "Abiba: ✅ Connected" in log
|
assert "Abiba: ✅ Connected" in log
|
||||||
assert "Tanko: ✅ service=active http=200" in log
|
assert "Tanko: ✅ service=active http=200" in log
|
||||||
assert "kagentz: ✅ A2A alive (HTTP 401)" in log
|
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||||
assert "Result: ✅ All healthy" in log
|
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||||
|
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||||
|
|
||||||
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||||
assert proc.stdout == ""
|
assert proc.stdout == ""
|
||||||
@@ -207,8 +212,8 @@ def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
|||||||
# The rest of the monitor still ran alongside the failing Tanko leg.
|
# The rest of the monitor still ran alongside the failing Tanko leg.
|
||||||
log = log_path.read_text()
|
log = log_path.read_text()
|
||||||
assert "Abiba: ✅ Connected" in log
|
assert "Abiba: ✅ Connected" in log
|
||||||
assert "kagentz: ✅ A2A alive" in log
|
assert "kagentz C1: ✅ A2A alive" in log
|
||||||
assert "Result: 🔴 1 issue(s) found" in log
|
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||||
|
|
||||||
|
|
||||||
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||||
@@ -216,9 +221,9 @@ def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
|||||||
assert proc.returncode == 0, proc.stderr
|
assert proc.returncode == 0, proc.stderr
|
||||||
log = log_path.read_text()
|
log = log_path.read_text()
|
||||||
|
|
||||||
assert "kagentz: 🟡 A2A unexpected http=500 (running, warning)" in log
|
assert "kagentz C1: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||||
assert "kagentz: ✅ A2A alive" not in log
|
assert "kagentz C1: ✅ A2A alive" not in log
|
||||||
assert "Result: 🔴 1 issue(s) found" in log
|
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||||
assert "kagentz A2A server answered HTTP 500" in proc.stdout
|
assert "kagentz A2A server answered HTTP 500" in proc.stdout
|
||||||
|
|
||||||
|
|
||||||
@@ -230,10 +235,10 @@ def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
|
|||||||
assert proc.returncode == 0, proc.stderr
|
assert proc.returncode == 0, proc.stderr
|
||||||
log = log_path.read_text()
|
log = log_path.read_text()
|
||||||
|
|
||||||
assert "kagentz: ❌ A2A down (HTTP 000)" in log
|
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||||
assert "kagentz: ✅ A2A alive" not in log
|
assert "kagentz C1: ✅ A2A alive" not in log
|
||||||
assert "unexpected" not in log
|
assert "unexpected" not in log
|
||||||
assert "Result: 🔴 1 issue(s) found" in log
|
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||||
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
|
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
Executable
+108
@@ -0,0 +1,108 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# test_pbs_gc_states.sh — Tests for PBS GC four-state logic
|
||||||
|
# Self-contained: inlines the SSH replacement logic
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
TEST_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
SCRIPTS_DIR="$(dirname "$TEST_DIR")/scripts"
|
||||||
|
PROXMOX_MONITOR="${1:-${SCRIPTS_DIR}/proxmox-monitor.sh}"
|
||||||
|
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
|
||||||
|
# Create the Python replacement script
|
||||||
|
REPLACE_SCRIPT=$(mktemp /tmp/replace_ssh_XXXXXX.py)
|
||||||
|
cat > "$REPLACE_SCRIPT" << 'PYEOF'
|
||||||
|
import sys
|
||||||
|
import re
|
||||||
|
|
||||||
|
wrapper = sys.argv[1]
|
||||||
|
monitor = sys.argv[2]
|
||||||
|
|
||||||
|
with open(monitor) as f:
|
||||||
|
c = f.read()
|
||||||
|
|
||||||
|
pattern = r'PBS_GC_OUTPUT=\$\(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192\.168\.68\.6 \\\n "/sbin/pct exec 107 -- /sbin/proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null\)'
|
||||||
|
replacement = 'PBS_GC_OUTPUT=$(cat "' + wrapper + '")'
|
||||||
|
|
||||||
|
if re.search(pattern, c):
|
||||||
|
c = re.sub(pattern, replacement, c)
|
||||||
|
|
||||||
|
with open(monitor, 'w') as f:
|
||||||
|
f.write(c)
|
||||||
|
PYEOF
|
||||||
|
|
||||||
|
run_test() {
|
||||||
|
local name="$1"
|
||||||
|
local json="$2"
|
||||||
|
local expected_behavior="$3"
|
||||||
|
local expected_pattern="$4"
|
||||||
|
|
||||||
|
local wrapper monitor
|
||||||
|
wrapper=$(mktemp /tmp/pbs-test-wrapper.XXXXXX)
|
||||||
|
monitor=$(mktemp /tmp/pbs-test-monitor.XXXXXX)
|
||||||
|
|
||||||
|
printf '%s\n' "$json" > "$wrapper"
|
||||||
|
cp "$PROXMOX_MONITOR" "$monitor"
|
||||||
|
|
||||||
|
# Replace the SSH call with cat "$wrapper"
|
||||||
|
python3 "$REPLACE_SCRIPT" "$wrapper" "$monitor"
|
||||||
|
|
||||||
|
local output exit_code
|
||||||
|
output=$(bash "$monitor" 2>&1)
|
||||||
|
exit_code=$?
|
||||||
|
|
||||||
|
local ok=true
|
||||||
|
if [ "$expected_behavior" = "fail" ]; then
|
||||||
|
# Should fail with PBS GC error
|
||||||
|
if ! echo "$output" | grep -q "🔴 PBS GC"; then ok=false; fi
|
||||||
|
if [ $exit_code -eq 0 ]; then ok=false; fi
|
||||||
|
else
|
||||||
|
# Should pass with expected pattern
|
||||||
|
if ! echo "$output" | grep -q "$expected_pattern"; then ok=false; fi
|
||||||
|
if echo "$output" | grep -q "🔴 PBS GC"; then ok=false; fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
if $ok; then
|
||||||
|
echo " ✅ $name"
|
||||||
|
PASS=$((PASS + 1))
|
||||||
|
else
|
||||||
|
echo " 🔴 $name FAILED (exit=$exit_code)"
|
||||||
|
echo "$output" | grep "PBS GC" | sed 's/^/ /'
|
||||||
|
FAIL=$((FAIL + 1))
|
||||||
|
fi
|
||||||
|
|
||||||
|
rm -f "$wrapper" "$monitor"
|
||||||
|
}
|
||||||
|
|
||||||
|
echo "=== PBS GC Four-State Tests ==="
|
||||||
|
echo "Script: $PROXMOX_MONITOR"
|
||||||
|
echo ""
|
||||||
|
|
||||||
|
echo "1. probe-failed (unparseable JSON)"
|
||||||
|
run_test "unparseable-json" 'NOT JSON {{{' "fail" "probe-failed"
|
||||||
|
|
||||||
|
echo "2. probe-failed (store not found)"
|
||||||
|
run_test "store-missing" '[{"store": "other", "last-run-endtime": 1000}]' "fail" "probe-failed"
|
||||||
|
|
||||||
|
echo "3. probe-failed (empty body)"
|
||||||
|
run_test "empty-body" "" "fail" "probe-failed"
|
||||||
|
|
||||||
|
echo "4. running (in progress - upid set, no last-run-endtime)"
|
||||||
|
run_test "running" '[{"store": "storepve-datastore", "upid": "UPID:123:1:456:gc:root@pam", "pending-bytes": 100}]' "pass" "⏳ PBS GC: running"
|
||||||
|
|
||||||
|
echo "5. stale (last run >48h)"
|
||||||
|
STALE=$(date -u -d "50 hours ago" +%s)
|
||||||
|
run_test "stale" "[{\"store\": \"storepve-datastore\", \"last-run-endtime\": $STALE, \"pending-bytes\": 200}]" "fail" "stale"
|
||||||
|
|
||||||
|
echo "6. healthy (completed <48h)"
|
||||||
|
HEALTHY=$(date -u -d "1 hour ago" +%s)
|
||||||
|
run_test "healthy" "[{\"store\": \"storepve-datastore\", \"last-run-endtime\": $HEALTHY, \"pending-bytes\": 0}]" "pass" "✅ PBS GC: healthy"
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
echo "=== Results: $PASS passed, $FAIL failed ==="
|
||||||
|
[ $FAIL -eq 0 ] && echo "✅ All passed" || echo "🔴 Some failed"
|
||||||
|
|
||||||
|
rm -f "$REPLACE_SCRIPT"
|
||||||
|
[ $FAIL -eq 0 ] && exit 0 || exit 1
|
||||||
@@ -104,6 +104,59 @@ def test_koby_ct111_is_on_storepve(ahc):
|
|||||||
assert ahc.AGENTS["koby"]["pve"] == "storepve"
|
assert ahc.AGENTS["koby"]["pve"] == "storepve"
|
||||||
|
|
||||||
|
|
||||||
|
def test_tanko_ct112_is_probed_on_minipve(ahc, monkeypatch, capsys):
|
||||||
|
# CT 112 (tanko) was live-migrated to minipve (.12) on 2026-09-27; the
|
||||||
|
# amdpve mapping made `pct status 112` fail and read as ct-unreachable.
|
||||||
|
# Execute the probe and assert the host the script actually contacts.
|
||||||
|
probes = []
|
||||||
|
monkeypatch.setattr(
|
||||||
|
ahc, "ssh",
|
||||||
|
lambda host, cmd, user="root": probes.append((host, cmd)) or "status: running",
|
||||||
|
)
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
ahc.REPORT_ONLY.clear()
|
||||||
|
try:
|
||||||
|
ahc.check_ct_liveness()
|
||||||
|
tanko_hosts = [h for h, cmd in probes if cmd == "pct status 112 2>/dev/null"]
|
||||||
|
assert tanko_hosts == ["192.168.68.12"]
|
||||||
|
finally:
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
ahc.REPORT_ONLY.clear()
|
||||||
|
|
||||||
|
|
||||||
|
def test_gpu_rtx3090_probe_uses_llmuser_not_root(ahc, monkeypatch, capsys):
|
||||||
|
# 2026-09-28: root SSH to .8 was lost when the guest was rebuilt; llmuser
|
||||||
|
# owns llama-server and can read systemctl status and the :8080 pid. A root
|
||||||
|
# probe reads as UNREACHABLE for a healthy host (the reported bug). Execute
|
||||||
|
# check_gpu_ports() against an SSH boundary that only accepts llmuser@.8 and
|
||||||
|
# assert the .8 leg does not produce the false UNREACHABLE failure.
|
||||||
|
seen = []
|
||||||
|
|
||||||
|
def fake_ssh(host, cmd, user="root"):
|
||||||
|
seen.append((host, user))
|
||||||
|
if host == "192.168.68.8" and user != "llmuser":
|
||||||
|
return None # root SSH denied -> baseline false UNREACHABLE
|
||||||
|
if cmd.startswith("systemctl is-active"):
|
||||||
|
return "active"
|
||||||
|
if cmd.startswith("ss -tlnp"):
|
||||||
|
return "48351"
|
||||||
|
if cmd.startswith("curl"):
|
||||||
|
return '{"status":"ok"}'
|
||||||
|
return None
|
||||||
|
|
||||||
|
monkeypatch.setattr(ahc, "ssh", fake_ssh)
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
try:
|
||||||
|
ahc.check_gpu_ports()
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert "gpu-unreachable:192.168.68.8" not in ahc.FAIL
|
||||||
|
assert "\u2705 gpu-rtx3090 (.8): healthy" in out
|
||||||
|
assert ("192.168.68.8", "llmuser") in seen
|
||||||
|
assert not any(host == "192.168.68.8" and user == "root" for host, user in seen)
|
||||||
|
finally:
|
||||||
|
ahc.FAIL.clear()
|
||||||
|
|
||||||
|
|
||||||
def test_report_only_legs_never_count_as_failures(ahc):
|
def test_report_only_legs_never_count_as_failures(ahc):
|
||||||
for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)):
|
for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)):
|
||||||
ahc.FAIL.clear()
|
ahc.FAIL.clear()
|
||||||
|
|||||||
Executable
+282
@@ -0,0 +1,282 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Behavioural tests for scripts/revision-preflight.sh
|
||||||
|
#
|
||||||
|
# Every case builds a throwaway clone with a real bare remote, so origin/master
|
||||||
|
# is genuine and the guard's fetch path is exercised. Nothing outside mktemp is
|
||||||
|
# touched.
|
||||||
|
#
|
||||||
|
# The pre-fix draft is kept at tests/fixtures/revision-preflight.prefix.sh and
|
||||||
|
# is run against the SAME cases, to prove these tests bite: the pre-fix guard
|
||||||
|
# exits 0 where the fixed guard exits 1.
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
REPO="$(cd "$HERE/.." && pwd)"
|
||||||
|
GUARD="$REPO/scripts/revision-preflight.sh"
|
||||||
|
PREFIX_GUARD="$HERE/fixtures/revision-preflight.prefix.sh"
|
||||||
|
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
FAILED_CASES=()
|
||||||
|
|
||||||
|
pass() { printf ' ✓ %s\n' "$1"; PASS=$((PASS + 1)); }
|
||||||
|
fail() { printf ' ✗ %s\n' "$1"; FAIL=$((FAIL + 1)); FAILED_CASES+=("$1"); }
|
||||||
|
|
||||||
|
# Build a clone with a real remote; echo the clone path.
|
||||||
|
make_clone() {
|
||||||
|
local tmp
|
||||||
|
tmp="$(mktemp -d)"
|
||||||
|
git init --bare -q "$tmp/remote.git"
|
||||||
|
git init -q "$tmp/clone"
|
||||||
|
(
|
||||||
|
cd "$tmp/clone" || exit 1
|
||||||
|
git config user.email test@example.invalid
|
||||||
|
git config user.name test
|
||||||
|
mkdir -p scripts
|
||||||
|
printf '#!/bin/bash\necho hello\n' > scripts/demo.sh
|
||||||
|
chmod +x scripts/demo.sh
|
||||||
|
git add -A
|
||||||
|
git commit -qm init
|
||||||
|
git branch -M master
|
||||||
|
git remote add origin "$tmp/remote.git"
|
||||||
|
git push -q origin master
|
||||||
|
git fetch -q origin
|
||||||
|
)
|
||||||
|
echo "$tmp/clone"
|
||||||
|
}
|
||||||
|
|
||||||
|
echo "== revision-preflight behavioural tests =="
|
||||||
|
|
||||||
|
# ── 1. match → exit 0 ────────────────────────────────────────────────────────
|
||||||
|
echo "1. matching copy"
|
||||||
|
C=$(make_clone)
|
||||||
|
if out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); then
|
||||||
|
pass "matching copy exits 0"
|
||||||
|
else
|
||||||
|
fail "matching copy should exit 0 (got $?, output: $out)"
|
||||||
|
fi
|
||||||
|
if [[ -z "$( "$GUARD" --quiet "$C/scripts/demo.sh" "$C" 2>&1 )" ]]; then
|
||||||
|
pass "--quiet prints nothing on a match"
|
||||||
|
else
|
||||||
|
fail "--quiet should print nothing on a match"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 2. mismatch → exit 1 and names both hashes ───────────────────────────────
|
||||||
|
echo "2. mismatched copy"
|
||||||
|
C=$(make_clone)
|
||||||
|
printf '#!/bin/bash\necho TAMPERED\n' > "$C/scripts/demo.sh"
|
||||||
|
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 1 ]]; then pass "mismatch exits 1"; else fail "mismatch should exit 1 (got $rc)"; fi
|
||||||
|
if [[ "$out" == *"MISMATCH"* ]]; then pass "mismatch says MISMATCH"; else fail "mismatch should say MISMATCH"; fi
|
||||||
|
if [[ "$out" == *"executed:"* && "$out" == *"merged:"* ]]; then
|
||||||
|
pass "mismatch prints both revisions"
|
||||||
|
else
|
||||||
|
fail "mismatch should print both revisions"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 3. paths under scripts/ resolve (the original basename defect) ───────────
|
||||||
|
echo "3. repo-relative path resolution"
|
||||||
|
C=$(make_clone)
|
||||||
|
if "$GUARD" --quiet "$C/scripts/demo.sh" "$C" >/dev/null 2>&1; then
|
||||||
|
pass "script under scripts/ resolves against origin/master"
|
||||||
|
else
|
||||||
|
fail "script under scripts/ must resolve (basename defect)"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 4. script that exists in NO revision → must fail ─────────────────────────
|
||||||
|
echo "4. untracked script present in no revision"
|
||||||
|
C=$(make_clone)
|
||||||
|
printf '#!/bin/bash\necho never committed\n' > "$C/scripts/ghost.sh"
|
||||||
|
out=$("$GUARD" "$C/scripts/ghost.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 1 ]]; then pass "ghost script exits 1"; else fail "ghost script must exit 1 (got $rc)"; fi
|
||||||
|
if [[ "$out" == *"does not exist in"* ]]; then
|
||||||
|
pass "ghost script says it is absent from the ref"
|
||||||
|
else
|
||||||
|
fail "ghost script should say it is absent from the ref"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
reason_of() { grep -m1 '^REASON=' <<<"$1" | cut -d= -f2-; }
|
||||||
|
|
||||||
|
# ── 5. unresolvable ref → CANNOT VERIFY (exit 2) ─────────────────────────────
|
||||||
|
echo "5. unresolvable ref"
|
||||||
|
C=$(make_clone)
|
||||||
|
out=$("$GUARD" --no-fetch --ref origin/nope "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 2 ]]; then pass "unresolvable ref exits 2 (cannot verify)"; else fail "unresolvable ref must exit 2 (got $rc)"; fi
|
||||||
|
if [[ "$(reason_of "$out")" == "cannot-verify:ref-unresolvable" ]]; then
|
||||||
|
pass "unresolvable ref names cannot-verify:ref-unresolvable"
|
||||||
|
else
|
||||||
|
fail "unresolvable ref should name its class (got: $(reason_of "$out"))"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 5b. fetch failure → CANNOT VERIFY, and it names that ─────────────────────
|
||||||
|
echo "5b. fetch failed"
|
||||||
|
C=$(make_clone)
|
||||||
|
( cd "$C" && git remote set-url origin /nonexistent/definitely-not-a-repo )
|
||||||
|
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 2 ]]; then pass "fetch failure exits 2 (cannot verify)"; else fail "fetch failure must exit 2 (got $rc)"; fi
|
||||||
|
if [[ "$(reason_of "$out")" == "cannot-verify:fetch-failed" ]]; then
|
||||||
|
pass "fetch failure names cannot-verify:fetch-failed"
|
||||||
|
else
|
||||||
|
fail "fetch failure should name its class (got: $(reason_of "$out"))"
|
||||||
|
fi
|
||||||
|
if [[ "$out" == *"bound:"* ]]; then pass "fetch failure reports the bound"; else fail "fetch failure should report the bound"; fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 5c. fetch timeout → CANNOT VERIFY, bounded (never hangs) ─────────────────
|
||||||
|
echo "5c. fetch timeout is bounded"
|
||||||
|
C=$(make_clone)
|
||||||
|
# a remote that will never answer: a fifo-backed git daemon is overkill, so use
|
||||||
|
# a black-hole address with a 1s bound and assert we return promptly.
|
||||||
|
( cd "$C" && git remote set-url origin http://10.255.255.1:9/never.git )
|
||||||
|
start=$(date +%s)
|
||||||
|
out=$("$GUARD" --fetch-timeout 1 "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||||
|
elapsed=$(( $(date +%s) - start ))
|
||||||
|
if [[ $rc -eq 2 ]]; then pass "timeout exits 2 (cannot verify)"; else fail "timeout must exit 2 (got $rc)"; fi
|
||||||
|
if [[ "$(reason_of "$out")" == "cannot-verify:fetch-failed" ]]; then
|
||||||
|
pass "timeout names cannot-verify:fetch-failed"
|
||||||
|
else
|
||||||
|
fail "timeout should name its class (got: $(reason_of "$out"))"
|
||||||
|
fi
|
||||||
|
if [[ $elapsed -le 10 ]]; then pass "timeout returned promptly (${elapsed}s, bound 1s)"; else fail "timeout did not bound the fetch (${elapsed}s)"; fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 6. missing script → MISMATCH:path-absent (exit 1) ────────────────────────
|
||||||
|
echo "6. missing script"
|
||||||
|
C=$(make_clone)
|
||||||
|
out=$("$GUARD" "$C/scripts/nope.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 1 ]]; then pass "missing script exits 1"; else fail "missing script must exit 1 (got $rc)"; fi
|
||||||
|
if [[ "$(reason_of "$out")" == "mismatch:path-absent" ]]; then
|
||||||
|
pass "missing script names mismatch:path-absent"
|
||||||
|
else
|
||||||
|
fail "missing script should name its class (got: $(reason_of "$out"))"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 7. script outside the clone → MISMATCH (exit 1) ──────────────────────────
|
||||||
|
echo "7. script outside the clone"
|
||||||
|
C=$(make_clone)
|
||||||
|
OUTSIDE=$(mktemp)
|
||||||
|
printf '#!/bin/bash\necho outside\n' > "$OUTSIDE"
|
||||||
|
out=$("$GUARD" "$OUTSIDE" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 1 ]]; then pass "outside script exits 1"; else fail "outside script must exit 1 (got $rc)"; fi
|
||||||
|
rm -f "$OUTSIDE"; rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 7b. detached HEAD is named, not reported as a raw content mismatch ───────
|
||||||
|
echo "7b. detached HEAD"
|
||||||
|
C=$(make_clone)
|
||||||
|
( cd "$C" && printf '#!/bin/bash\necho TAMPERED\n' > scripts/demo.sh \
|
||||||
|
&& git add scripts/demo.sh && git commit -qm tamper && git checkout -q --detach HEAD )
|
||||||
|
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 1 ]]; then pass "detached HEAD exits 1"; else fail "detached HEAD must exit 1 (got $rc)"; fi
|
||||||
|
if [[ "$(reason_of "$out")" == "mismatch:detached-head" ]]; then
|
||||||
|
pass "detached HEAD names mismatch:detached-head"
|
||||||
|
else
|
||||||
|
fail "detached HEAD should name its class (got: $(reason_of "$out"))"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 7c. a branch ahead of the ref is named as such, not as a raw mismatch ────
|
||||||
|
echo "7c. clone ahead of the ref"
|
||||||
|
C=$(make_clone)
|
||||||
|
( cd "$C" && git checkout -q -b feature \
|
||||||
|
&& printf '#!/bin/bash\necho FEATURE\n' > scripts/demo.sh \
|
||||||
|
&& git add scripts/demo.sh && git commit -qm feature )
|
||||||
|
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 1 ]]; then pass "clone ahead exits 1"; else fail "clone ahead must exit 1 (got $rc)"; fi
|
||||||
|
if [[ "$(reason_of "$out")" == "mismatch:clone-ahead" ]]; then
|
||||||
|
pass "clone ahead names mismatch:clone-ahead"
|
||||||
|
else
|
||||||
|
fail "clone ahead should name its class (got: $(reason_of "$out"))"
|
||||||
|
fi
|
||||||
|
if [[ "$out" == *"mid-review"* ]]; then pass "clone ahead explains it is a legitimate state"; else fail "clone ahead should explain the state"; fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 7d. a genuine content mismatch is named as content ──────────────────────
|
||||||
|
echo "7d. genuine content mismatch"
|
||||||
|
C=$(make_clone)
|
||||||
|
printf '#!/bin/bash\necho TAMPERED\n' > "$C/scripts/demo.sh"
|
||||||
|
out=$("$GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ "$(reason_of "$out")" == "mismatch:content" ]]; then
|
||||||
|
pass "hand-edit names mismatch:content"
|
||||||
|
else
|
||||||
|
fail "hand-edit should name mismatch:content (got: $(reason_of "$out"))"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 8. --no-fetch states the freshness assumption ────────────────────────────
|
||||||
|
echo "8. --no-fetch states its assumption"
|
||||||
|
C=$(make_clone)
|
||||||
|
out=$("$GUARD" --no-fetch "$C/scripts/demo.sh" "$C" 2>&1)
|
||||||
|
if [[ "$out" == *"freshness is assumed"* ]]; then
|
||||||
|
pass "--no-fetch states the freshness assumption"
|
||||||
|
else
|
||||||
|
fail "--no-fetch should state the freshness assumption"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# ── 8b. contract-run.sh defaults to warn, not enforce ───────────────────────
|
||||||
|
echo "8b. contract-run.sh default mode"
|
||||||
|
DEFAULT=$(grep -m1 'CONTRACT_REVISION_PREFLIGHT:-' "$REPO/scripts/contract-run.sh" | sed 's/.*:-//; s/}.*//')
|
||||||
|
if [[ "$DEFAULT" == "warn" ]]; then
|
||||||
|
pass "contract-run.sh defaults to warn"
|
||||||
|
else
|
||||||
|
fail "contract-run.sh default must be warn (found: '$DEFAULT')"
|
||||||
|
fi
|
||||||
|
if grep -q 'CONTRACT_REVISION_PREFLIGHT=enforce' "$REPO/scripts/contract-run.sh"; then
|
||||||
|
pass "enforce remains available and documented"
|
||||||
|
else
|
||||||
|
fail "enforce must remain documented"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── 9. the pre-fix guard must FAIL these same cases (proves the tests bite) ──
|
||||||
|
echo "9. pre-fix draft fails the same cases (bite proof)"
|
||||||
|
if [[ ! -f "$PREFIX_GUARD" ]]; then
|
||||||
|
fail "pre-fix fixture missing: $PREFIX_GUARD"
|
||||||
|
else
|
||||||
|
# 9a. repo-relative path: pre-fix drops scripts/ and cannot resolve
|
||||||
|
C=$(make_clone)
|
||||||
|
out=$("$PREFIX_GUARD" "$C/scripts/demo.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 0 && "$out" == *"could not resolve"* ]]; then
|
||||||
|
pass "pre-fix: exits 0 and cannot resolve scripts/demo.sh (defect confirmed)"
|
||||||
|
else
|
||||||
|
fail "pre-fix should exit 0 with 'could not resolve' (got rc=$rc)"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# 9b. ghost script: pre-fix passes a script that exists in no revision
|
||||||
|
C=$(make_clone)
|
||||||
|
printf '#!/bin/bash\necho never committed\n' > "$C/scripts/ghost.sh"
|
||||||
|
out=$("$PREFIX_GUARD" "$C/scripts/ghost.sh" "$C" 2>&1); rc=$?
|
||||||
|
if [[ $rc -eq 0 ]]; then
|
||||||
|
pass "pre-fix: PASSES a ghost script that exists in no revision (defect confirmed)"
|
||||||
|
else
|
||||||
|
fail "pre-fix was expected to wrongly pass the ghost script (got rc=$rc)"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
|
||||||
|
# 9c. a file that DOES exist at the repo root still works pre-fix, showing
|
||||||
|
# the defect is specific to nested paths
|
||||||
|
C=$(make_clone)
|
||||||
|
printf '#!/bin/bash\necho root\n' > "$C/rootlevel.sh"
|
||||||
|
( cd "$C" && git add rootlevel.sh && git commit -qm root && git push -q origin master && git fetch -q origin )
|
||||||
|
if "$PREFIX_GUARD" "$C/rootlevel.sh" "$C" >/dev/null 2>&1; then
|
||||||
|
pass "pre-fix: root-level path resolves (so the defect is the basename, not git)"
|
||||||
|
else
|
||||||
|
fail "pre-fix should resolve a root-level tracked file"
|
||||||
|
fi
|
||||||
|
rm -rf "$(dirname "$C")"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo
|
||||||
|
echo " passed: $PASS failed: $FAIL"
|
||||||
|
if [[ $FAIL -gt 0 ]]; then
|
||||||
|
printf ' FAILED: %s\n' "${FAILED_CASES[@]}"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "All revision-preflight tests passed."
|
||||||
Executable
+153
@@ -0,0 +1,153 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# test_secret_scan.sh — self-test for the commit-time secret guard.
|
||||||
|
#
|
||||||
|
# Run: bash tests/test_secret_scan.sh
|
||||||
|
# Exit: 0 all cases passed, 1 a case failed.
|
||||||
|
#
|
||||||
|
# WHY THIS FILE EXISTS: a scanner that is never observed to fail is not a guard.
|
||||||
|
# Every fixture below is fabricated and pattern-shaped; the test writes it to a
|
||||||
|
# temp tree (a path no allowlist entry covers) and asserts the guard FAILS. The
|
||||||
|
# same fixtures are deliberately listed in scripts/secret-allowlist.tsv, so the
|
||||||
|
# repo-wide tree scan stays quiet while a planted copy still bites — that is the
|
||||||
|
# difference between an explicit, reasoned exception and a guard trained to
|
||||||
|
# ignore a word.
|
||||||
|
#
|
||||||
|
# Only bash + coreutils + grep. No python/node: the Gitea runner executes job
|
||||||
|
# steps inside the runner container, which has neither.
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
HERE=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||||
|
ROOT=$(cd -- "$HERE/.." && pwd)
|
||||||
|
SCAN="$ROOT/scripts/secret-scan.sh"
|
||||||
|
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
LAST_OUT=""
|
||||||
|
|
||||||
|
ok() { PASS=$((PASS + 1)); echo " ✅ $1"; }
|
||||||
|
bad() { FAIL=$((FAIL + 1)); echo " ❌ $1"; }
|
||||||
|
|
||||||
|
expect_exit() { # expect_exit <want-code> <label> <cmd...>
|
||||||
|
local want="$1" label="$2"; shift 2
|
||||||
|
local rc
|
||||||
|
LAST_OUT=$("$@" 2>&1); rc=$?
|
||||||
|
if [ "$rc" -eq "$want" ]; then ok "$label (exit $rc)"; else
|
||||||
|
bad "$label (wanted exit $want, got $rc)"
|
||||||
|
printf '%s\n' "$LAST_OUT" | sed 's/^/ /' | head -8
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
expect_contains() { # expect_contains <label> <needle>
|
||||||
|
if printf '%s' "$LAST_OUT" | grep -qF -- "$2"; then ok "$1"; else
|
||||||
|
bad "$1 (output did not mention: $2)"
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
TMPROOT=$(mktemp -d)
|
||||||
|
trap 'rm -rf "$TMPROOT"' EXIT
|
||||||
|
|
||||||
|
echo "── secret-scan self-test ──"
|
||||||
|
|
||||||
|
# ── 1. Guard syntax ───────────────────────────────────────────────────────
|
||||||
|
expect_exit 0 "scanner parses with bash -n" bash -n "$SCAN"
|
||||||
|
|
||||||
|
# ── 2. Guard FAILS on planted, pattern-matching fixtures ──────────────────
|
||||||
|
mkdir -p "$TMPROOT/planted"
|
||||||
|
cat > "$TMPROOT/planted/ops.env" <<'EOF'
|
||||||
|
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
|
||||||
|
EOF
|
||||||
|
expect_exit 1 "planted sk-or-v1 key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||||
|
expect_contains "planted sk-or-v1 key names the openrouter-key rule" "[openrouter-key]"
|
||||||
|
|
||||||
|
rm -f "$TMPROOT/planted/"*
|
||||||
|
cat > "$TMPROOT/planted/curl.sh" <<'EOF'
|
||||||
|
curl -s -H "Authorization: Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb" http://example.invalid/
|
||||||
|
EOF
|
||||||
|
expect_exit 1 "planted literal Bearer token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||||
|
expect_contains "planted Bearer token names the bearer-token rule" "[bearer-token]"
|
||||||
|
|
||||||
|
rm -f "$TMPROOT/planted/"*
|
||||||
|
cat > "$TMPROOT/planted/pve.sh" <<'EOF'
|
||||||
|
AUTH="Authorization: PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555"
|
||||||
|
EOF
|
||||||
|
expect_exit 1 "planted Proxmox token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||||
|
expect_contains "planted Proxmox token names the proxmox-token rule" "[proxmox-token]"
|
||||||
|
|
||||||
|
rm -f "$TMPROOT/planted/"*
|
||||||
|
cat > "$TMPROOT/planted/deploy-key.pem" <<'EOF'
|
||||||
|
-----BEGIN OPENSSH PRIVATE KEY-----
|
||||||
|
b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW
|
||||||
|
-----END OPENSSH PRIVATE KEY-----
|
||||||
|
EOF
|
||||||
|
expect_exit 1 "planted PEM private key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||||
|
expect_contains "planted PEM key names the private-key rule" "[private-key]"
|
||||||
|
|
||||||
|
# Prose is scanned exactly like code — the original exposures were in .md files.
|
||||||
|
rm -f "$TMPROOT/planted/"*
|
||||||
|
cat > "$TMPROOT/planted/handover.md" <<'EOF'
|
||||||
|
- Admin credentials: `admin` / `correct-horse-battery-staple`
|
||||||
|
EOF
|
||||||
|
expect_exit 1 "planted prose credential line fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||||
|
expect_contains "planted prose line names the cred-prose rule" "[cred-prose]"
|
||||||
|
|
||||||
|
rm -f "$TMPROOT/planted/"*
|
||||||
|
cat > "$TMPROOT/planted/config.env" <<'EOF'
|
||||||
|
DB_PASSWORD=correct-horse-battery-staple
|
||||||
|
EOF
|
||||||
|
expect_exit 1 "planted password assignment fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||||
|
expect_contains "planted password assignment names the secret-assign rule" "[secret-assign]"
|
||||||
|
|
||||||
|
# ── 3. Guard stays QUIET on inert values and on the real tree ─────────────
|
||||||
|
mkdir -p "$TMPROOT/inert"
|
||||||
|
cat > "$TMPROOT/inert/config.yaml" <<'EOF'
|
||||||
|
api_key: not-needed
|
||||||
|
bearer_token=monitor_key
|
||||||
|
api_key: $LITELLM_API_KEY
|
||||||
|
EOF
|
||||||
|
expect_exit 0 "env refs, sentinels and variable names are not credentials" bash "$SCAN" --path "$TMPROOT/inert" --quiet
|
||||||
|
|
||||||
|
expect_exit 0 "current repo tree passes the guard" bash "$SCAN" --tree
|
||||||
|
expect_contains "tree run reports the allowlisted exceptions it applied" "allowlisted exception(s)"
|
||||||
|
|
||||||
|
# ── 4. Allowlist entries are path-explicit, not word-based ────────────────
|
||||||
|
# This exact line is allowlisted in infrastructure-control.prose.md; the same
|
||||||
|
# text at an unlisted path must still fail, proving the exception is per-file
|
||||||
|
# and reviewed, not a blanket "ignore the word vault".
|
||||||
|
rm -f "$TMPROOT/planted/"*
|
||||||
|
cat > "$TMPROOT/planted/unlisted.md" <<'EOF'
|
||||||
|
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||||
|
EOF
|
||||||
|
expect_exit 1 "allowlisted text at an unlisted path still fails" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||||
|
|
||||||
|
# ── 5. Commit-time mode: the guard blocks a STAGED credential ─────────────
|
||||||
|
# A throwaway git repo with its own copy of the scanner, so this exercises the
|
||||||
|
# real pre-commit path (--staged) without touching this repo's index.
|
||||||
|
mkdir -p "$TMPROOT/repo/scripts"
|
||||||
|
cp "$SCAN" "$TMPROOT/repo/scripts/secret-scan.sh"
|
||||||
|
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/repo/scripts/secret-patterns.tsv"
|
||||||
|
cp "$ROOT/scripts/secret-allowlist.tsv" "$TMPROOT/repo/scripts/secret-allowlist.tsv"
|
||||||
|
git -C "$TMPROOT/repo" init -q
|
||||||
|
git -C "$TMPROOT/repo" -c user.email=t@example.invalid -c user.name=test commit -q --allow-empty -m base
|
||||||
|
cat > "$TMPROOT/repo/planted.env" <<'EOF'
|
||||||
|
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
|
||||||
|
EOF
|
||||||
|
git -C "$TMPROOT/repo" add planted.env
|
||||||
|
expect_exit 1 "staged credential fails at commit time (--staged)" bash "$TMPROOT/repo/scripts/secret-scan.sh" --staged --quiet
|
||||||
|
expect_contains "staged credential names the openrouter-key rule" "[openrouter-key]"
|
||||||
|
|
||||||
|
# ── 6. Fail closed: an allowlist entry without a reason is a hard error ───
|
||||||
|
mkdir -p "$TMPROOT/scanner" "$TMPROOT/clean"
|
||||||
|
cp "$SCAN" "$TMPROOT/scanner/secret-scan.sh"
|
||||||
|
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/scanner/secret-patterns.tsv"
|
||||||
|
printf '*\t*.md\twhatever\n' > "$TMPROOT/scanner/secret-allowlist.tsv"
|
||||||
|
echo "placeholder" > "$TMPROOT/clean/ok.md"
|
||||||
|
expect_exit 2 "allowlist entry with no reason fails closed" bash "$TMPROOT/scanner/secret-scan.sh" --path "$TMPROOT/clean" --quiet
|
||||||
|
|
||||||
|
# ── Verdict ───────────────────────────────────────────────────────────────
|
||||||
|
echo ""
|
||||||
|
if [ "$FAIL" -gt 0 ]; then
|
||||||
|
echo "❌ secret-scan self-test FAILED — $PASS passed, $FAIL failed"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "✅ secret-scan self-test passed ($PASS cases)"
|
||||||
@@ -0,0 +1,209 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Behavioural tests for zulip-monitor.sh kagentz C1/C3 legs and the Result verdict.
|
||||||
|
|
||||||
|
WHY THIS FILE EXISTS: the 2026-09-19 kagentz A2A outage was correctly detected
|
||||||
|
by the monitor (C1 returned 000, the run logged an issue) but the lane's own
|
||||||
|
summarization in ops.status wrote "OK" with the note "A2A server DOWN — expected
|
||||||
|
(no credentials configured)". The optimistic verdict came from the lane, not the
|
||||||
|
script. The fix adds a C3 public-access-path leg and makes the Result line say
|
||||||
|
"INCIDENT" when issues are found, so the lane can quote it verbatim.
|
||||||
|
|
||||||
|
CONTRACT UNDER TEST:
|
||||||
|
* C1 (A2A liveness, no credential): 000 → INCIDENT.
|
||||||
|
* C3 (public access path, no credential): 502 → INCIDENT, 000 → INCIDENT,
|
||||||
|
200/302/401 → alive.
|
||||||
|
* Result verdict: when ISSUES > 0, the log's Result line says "INCIDENT",
|
||||||
|
not just "issues found".
|
||||||
|
* Healthy control: C1 401 + C3 302 → 0 issues, "all healthy".
|
||||||
|
|
||||||
|
HOW: behavioural execution using the sandbox pattern already in this repo
|
||||||
|
(tests/test_mumuni_monitor_removal.py). The sandbox copies the shipped monitor
|
||||||
|
verbatim, rewrites only its LOG constant, and runs it with stub ssh/curl on
|
||||||
|
PATH. Each test asserts from the run's own log/verdict, not from file text.
|
||||||
|
|
||||||
|
Usage: python3 -m pytest tests/test_zulip_kagentz_legs.py
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
import pathlib
|
||||||
|
import stat
|
||||||
|
import subprocess
|
||||||
|
|
||||||
|
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||||
|
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||||
|
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||||
|
|
||||||
|
TANKO_VANTAGE = "192.168.68.12" # minipve — Tanko CT 112 via pct exec
|
||||||
|
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||||
|
|
||||||
|
|
||||||
|
# ── Stub ssh: answers Tanko and Agent Zero probes by env vars ──────────
|
||||||
|
|
||||||
|
SSH_STUB = r"""#!/usr/bin/env bash
|
||||||
|
# Stub ssh: record the target host, then answer by host + remote command.
|
||||||
|
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||||
|
host=""
|
||||||
|
for a in "$@"; do
|
||||||
|
case "$a" in
|
||||||
|
*@192.168.*) host="${a##*@}" ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||||
|
cmd="${*: -1}"
|
||||||
|
case "$host" in
|
||||||
|
192.168.68.12)
|
||||||
|
case "$cmd" in
|
||||||
|
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||||
|
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||||
|
esac ;;
|
||||||
|
192.168.68.14)
|
||||||
|
case "$cmd" in
|
||||||
|
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "${AZ_A2A_EXIT:-0}" ;;
|
||||||
|
esac ;;
|
||||||
|
*)
|
||||||
|
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||||
|
esac
|
||||||
|
exit 0
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
# ── Stub curl: serves Zulip server, Abiba health, and C3 public URL ───
|
||||||
|
|
||||||
|
CURL_STUB = r"""#!/usr/bin/env bash
|
||||||
|
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||||
|
# C3 public URL probe (https://kagentz.sysloggh.net/). Record every call.
|
||||||
|
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||||
|
case "$*" in
|
||||||
|
*:9200/health*)
|
||||||
|
case " $* " in
|
||||||
|
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||||
|
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||||
|
esac ;;
|
||||||
|
*server_settings*)
|
||||||
|
printf '%s' "$SERVER_HTTP" ;;
|
||||||
|
*kagentz.sysloggh.net*)
|
||||||
|
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||||
|
esac
|
||||||
|
exit 0
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||||
|
path.write_text(body)
|
||||||
|
path.chmod(path.stat().st_mode
|
||||||
|
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||||
|
|
||||||
|
|
||||||
|
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||||
|
az_a2a_code="401", az_a2a_exit=0,
|
||||||
|
kagentz_public_code="302"):
|
||||||
|
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||||
|
|
||||||
|
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||||
|
Everything else — legs, labels, notify logic — is the shipped script.
|
||||||
|
"""
|
||||||
|
sandbox = tmp_path / "sandbox"
|
||||||
|
bindir = sandbox / "bin"
|
||||||
|
record = sandbox / "record"
|
||||||
|
bindir.mkdir(parents=True)
|
||||||
|
record.mkdir()
|
||||||
|
|
||||||
|
_write_exec(bindir / "ssh", SSH_STUB)
|
||||||
|
_write_exec(bindir / "curl", CURL_STUB)
|
||||||
|
|
||||||
|
source = ZULIP_MONITOR.read_text()
|
||||||
|
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||||
|
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||||
|
log_path = sandbox / "zulip-health-monitor.log"
|
||||||
|
script = sandbox / "zulip-monitor.sh"
|
||||||
|
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||||
|
|
||||||
|
env = dict(os.environ)
|
||||||
|
env.update({
|
||||||
|
"PATH": f"{bindir}:{env['PATH']}",
|
||||||
|
"RECORD_DIR": str(record),
|
||||||
|
"TANKO_SVC": tanko_svc,
|
||||||
|
"TANKO_HTTP": tanko_http,
|
||||||
|
"AZ_A2A_CODE": az_a2a_code,
|
||||||
|
"AZ_A2A_EXIT": str(az_a2a_exit),
|
||||||
|
"PI_HTTP": "200",
|
||||||
|
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||||
|
"SERVER_HTTP": "200",
|
||||||
|
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||||
|
})
|
||||||
|
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||||
|
capture_output=True, text=True)
|
||||||
|
return proc, record, log_path
|
||||||
|
|
||||||
|
|
||||||
|
# ── Required behavioural cases ─────────────────────────────────────────
|
||||||
|
|
||||||
|
def test_c3_502_is_incident(tmp_path):
|
||||||
|
"""C3 public leg returns 502 → the run's verdict is an INCIDENT,
|
||||||
|
the C3 line names the 502, and the run is not summarised as healthy."""
|
||||||
|
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="502")
|
||||||
|
assert proc.returncode == 0, proc.stderr
|
||||||
|
log = log_path.read_text()
|
||||||
|
|
||||||
|
# The C3 line names the 502.
|
||||||
|
assert "kagentz C3: ❌ public URL 502 (upstream refused)" in log
|
||||||
|
# The verdict is an INCIDENT, not healthy.
|
||||||
|
assert "Result: 🔴 INCIDENT" in log
|
||||||
|
assert "all healthy" not in log
|
||||||
|
# The run is not summarised as healthy.
|
||||||
|
assert "✅ 0 issues" not in log
|
||||||
|
|
||||||
|
|
||||||
|
def test_c3_000_is_incident(tmp_path):
|
||||||
|
"""C3 public leg returns 000 → INCIDENT."""
|
||||||
|
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="000")
|
||||||
|
assert proc.returncode == 0, proc.stderr
|
||||||
|
log = log_path.read_text()
|
||||||
|
|
||||||
|
# The C3 line reports the connection failure.
|
||||||
|
assert "kagentz C3: ❌ public URL down (HTTP 000)" in log
|
||||||
|
# The verdict is an INCIDENT.
|
||||||
|
assert "Result: 🔴 INCIDENT" in log
|
||||||
|
assert "all healthy" not in log
|
||||||
|
assert "✅ 0 issues" not in log
|
||||||
|
|
||||||
|
|
||||||
|
def test_healthy_control_c1_401_c3_302(tmp_path):
|
||||||
|
"""Healthy control: C1 401 plus C3 302 → 0 issues and a healthy verdict,
|
||||||
|
proving the new leg cannot cry wolf."""
|
||||||
|
proc, record, log_path = _run_monitor(tmp_path,
|
||||||
|
az_a2a_code="401",
|
||||||
|
kagentz_public_code="302")
|
||||||
|
assert proc.returncode == 0, proc.stderr
|
||||||
|
log = log_path.read_text()
|
||||||
|
|
||||||
|
# Both legs report alive.
|
||||||
|
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||||
|
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||||
|
# Zero issues, healthy verdict.
|
||||||
|
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||||
|
# No INCIDENT.
|
||||||
|
assert "INCIDENT" not in log
|
||||||
|
# No notify fired for kagentz.
|
||||||
|
assert "kagentz public URL" not in proc.stdout
|
||||||
|
assert "kagentz A2A server" not in proc.stdout
|
||||||
|
|
||||||
|
|
||||||
|
def test_c1_000_is_incident(tmp_path):
|
||||||
|
"""C1 returns 000 → INCIDENT, keeping the leg that actually caught this
|
||||||
|
outage covered behaviourally."""
|
||||||
|
proc, record, log_path = _run_monitor(tmp_path,
|
||||||
|
az_a2a_code="000",
|
||||||
|
az_a2a_exit=7,
|
||||||
|
kagentz_public_code="302")
|
||||||
|
assert proc.returncode == 0, proc.stderr
|
||||||
|
log = log_path.read_text()
|
||||||
|
|
||||||
|
# The C1 line reports the A2A down.
|
||||||
|
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||||
|
# The verdict is an INCIDENT (even though C3 is healthy).
|
||||||
|
assert "Result: 🔴 INCIDENT" in log
|
||||||
|
assert "all healthy" not in log
|
||||||
|
# The notify fired for the A2A down.
|
||||||
|
assert "kagentz A2A server DOWN" in proc.stdout
|
||||||
+73
-43
@@ -1,9 +1,9 @@
|
|||||||
---
|
---
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: zulip-health
|
name: zulip-health
|
||||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||||
version: 3.3.0
|
version: 3.4.0
|
||||||
runtime_contract: 2
|
runtime_contract: 2
|
||||||
agent: abiba
|
agent: abiba
|
||||||
report_only_agents:
|
report_only_agents:
|
||||||
@@ -28,9 +28,9 @@ session start.
|
|||||||
## Requires
|
## Requires
|
||||||
|
|
||||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
- **SSH access** to minipve (192.168.68.12) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||||
- **PM2** on localhost for pi process management
|
- **PM2** on localhost for pi process management
|
||||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200`
|
||||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||||
- **Relay access** via RA-H OS MCP for alert delivery
|
- **Relay access** via RA-H OS MCP for alert delivery
|
||||||
|
|
||||||
@@ -126,13 +126,20 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
|
|||||||
- **On critical alert**: Escalate to relay message immediately, don't wait for schedule
|
- **On critical alert**: Escalate to relay message immediately, don't wait for schedule
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
||||||
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
||||||
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
||||||
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
||||||
|
execution — it only proves the agent read the result and reported it. The actual
|
||||||
|
monitoring work happens in the host cron job.
|
||||||
|
|
||||||
|
|
||||||
### Liveness rule (scoped)
|
### Liveness rule (scoped)
|
||||||
|
|
||||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN.
|
||||||
|
|
||||||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above).
|
||||||
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
||||||
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
||||||
|
|
||||||
@@ -221,20 +228,20 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
|
|||||||
| Crash loop >10/h | Alert user |
|
| Crash loop >10/h | Alert user |
|
||||||
|
|
||||||
|
|
||||||
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
|
### Step 3: Platform B — Tanko (DSH on minipve CT 112)
|
||||||
|
|
||||||
Mumuni is out of scope for this host (see the note above): she runs on her own
|
Mumuni is out of scope for this host (see the note above): she runs on her own
|
||||||
container and is monitored on her side.
|
container and is monitored on her side.
|
||||||
|
|
||||||
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
|
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
|
||||||
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
|
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
|
||||||
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **amdpve**
|
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **minipve**
|
||||||
PVE host (**192.168.68.15**). Direct SSH to 192.168.68.122 is not a dependency
|
PVE host (**192.168.68.12**). Direct SSH to 192.168.68.122 is not a dependency
|
||||||
of this contract — per-worker key availability varies — so CT 112 probes run
|
of this contract — per-worker key availability varies — so CT 112 probes run
|
||||||
from the amdpve vantage via `pct exec`:
|
from the minipve vantage via `pct exec`:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- <command>"
|
ssh root@192.168.68.12 "pct exec 112 -- <command>"
|
||||||
```
|
```
|
||||||
|
|
||||||
> **By design (verified 2026-09-08):** the `dsh-web` gateway binds
|
> **By design (verified 2026-09-08):** the `dsh-web` gateway binds
|
||||||
@@ -246,7 +253,7 @@ ssh root@192.168.68.15 "pct exec 112 -- <command>"
|
|||||||
**B1: Gateway Service State (Tanko)**
|
**B1: Gateway Service State (Tanko)**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web"
|
ssh root@192.168.68.12 "pct exec 112 -- systemctl is-active dsh-web"
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: `active`. Anything else → gateway service down → apply the Tanko heal
|
Expected: `active`. Anything else → gateway service down → apply the Tanko heal
|
||||||
@@ -255,7 +262,7 @@ Expected: `active`. Anything else → gateway service down → apply the Tanko h
|
|||||||
**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)**
|
**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
|
ssh root@192.168.68.12 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
|
||||||
```
|
```
|
||||||
|
|
||||||
Alive = **ANY** HTTP status response from the endpoint — the expected set is
|
Alive = **ANY** HTTP status response from the endpoint — the expected set is
|
||||||
@@ -266,13 +273,13 @@ process answering `503` is running and self-heal must NOT restart-loop it.
|
|||||||
Down = connection refused (`000`) or timeout only. Statuses outside the
|
Down = connection refused (`000`) or timeout only. Statuses outside the
|
||||||
expected set are logged/reported as a warning — reported, never healed on.
|
expected set are logged/reported as a warning — reported, never healed on.
|
||||||
|
|
||||||
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)**
|
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to minipve)**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/
|
curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/
|
||||||
```
|
```
|
||||||
|
|
||||||
Fallback only — used when the monitoring node has no pct/SSH path to amdpve.
|
Fallback only — used when the monitoring node has no pct/SSH path to minipve.
|
||||||
Alive = **ANY** HTTP status response from the endpoint — healthy signals are
|
Alive = **ANY** HTTP status response from the endpoint — healthy signals are
|
||||||
`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other
|
`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other
|
||||||
status, including `404`/`5xx`, also counts alive: the endpoint is up and
|
status, including `404`/`5xx`, also counts alive: the endpoint is up and
|
||||||
@@ -397,29 +404,29 @@ ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
|
|||||||
4. Every later request through `/` presents that cookie; the token is not needed
|
4. Every later request through `/` presents that cookie; the token is not needed
|
||||||
again until the cookie expires or a new browser is used.
|
again until the cookie expires or a new browser is used.
|
||||||
|
|
||||||
**Verification** (amdpve vantage):
|
**Verification** (minipve vantage):
|
||||||
```bash
|
```bash
|
||||||
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
|
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
||||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
|
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
|
||||||
# Expected: 302
|
# Expected: 302
|
||||||
|
|
||||||
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
|
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
ssh root@192.168.68.12 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
||||||
-w '%{http_code}\n' http://192.168.68.122:8081/"
|
-w '%{http_code}\n' http://192.168.68.122:8081/"
|
||||||
# Expected: 000
|
# Expected: 000
|
||||||
|
|
||||||
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
|
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
|
||||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
||||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||||
# Expected: 200 — the minted dsh-auth-... cookie (authority
|
# Expected: 200 — the minted dsh-auth-... cookie (authority
|
||||||
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
|
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
|
||||||
|
|
||||||
# 4. Token refresh is non-disruptive and idempotent.
|
# 4. Token refresh is non-disruptive and idempotent.
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
ssh root@192.168.68.12 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
||||||
# Expected: "token unchanged; nginx not reloaded" when nothing changed
|
# Expected: "token unchanged; nginx not reloaded" when nothing changed
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -430,32 +437,32 @@ fresh cookie. Both verified live 2026-09-11.
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
|
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
|
ssh root@192.168.68.12 "pct exec 112 -- systemctl restart dsh-web"
|
||||||
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
|
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
|
||||||
# until the socket answers (any status but 000) before asserting the cookie.
|
# until the socket answers (any status but 000) before asserting the cookie.
|
||||||
for i in $(seq 1 60); do
|
for i in $(seq 1 60); do
|
||||||
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
UP=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
|
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
|
||||||
[ "$UP" != "000" ] && break
|
[ "$UP" != "000" ] && break
|
||||||
sleep 2
|
sleep 2
|
||||||
done
|
done
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||||
# Expected: 200 — the pre-restart cookie is still accepted.
|
# Expected: 200 — the pre-restart cookie is still accepted.
|
||||||
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
|
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
|
||||||
# manual run may no-op on the flock, so poll until the include carries a token
|
# manual run may no-op on the flock, so poll until the include carries a token
|
||||||
# the running process accepts (bounded wait) before the mint+reuse check.
|
# the running process accepts (bounded wait) before the mint+reuse check.
|
||||||
for i in $(seq 1 60); do
|
for i in $(seq 1 60); do
|
||||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
||||||
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
CODE=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
|
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
|
||||||
[ "$CODE" = "303" ] && break
|
[ "$CODE" = "303" ] && break
|
||||||
sleep 2
|
sleep 2
|
||||||
done
|
done
|
||||||
# Expected: 303 — the include now holds the token the running process accepts.
|
# Expected: 303 — the include now holds the token the running process accepts.
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
||||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
||||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||||
# Expected: 200 — the refreshed token minted a fresh cookie.
|
# Expected: 200 — the refreshed token minted a fresh cookie.
|
||||||
```
|
```
|
||||||
@@ -469,9 +476,10 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
|
|||||||
> monitor issued a restart for something that could not start, posting a false
|
> monitor issued a restart for something that could not start, posting a false
|
||||||
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
|
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
|
||||||
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
|
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
|
||||||
> liveness only, and a probe must never restart a platform.
|
> liveness/response and public-path access only, and a probe must never restart
|
||||||
|
> a platform.
|
||||||
|
|
||||||
**C1: A2A Server Health**
|
**C1: A2A Server Health (no credential needed)**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
|
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
|
||||||
@@ -479,9 +487,9 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
|
|||||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
|
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
|
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down (INCIDENT). Any other status → running but unexpected: log/report it, never restart.
|
||||||
|
|
||||||
**C2: A2A Response Verification**
|
**C2: A2A Response Verification (requires LITELLM_KEY)**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
|
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
|
||||||
@@ -491,15 +499,30 @@ ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:
|
|||||||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
|
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set (this is a credential issue, NOT a server-down incident).
|
||||||
|
|
||||||
|
**C3: Public Access Path (no credential needed)**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Probes the public URL that NetBird proxies to the agent-zero container.
|
||||||
|
# This is the captain's point of view: if the captain can't reach it, it's down.
|
||||||
|
# 200/302/401 = alive, 502 = proxy's "upstream refused" page (INCIDENT),
|
||||||
|
# connection failed (000) = INCIDENT. Never restarts anything.
|
||||||
|
curl -s -o /dev/null --connect-timeout 10 --max-time 15 -w '%{http_code}' https://kagentz.sysloggh.net/
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: `200` (Agent Zero login page), `302` (redirect), or `401` (auth-gated) = alive. `502` = NetBird proxy's "upstream refused" page (the container's port 80 is not listening) = INCIDENT. `000` (connection failed) = INCIDENT. Any other status → running but unexpected: log/report it, never restart.
|
||||||
|
|
||||||
**Platform C Actions**
|
**Platform C Actions**
|
||||||
|
|
||||||
| Condition | Action |
|
| Condition | Action |
|
||||||
|-----------|--------|
|
|-----------|--------|
|
||||||
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
| C1 A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||||
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
| C1 A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||||
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
|
| C2 A2A returns 401 | Check `LITELLM_KEY` is set (credential issue, NOT server-down) |
|
||||||
|
| C3 public URL returns `502` (upstream refused) | Alert only — the container's port 80 is not listening; investigate the agent-zero container |
|
||||||
|
| C3 public URL returns `000` (connection failed) | Alert only — the public path is down; investigate the NetBird proxy or the container |
|
||||||
|
| C3 public URL returns a status other than `200`/`302`/`401` | Log/report as a warning — reported, never healed on |
|
||||||
|
|
||||||
### Step 5: Global Checks
|
### Step 5: Global Checks
|
||||||
|
|
||||||
@@ -513,12 +536,19 @@ If any bot processes >50 bot-originated messages in 15min → warning.
|
|||||||
|
|
||||||
### Step 6: Compile and Report
|
### Step 6: Compile and Report
|
||||||
|
|
||||||
1. Compile all platform checks and severity
|
1. Run `scripts/zulip-monitor.sh` and take its final `Result:` line as the
|
||||||
2. Determine `overall_severity` from worst per-agent severity
|
authoritative run verdict. The verdict line is either
|
||||||
3. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
`Result: ✅ 0 issues (all healthy)` or
|
||||||
4. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
`Result: 🔴 INCIDENT — N issue(s) found`.
|
||||||
5. If any agent critical or >2 degraded: send relay message to user
|
2. Quote that `Result:` line verbatim in the status report. When it says
|
||||||
6. Update `last_check` timestamp in `### Maintains` snapshot
|
`INCIDENT`, the run MUST be reported as an incident — never summarised as
|
||||||
|
OK/healthy and never annotated as "expected".
|
||||||
|
3. Compile all platform checks and severity
|
||||||
|
4. Determine `overall_severity` from worst per-agent severity
|
||||||
|
5. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||||
|
6. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||||
|
7. If any agent critical or >2 degraded: send relay message to user
|
||||||
|
8. Update `last_check` timestamp in `### Maintains` snapshot
|
||||||
|
|
||||||
### Restart Debounce
|
### Restart Debounce
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user