Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8f1e5eebc4 | ||
|
|
bb1b65340e | ||
|
|
9e87927444 | ||
|
|
b0683e9566 | ||
|
|
dd11c8f14f | ||
|
|
0732eed329 | ||
|
|
59ed7cdbf7 | ||
|
|
5d70bbf25b | ||
|
|
ccc916d1ec | ||
|
|
48fb263d4b | ||
|
|
64790ebd19 | ||
|
|
6fa0a255df | ||
|
|
c72436b406 | ||
|
|
a17379676d | ||
|
|
63990b84f7 | ||
|
|
7a5ddb46a9 | ||
|
|
568fec2efa | ||
|
|
641a52c6da | ||
|
|
58033f39c5 | ||
|
|
6c59988e7f | ||
|
|
9a2ee6faec | ||
|
|
e71ded3c8c | ||
|
|
a12abbeb14 | ||
|
|
0b9aebca37 | ||
|
|
fa458afa26 | ||
|
|
aa3da83af5 | ||
|
|
aee2de25ac | ||
|
|
76653381ec | ||
|
|
3b32cc9658 | ||
|
|
d07c4494b5 | ||
|
|
66ad5ac89d | ||
|
|
b10fd6fc98 | ||
|
|
8ff13d38f3 | ||
|
|
a13457bcd6 | ||
|
|
f59d1a2159 | ||
|
|
da8f5f43c9 | ||
|
|
8ae4b59150 | ||
|
|
ef7f90ef5a | ||
|
|
43e891e679 | ||
|
|
933cfd223b | ||
|
|
f77d6ca1d1 | ||
|
|
f4f8a4cab8 | ||
|
|
7e257ce512 | ||
|
|
c0454811bb | ||
|
|
3c7f5d7d65 | ||
|
|
03be9b13d0 | ||
|
|
c295322c85 | ||
|
|
385f7e0623 | ||
|
|
7efbfffe44 | ||
|
|
315fcbae23 | ||
|
|
93f15709d1 | ||
|
|
1137dd4582 | ||
|
|
7400dfd833 | ||
|
|
c65f5219e1 | ||
|
|
077972fa2b | ||
|
|
d2bca5405a | ||
|
|
8a5cba8515 | ||
|
|
20f882412f | ||
|
|
30b2fe3fdc | ||
|
|
5112c566c8 | ||
|
|
83307eb9b2 | ||
|
|
8245716286 | ||
|
|
cfb6c03572 | ||
|
|
85f70f65bc | ||
|
|
a820b3f7dd | ||
|
|
0b92ab17b1 | ||
|
|
9edefe036e | ||
|
|
dd6e1e8b22 | ||
|
|
c712d4faf0 | ||
|
|
7f62f19c24 | ||
|
|
57bfe7e06a | ||
|
|
9100ea3326 | ||
|
|
0f26119859 | ||
|
|
5c1c8d7c19 | ||
|
|
8a2ea2d0d7 | ||
|
|
dae8d14880 | ||
|
|
39209c7ac9 | ||
|
|
8ed3b9c606 | ||
|
|
6c616a9e58 | ||
|
|
b9b1712ac6 | ||
|
|
6fb411613e | ||
|
|
4bdd88613b | ||
|
|
bc7a55122f | ||
|
|
713b9ce80c | ||
|
|
b451d6f81a | ||
|
|
fe6eb35291 | ||
|
|
40cf057370 | ||
|
|
5053a33e2c | ||
|
|
b4b5321011 | ||
|
|
69940bc9eb | ||
|
|
65eaffe1c6 | ||
|
|
b386bd0c19 | ||
|
|
8a4dd08b05 | ||
|
|
88b6decb31 | ||
|
|
7bf9f78fc6 | ||
|
|
dd9e68329e | ||
|
|
cd9ec6a0df | ||
|
|
1d20fbaa7f | ||
|
|
2f961d7e7a | ||
|
|
4fe4f3621d | ||
|
|
86d2987ad8 | ||
|
|
efe9381283 | ||
|
|
6a1c4967db | ||
|
|
5ed6f8179c | ||
|
|
b9322973ce | ||
|
|
fb1916707d | ||
|
|
0c7d7be5ad | ||
|
|
d28df4f4de | ||
|
|
cb26ee06d6 | ||
|
|
9a789ab76d | ||
|
|
71ceda0042 | ||
|
|
4e34b7a2a2 | ||
|
|
f9f6661dd5 | ||
|
|
cb36ff1ea5 | ||
|
|
05366bd58d | ||
|
|
e648b5ac0e | ||
|
|
38d7e8b064 | ||
|
|
e94fadfa60 | ||
|
|
ba76f2c7d3 | ||
|
|
7906b2d52d | ||
|
|
c81cf5b6f0 | ||
|
|
ef168d9690 | ||
|
|
edcf465831 | ||
|
|
89651cf37c | ||
|
|
6f40a3be60 | ||
|
|
d9368467ff | ||
|
|
e38598eea4 | ||
|
|
876b011359 | ||
|
|
d23cce89e1 | ||
|
|
9ada2b23c7 | ||
|
|
bd0065bb31 | ||
|
|
f99f7e1e34 | ||
|
|
8210fd905c | ||
|
|
b3644f0292 | ||
|
|
672bf8a912 | ||
|
|
6e612ce37b | ||
|
|
357f808a82 | ||
|
|
6272d29978 | ||
|
|
4bd6132cf5 | ||
|
|
6d65cba064 | ||
|
|
de1428b4ae | ||
|
|
4ea2d0309f | ||
|
|
7b8cc5f9ac | ||
|
|
d9e06863d8 | ||
|
|
a2edc2f56f | ||
|
|
78b501798f | ||
|
|
1f1b47f59d | ||
|
|
3d55764799 | ||
|
|
ce48070f21 | ||
|
|
9e581ab203 | ||
|
|
9cac3589cf | ||
|
|
221f9f79f3 | ||
|
|
1dc040251d | ||
|
|
21f7b6171c | ||
|
|
bb17c2120f | ||
|
|
dc42ecc235 | ||
|
|
baaac9d7c6 | ||
|
|
2f65c38213 | ||
|
|
d9eb18c024 | ||
|
|
3f07b9bbcc | ||
|
|
1a598d0fcb | ||
|
|
e3752af162 | ||
|
|
22d2c3acac | ||
|
|
5f582e2c9c | ||
|
|
b1b3b4c010 | ||
|
|
9312c1a806 | ||
|
|
5d9b9847bc | ||
|
|
d29da3cc69 | ||
|
|
2ba1016ca8 | ||
|
|
9c8637bcaf | ||
|
|
aec62f7e77 | ||
|
|
b326a8944a | ||
|
|
7730cc7c03 | ||
|
|
af27530edc | ||
|
|
832b184af6 | ||
|
|
862356bcac | ||
|
|
c26255f5ff | ||
|
|
767bd22d9c | ||
|
|
64ccf65eaf | ||
|
|
e97145c88f | ||
|
|
55f1208eb8 | ||
|
|
b9adf353ee | ||
|
|
aeb66ea22d | ||
|
|
288f74cf84 | ||
|
|
c66671dbee | ||
|
|
b80d3142aa | ||
|
|
c7af7c0689 | ||
|
|
266fa1f835 | ||
|
|
85ea1f4f3d | ||
|
|
b2a259fa23 | ||
|
|
fb185ed90a | ||
|
|
b001b657d4 | ||
|
|
a1ffeaad34 | ||
|
|
287657a77a | ||
|
|
21f9073e0b | ||
|
|
32fe7c0652 | ||
|
|
25cf2f5eef | ||
|
|
26f2301188 | ||
|
|
a6a459acc0 | ||
|
|
14bed6e916 | ||
|
|
2e70c834cb | ||
|
|
4f59b82404 | ||
|
|
8a4ee4cf5a | ||
|
|
952fca9c92 | ||
|
|
b1462f3e79 | ||
|
|
19ed186d0a | ||
|
|
88f27e75ed | ||
|
|
bbdf6c1249 | ||
|
|
1974959cc9 | ||
|
|
c59c9fb174 | ||
|
|
194e256ac5 | ||
|
|
c66d9c1e20 | ||
|
|
532250b017 | ||
|
|
f57923b4fa | ||
|
|
7ed2e4e923 | ||
|
|
91d16d2693 | ||
|
|
7ff7ce5b33 | ||
|
|
36f218e255 | ||
|
|
947e8b24e0 | ||
|
|
95b4a0e6b0 | ||
|
|
028f276be4 | ||
|
|
17751e24d1 | ||
|
|
79eeb457fc | ||
|
|
1b8186f6b9 | ||
|
|
9dd0cb18d5 | ||
|
|
4ac60e14a3 | ||
|
|
2e0b737f2d | ||
|
|
bbe9ee533f | ||
|
|
4684ee64e0 | ||
|
|
af3d364242 | ||
|
|
2716e55c16 | ||
|
|
5313e6b9ba | ||
|
|
bfdff13ae7 | ||
|
|
0e4eda0abb | ||
|
|
801de0a25c | ||
|
|
8bf32f6f0f | ||
|
|
36ae464f59 | ||
|
|
a3e97ce72b | ||
|
|
cc5fe0991c | ||
|
|
dc78604360 | ||
|
|
7bbf148778 | ||
|
|
3a25c7cce5 | ||
|
|
0aa0ea4906 | ||
|
|
19821ed6b5 | ||
|
|
403fbcdd9f | ||
|
|
0753f38cf9 | ||
|
|
274596fdd1 | ||
|
|
8c4df63db4 | ||
|
|
79a1d22c99 | ||
|
|
782831f548 | ||
|
|
143dd3f16b | ||
|
|
11076ad174 | ||
|
|
b079c02d0c | ||
|
|
09065e7dee | ||
|
|
e8b9f990b2 | ||
|
|
2dfc3e1530 | ||
|
|
79af0ae7a3 | ||
|
|
30c821469b | ||
|
|
3dcbbf1d76 | ||
|
|
aac4c7eac3 | ||
|
|
031ad814a0 | ||
|
|
ca39fead74 | ||
|
|
9b280060b7 | ||
|
|
e898048baf | ||
|
|
b01469ba18 | ||
|
|
994ae1b7ac | ||
|
|
b835986d44 | ||
|
|
d6ad016ac9 | ||
|
|
c3306e87e4 | ||
|
|
20fe5adcbc | ||
|
|
82eae77cc5 | ||
|
|
8b00a4beea | ||
|
|
19a67c6815 | ||
|
|
44f7008302 | ||
|
|
a179164f1f | ||
|
|
c4626c3512 | ||
|
|
b799d46596 | ||
|
|
3abf784538 | ||
|
|
d73084d481 | ||
|
|
9f0a940cee | ||
|
|
5c3ba31b17 | ||
|
|
d0144d6db6 | ||
|
|
1d4f6c8ebb | ||
|
|
38a32f8b32 | ||
|
|
96b0caa0b3 | ||
|
|
ea17f64bb4 |
@@ -1,19 +1,18 @@
|
||||
name: PR Pipeline — Authorize → Validate → Review → Merge
|
||||
# TRIGGER IS INTENTIONALLY UNFILTERED — DO NOT RE-ADD A `paths:` FILTER.
|
||||
#
|
||||
# This workflow previously carried `paths: ['**.prose.md', 'scripts/**.sh',
|
||||
# '**.yaml', '**.yml']` on both `push` and `pull_request`. Any PR whose diff
|
||||
# touched none of those patterns (for example a `deliverables/`-only PR, or a
|
||||
# `scripts/*.py` / `bin/*` change) therefore produced NO Gitea Actions run at
|
||||
# all: validation, lint, ai-review and the merge gate were silently skipped.
|
||||
# Validation must run for every pull request and every push to master, so the
|
||||
# trigger is deliberately unconditional.
|
||||
on:
|
||||
push:
|
||||
branches: [master]
|
||||
paths:
|
||||
- '**.prose.md'
|
||||
- 'scripts/**.sh'
|
||||
- '**.yaml'
|
||||
- '**.yml'
|
||||
pull_request:
|
||||
types: [opened, synchronize, reopened]
|
||||
paths:
|
||||
- '**.prose.md'
|
||||
- 'scripts/**.sh'
|
||||
- '**.yaml'
|
||||
- '**.yml'
|
||||
|
||||
jobs:
|
||||
auth:
|
||||
@@ -43,17 +42,21 @@ jobs:
|
||||
echo "=== Prose Contract Frontmatter Validation ==="
|
||||
FAILED=0
|
||||
for f in $(find . -name "*.prose.md" -not -path "./.git/*" -not -path "./runs/*"); do
|
||||
# NOTE: use herestrings, not `echo "$FM" | grep ...`. Under the runner's
|
||||
# `-e -o pipefail`, `grep -q` exits on first match and can SIGPIPE the
|
||||
# producer, making the pipeline report non-zero and raising a false
|
||||
# "Missing name/description" whose file set varies run to run.
|
||||
FM=$(sed -n '/^---$/,/^---$/p' "$f" | sed '1d;$d')
|
||||
[ -z "$FM" ] && { echo " ❌ $f: No YAML frontmatter"; FAILED=$((FAILED+1)); continue; }
|
||||
|
||||
KIND=$(echo "$FM" | grep '^kind:' | awk '{print $2}')
|
||||
KIND=$(grep '^kind:' <<< "$FM" | awk '{print $2}')
|
||||
case "$KIND" in
|
||||
function|responsibility|gateway|pattern|test|template|architecture|enforcement) echo " ✅ $f: kind=$KIND" ;;
|
||||
*) echo " ❌ $f: Invalid kind='$KIND'"; FAILED=$((FAILED+1)) ;;
|
||||
esac
|
||||
|
||||
echo "$FM" | grep -q '^name:' || { echo " ❌ $f: Missing name"; FAILED=$((FAILED+1)); }
|
||||
echo "$FM" | grep -q '^description:' || { echo " ❌ $f: Missing description"; FAILED=$((FAILED+1)); }
|
||||
grep -q '^name:' <<< "$FM" || { echo " ❌ $f: Missing name"; FAILED=$((FAILED+1)); }
|
||||
grep -q '^description:' <<< "$FM" || { echo " ❌ $f: Missing description"; FAILED=$((FAILED+1)); }
|
||||
done
|
||||
[ $FAILED -gt 0 ] && { echo "❌ FRONTMATTER FAILED ($FAILED error(s))"; exit 1; }
|
||||
echo "✅ Frontmatter validation passed"
|
||||
@@ -68,6 +71,18 @@ jobs:
|
||||
git fetch origin "${{ gitea.ref }}" --depth=50
|
||||
git checkout "${{ gitea.sha }}"
|
||||
|
||||
- name: Committed-credential scan (secret guard)
|
||||
run: |
|
||||
# Fails the build on a credential-shaped string in the tree. Patterns
|
||||
# live in scripts/secret-patterns.tsv; the only tolerated literal
|
||||
# examples are in scripts/secret-allowlist.tsv, each with a reason.
|
||||
# Do not turn this into a warning: a warning in a stream nobody reads
|
||||
# is how six live credentials sat in this repo for weeks.
|
||||
bash scripts/secret-scan.sh
|
||||
|
||||
- name: Secret guard self-test
|
||||
run: bash tests/test_secret_scan.sh
|
||||
|
||||
- name: Structure + regression + consistency lint
|
||||
run: bash scripts/prose-lint.sh
|
||||
|
||||
|
||||
@@ -1 +1,2 @@
|
||||
__pycache__/
|
||||
state/host-disk-bands.json
|
||||
|
||||
@@ -51,11 +51,17 @@ Two incidents taught us this:
|
||||
- `/grafana/` nginx route — was reverted Jul 2, must not reappear
|
||||
- `CT 122` or `CT 123` as CT ID labels — don't exist in the cluster
|
||||
- These rules are hardcoded in `scripts/prose-lint.sh`
|
||||
- **Committed-credential guard:** `scripts/secret-scan.sh` FAILS the build on
|
||||
credential-shaped strings (patterns in `scripts/secret-patterns.tsv`, prose
|
||||
included). Tolerated literals are listed one-per-example with a reason in
|
||||
`scripts/secret-allowlist.tsv`; never allowlist a live credential. It runs in
|
||||
the CI lint job, in `scripts/prose-lint.sh`, and via
|
||||
`bash scripts/secret-scan.sh --staged` before committing.
|
||||
|
||||
### Stage 3 — AI Review
|
||||
- Diff is sent to `syslog-auto` model via LiteLLM
|
||||
- Review checks against infrastructure-control ground truth:
|
||||
- CT IDs match PVE cluster (100-117, no 122/123)
|
||||
- CT IDs match the PVE cluster inventory in `infrastructure-control.prose.md` Appendix B (100-120 with gaps; no 122/123)
|
||||
- Grafana is direct LAN :3001, NOT behind nginx
|
||||
- Zulip is CT 117 on storepve (bridge IP .19)
|
||||
- Strix Halo :8080 is firewalled to .116 only
|
||||
@@ -67,10 +73,10 @@ Two incidents taught us this:
|
||||
|
||||
| Contract | Sensitivity | Who can change |
|
||||
|----------|------------|----------------|
|
||||
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba only (after live verification) |
|
||||
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba, Tanko (Tanko maintains its own CT row) |
|
||||
| `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only |
|
||||
| `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko |
|
||||
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni |
|
||||
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni, Tanko |
|
||||
| Other contracts | Normal | Any registered agent |
|
||||
| `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only |
|
||||
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
<!-- Points Claude at AGENTS.md via import; edit AGENTS.md, not this file. -->
|
||||
@AGENTS.md
|
||||
@@ -20,7 +20,8 @@ An agent ran `pct set` without checking `pct config` first and changed the IP
|
||||
to the wrong value, breaking Zulip. The staleness was harmless until acted on.
|
||||
|
||||
**Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents:
|
||||
Abiba, Tanko, Mumuni, Koby, Koonimo). ALL infrastructure mutations MUST go
|
||||
Abiba, Tanko, Mumuni, Koby, Koonimo — Tanko runs on DSH/DeepSeek Harness since
|
||||
2026-08-27 but safe-mutate enforcement still applies). ALL infrastructure mutations MUST go
|
||||
through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up
|
||||
--force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an
|
||||
auditable protocol violation. The wrapper runs a verify command, optionally
|
||||
@@ -87,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
|
||||
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
|
||||
|
||||
# Configure an agent with a different auxiliary model
|
||||
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b
|
||||
prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision
|
||||
```
|
||||
|
||||
### Option B: Manual Execution
|
||||
@@ -115,7 +116,7 @@ Run on trigger or schedule. Maintain persistent world-model state across runs.
|
||||
| `zulip-health` | Zulip | Checks Zulip connectivity, message flow, and bot responsiveness. |
|
||||
| `zulip-mention-reliability` | Zulip | Diagnoses and fixes @mention detection issues in Zulip. |
|
||||
| `zulip-approval-fix` | Zulip | Fixes broken /approve and /deny slash commands for Hermes agents. |
|
||||
| `litellm-self-heal` | LiteLLM | Consolidated health check + self-healing for the full nginx → LiteLLM → GPU chain. Verifies 8 containers, 3 GPUs, model inference, and agent keys. Applies 9 remediation rules. (litellm-health merged into this contract 2026-07-09.) |
|
||||
| `litellm-self-heal` | LiteLLM | Applies remediation rules for LiteLLM stack failures detected by `litellm-health` (full nginx → LiteLLM → GPU chain). 9 remediation rules. |
|
||||
| `gpu-fleet` | GPU | Manages the GPU inference fleet: model deployment, registration, health checks, LiteLLM sync. |
|
||||
| `gpu-monitor` | GPU | Comprehensive GPU fleet monitor — polls sidecars, router, LiteLLM every 15s, renders SSE dashboard. |
|
||||
| `proxmox-monitor` | Infra | Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. |
|
||||
@@ -148,7 +149,7 @@ Called on-demand as single-render tools.
|
||||
| Contract | Description |
|
||||
|---|---|
|
||||
| `litellm-api-keys` | Manages LiteLLM API keys for agent identity. Create, rotate, verify, and list agent keys. References gpu-fleet for current key inventory. |
|
||||
| `litellm-health` | ⚠️ **DEPRECATED** — consolidated into `litellm-self-heal` (2026-07-09). Retained for reference only. |
|
||||
| `litellm-health` | LiteLLM health check: public vs backend surfaces, CT 116 containers, GPU fleet, one model per GPU host, and agent keys. Owner of the probes; `litellm-self-heal` owns remediation. |
|
||||
| `infrastructure-monitoring` | Target-state for Prometheus + GPU exporters + Grafana. Core stack deployed, GPU exporters NOT live. |
|
||||
| `stirling-pdf-agent-access` | Documents the Stirling-PDF API access pattern for agents — global API key, 12 operations, curl examples. Agents use the `stirling-pdf-api` shared skill for templates. |
|
||||
| `hello-world` | Minimal test contract — verifies the OpenProse execution pipeline works. |
|
||||
@@ -161,7 +162,7 @@ Companion shell scripts that contracts delegate to.
|
||||
|---|---|
|
||||
| `daily-infra-report.py` | Generates the daily infrastructure dashboard (HTML email to jerome@sysloggh.com). |
|
||||
| `pm2-self-heal.sh` | Shell companion to the pm2-self-heal contract — restarts crashed PM2 processes (runs every 5 min). |
|
||||
| `agent-health-check.py` | Consolidated agent health: LiteLLM key validation + GPU port conflict + streaming checks (every 10 min). Replaced zulip-monitor.sh. |
|
||||
| `agent-health-check.py` | Consolidated agent health: LiteLLM key validation + GPU port conflict + streaming checks (every 10 min). |
|
||||
|
||||
## Contract Structure
|
||||
|
||||
|
||||
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: function
|
||||
name: abiba-zulip-restore
|
||||
description: >
|
||||
@@ -11,6 +13,7 @@ version: 1.0.0
|
||||
status: active
|
||||
runtime_contract: 2
|
||||
---
|
||||
---
|
||||
|
||||
# Abiba Zulip Restore — Resume pi Zulip Communication
|
||||
|
||||
@@ -31,7 +34,7 @@ verification, and DM loopback testing.
|
||||
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
|
||||
| PM2 process name | abiba-zulip | ✅ |
|
||||
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
|
||||
| Default model | deepseek-v4-pro | ✅ (settings.json) |
|
||||
| Default model | syslog-auto | ✅ (settings.json) |
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -304,6 +307,7 @@ module.exports = {
|
||||
};
|
||||
```
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
|
||||
|
||||
@@ -0,0 +1,135 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: agent-health-check
|
||||
description: >
|
||||
Consolidated agent health verification for the LiteLLM + GPU + Zulip +
|
||||
gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours
|
||||
via cron and on-demand via "run contract: agent-health-check". Verifies:
|
||||
LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway
|
||||
liveness, CT liveness, gateway log health, config YAML integrity,
|
||||
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
|
||||
title: Agent Health Check — Consolidated
|
||||
version: 1.0.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
---
|
||||
|
||||
# Agent Health Check
|
||||
|
||||
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
|
||||
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only.
|
||||
|
||||
**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
|
||||
|
||||
## Requires
|
||||
|
||||
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
|
||||
- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
|
||||
- **Python 3** for script execution
|
||||
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
|
||||
|
||||
## Maintains
|
||||
|
||||
- last_check: timestamp — When the last full diagnostic ran
|
||||
- overall_severity: "healthy" | "degraded" | "critical"
|
||||
- liteLLM_keys: map of agent → key validity
|
||||
- gpu_ports: map of host → port conflict status
|
||||
- agents: map of agent → streaming health + gateway liveness
|
||||
- ct_liveness: map of CT → active status
|
||||
- config_integrity: map of config file → valid/invalid
|
||||
|
||||
## Execution
|
||||
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool
|
||||
calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Run the consolidated health check script
|
||||
python3 /root/scripts/agent-health-check.py --json
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the script
|
||||
executed from** so a stale-consumer report is distinguishable from a real fault
|
||||
at read time. Summarize actual results from each check. Apply the standing probe
|
||||
rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
|
||||
|
||||
**Expected output**: JSON with `overall_severity` field. If `healthy`, report
|
||||
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
||||
failures and their severity.
|
||||
|
||||
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
|
||||
MUST include one clause per check leg, in every state: healthy, degraded/warn,
|
||||
skipped, or failed. A missing leg must never look the same as a healthy leg.
|
||||
Required legs and their templates in every state:
|
||||
|
||||
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
|
||||
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
|
||||
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
|
||||
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
|
||||
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
|
||||
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
|
||||
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
|
||||
- The GPU leg has six non-healthy states the code can produce:
|
||||
(i) `gpu-unreachable:{host}` — SSH probe failed;
|
||||
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
|
||||
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
|
||||
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
|
||||
(v) unit active, /health body contains "error" — error response;
|
||||
(vi) unit active, /health body unrecognised — unknown health.
|
||||
In every case the failing host and reason must be named.
|
||||
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
|
||||
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
|
||||
- skipped: `CTs: SKIPPED (SSH access unavailable)`
|
||||
- `Vault secrets: 3/3 present`
|
||||
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
|
||||
- skipped: `Vault secrets: SKIPPED (vault not configured)`
|
||||
|
||||
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
|
||||
long as the host is identifiable from context; the full `probe-failed: <target>
|
||||
<kind>` form is required when a leg reports a failure in the detail section.
|
||||
|
||||
### Probe Shape (per standing rules from 1150.msg)
|
||||
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||
service answered — report the code, never "down". A redirect is not a failure.
|
||||
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
2. **A failed probe is never a service verdict.** Print
|
||||
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
||||
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
||||
then report.
|
||||
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
||||
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
||||
(retried at 25s: also timeout)" is actionable.
|
||||
|
||||
## Strategies
|
||||
|
||||
### When LiteLLM keys are invalid
|
||||
Report the specific agent + key name. Do not attempt to fix — credential
|
||||
rotation is a separate operation.
|
||||
|
||||
### When GPU port conflicts are detected
|
||||
Report the conflicting ports and processes. Do not kill processes — that's a
|
||||
destructive action requiring captain approval.
|
||||
|
||||
### When gateway liveness is degraded
|
||||
Report the specific CT + gateway status. Do not restart unless the restart
|
||||
debounce window has passed.
|
||||
|
||||
### When CT liveness is down
|
||||
Report the specific CT. Do not restart — that's a destructive action.
|
||||
|
||||
### When config YAML is invalid
|
||||
Report the specific file + parse error. Do not fix — that's a config change.
|
||||
|
||||
### When gateway log health is degraded
|
||||
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
|
||||
|
||||
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
|
||||
|
||||
## Continuity
|
||||
|
||||
- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
|
||||
- **On `agent-health` command**: Run on-demand and report to user
|
||||
- **On critical alert**: Escalate to relay message immediately
|
||||
@@ -0,0 +1,186 @@
|
||||
# Agent Zero Issue Fix Summary
|
||||
|
||||
**Date**: 2026-09-01
|
||||
**Agent**: Agent Zero (Docker container on kagentz CT105)
|
||||
**Issue**: AuthenticationError + Telegram conflicts
|
||||
**Status**: ✅ RESOLVED
|
||||
|
||||
---
|
||||
|
||||
## Problems Identified
|
||||
|
||||
### 1. OpenRouter Authentication Error (CRITICAL)
|
||||
```
|
||||
litellm.exceptions.AuthenticationError: OpenrouterException -
|
||||
{"error":{"message":"User not found.","code":401}}
|
||||
```
|
||||
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
||||
|
||||
**Old Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
**New Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
|
||||
### 2. Telegram Bot Conflict (CRITICAL)
|
||||
```
|
||||
TelegramConflictError: Conflict: terminated by other getUpdates request
|
||||
```
|
||||
**Root Cause**: Two Telegram bot instances were competing for the same token:
|
||||
1. Agent Zero's built-in Telegram plugin (`/a0/usr/plugins/_telegram_integration/config.json`)
|
||||
2. Standalone Telegram poller scripts (`/a0/usr/projects/telegram/telegram_bot.py`)
|
||||
|
||||
Both were using token `8476855065:***` in polling mode.
|
||||
|
||||
**Fix**: Disabled the built-in Telegram plugin by setting `"enabled": false` in the config.
|
||||
|
||||
### 3. MCP Service Connectivity Issues (SEVERE)
|
||||
```
|
||||
McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 seconds.
|
||||
```
|
||||
**Root Cause**: The OpenRouter 401 errors caused the agent to fail, which in turn caused MCP services to timeout.
|
||||
|
||||
**Status**: ✅ RESOLVED with OpenRouter key fix.
|
||||
|
||||
---
|
||||
|
||||
## Fixes Applied
|
||||
|
||||
### Fix 1: Update OpenRouter Key
|
||||
```bash
|
||||
# Container .env update
|
||||
sudo docker exec agent-zero bash -c '
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»|" /a0/usr/.env
|
||||
'
|
||||
```
|
||||
|
||||
**Verification**:
|
||||
```bash
|
||||
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||
-H "Authorization: Bearer «vault: agents/production OPENROUTER_API_KEY»" | python3 -m json.tool
|
||||
```
|
||||
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
||||
|
||||
### Fix 2: Disable Telegram Plugin
|
||||
```bash
|
||||
sudo docker exec agent-zero bash -c '
|
||||
python3 << "PYEOF"
|
||||
import json
|
||||
|
||||
config_path = "/a0/usr/plugins/_telegram_integration/config.json"
|
||||
with open(config_path) as f:
|
||||
config = json.load(f)
|
||||
|
||||
config["bots"][0]["enabled"] = False
|
||||
|
||||
with open(config_path, "w") as f:
|
||||
json.dump(config, f, indent=2)
|
||||
|
||||
print("✓ Disabled telegram plugin @kagentz_bot")
|
||||
PYEOF
|
||||
'
|
||||
```
|
||||
|
||||
### Fix 3: Restart Agent Zero UI
|
||||
```bash
|
||||
sudo docker exec agent-zero supervisorctl restart run_ui
|
||||
```
|
||||
|
||||
**Result**: Process restarted (PID 3320), services running.
|
||||
|
||||
### Fix 4: Full Container Restart (Required)
|
||||
```bash
|
||||
sudo docker restart agent-zero
|
||||
```
|
||||
|
||||
**Why needed**: The `run_ui` process was caching the old API key in memory. A full container restart was required to force Agent Zero to reload the `.env` file with the new OpenRouter key.
|
||||
|
||||
**Result**: All services restarted cleanly, no more 401 errors.
|
||||
|
||||
### Fix 5: Update Stale `.env.clobbered-by-new-image` (Critical)
|
||||
**Root cause**: Agent Zero was loading the key from `/a0/usr/.env.clobbered-by-new-image` (line 28) instead of the main `/a0/usr/.env` (line 72). The clobbered file still had the old, stale key.
|
||||
|
||||
**Fix**:
|
||||
```bash
|
||||
KEY=$(grep "^API_KEY_OPENROUTER=" /a0/usr/.env | cut -d"=" -f2-)
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=$KEY|" /a0/usr/.env.clobbered-by-new-image
|
||||
```
|
||||
|
||||
**Lesson**: When updating Agent Zero's `.env`, check BOTH files:
|
||||
- `/a0/usr/.env` (main)
|
||||
- `/a0/usr/.env.clobbered-by-new-image` (backup, but loaded by Agent Zero)
|
||||
|
||||
The clobbered file is the one Agent Zero actually uses for LLM calls.
|
||||
|
||||
---
|
||||
|
||||
## Infrastructure Documentation
|
||||
|
||||
### New Contract Created
|
||||
**File**: `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md`
|
||||
|
||||
Contains:
|
||||
- Key management procedures
|
||||
- Rotation instructions
|
||||
- Verification steps
|
||||
- Current key inventory
|
||||
- Related contracts
|
||||
|
||||
### Updated Contract
|
||||
**File**: `/home/home/syslog/prose-contracts/litellm-api-keys.prose.md`
|
||||
|
||||
Added section:
|
||||
- Agent Zero OpenRouter integration
|
||||
- Key storage locations
|
||||
- Model configuration
|
||||
- Why not LiteLLM proxy
|
||||
- Rotation procedure
|
||||
|
||||
---
|
||||
|
||||
## Current State
|
||||
|
||||
| Component | Status | Details |
|
||||
|-----------|--------|---------|
|
||||
| **OpenRouter Key** | ✅ Valid | `«vault: agents/production OPENROUTER_API_KEY»` user verified, |
|
||||
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
||||
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
||||
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
||||
| **Services** | ✅ All UP | run_ui, run_tunnel_api, run_searxng, run_cron, the_listener |
|
||||
|
||||
---
|
||||
|
||||
## Related Files
|
||||
|
||||
| Path | Purpose |
|
||||
|------|---------|
|
||||
| `/a0/usr/.env` | Container key storage |
|
||||
| `/a0/usr/plugins/_telegram_integration/config.json` | Telegram plugin config |
|
||||
| `/a0/usr/plugins/_model_config/presets.yaml` | Model selection (moonshotai/kimi-k3) |
|
||||
| `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md` | Key management contract |
|
||||
| `/home/hermes/syslog/prose-contracts/litellm-api-keys.prose.md` | Fleet key inventory |
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. **Sync key to Infisical vault** (optional, currently .env fallback only)
|
||||
2. **Monitor usage** — Check OpenRouter dashboard for daily/weekly spend
|
||||
3. **Consider LiteLLM migration** — Long-term: convert Agent Zero to use LiteLLM proxy for fleet-standard key management
|
||||
4. **Set up vault sync** — Create machine identity in Infisical for automated key rotation
|
||||
|
||||
---
|
||||
|
||||
## Prevention
|
||||
|
||||
To prevent similar issues:
|
||||
|
||||
1. **Always verify API keys** against their providers before using
|
||||
2. **Keep fleet-wide key inventory** updated in prose contracts
|
||||
3. **Rotate keys on schedule** (quarterly hygiene, not on-demand only)
|
||||
4. **Test key changes** in staging before production rollout
|
||||
5. **Document key locations** in both code and prose contracts
|
||||
|
||||
---
|
||||
|
||||
**Verified by**: Mumuni 🦅
|
||||
**Last updated**: 2026-09-01
|
||||
**Session**: 1
|
||||
@@ -0,0 +1,129 @@
|
||||
---
|
||||
kind: function
|
||||
name: agent-zero-openrouter-key
|
||||
description: >
|
||||
Manages the OpenRouter API key for Agent Zero (Docker container on kagentz .14).
|
||||
Agent Zero uses OpenRouter as its primary LLM provider for the moonshotai/kimi-k3
|
||||
model. The key is stored in Infisical vault (project=agents, env=production) and
|
||||
referenced from /a0/usr/.env in the container. Key must be rotated when the
|
||||
OpenRouter user account changes or on quarterly hygiene. Last verified: 2026-09-01.
|
||||
---
|
||||
|
||||
## Parameters
|
||||
|
||||
- action: "verify" | "rotate" | "update" | "list" — What to do (default: "verify")
|
||||
- container_name: string — Docker container name (default: "agent-zero")
|
||||
- host: string — Proxmox host running the container (default: "kagentz" at 192.168.68.14)
|
||||
- env_path: string — Path to .env file in container (default: "/a0/usr/.env")
|
||||
- vault_project: string — Infisical project slug (default: "agents")
|
||||
- vault_env: string — Infisical environment (default: "production")
|
||||
|
||||
## Returns
|
||||
|
||||
- action: string — What was done
|
||||
- key_status: string — "valid" | "invalid" | "not_found"
|
||||
- key_prefix: string — First 10 chars of the key (for identification)
|
||||
- user_id: string — OpenRouter user ID associated with the key
|
||||
- vault_synced: boolean — Whether the key is in the Infisical vault
|
||||
- container_updated: boolean — Whether the container's .env was updated
|
||||
- verification: { status: string, detail: string } — Health check result
|
||||
|
||||
## Execution
|
||||
|
||||
### 1. Verify the key
|
||||
|
||||
1. **Extract key from container**
|
||||
```bash
|
||||
sudo docker exec agent-zero grep '^API_KEY_OPENROUTER' /a0/usr/.env | cut -d'=' -f2-
|
||||
```
|
||||
|
||||
2. **Test against OpenRouter API**
|
||||
```bash
|
||||
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||
-H "Authorization: Bearer <key>" | python3 -m json.tool
|
||||
```
|
||||
Expected: HTTP 200, JSON with `data.label` and `data.is_free_tier`
|
||||
|
||||
3. **Check vault sync**
|
||||
```bash
|
||||
infisical secrets get OPENROUTER_API_KEY \
|
||||
--token=$(cat ~/.infisical-token) \
|
||||
--projectId=agents \
|
||||
--env=production \
|
||||
--domain=https://vault.sysloggh.net
|
||||
```
|
||||
|
||||
4. **Return status**
|
||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-synthetic...", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
||||
- If vault secret is missing: `{ vault_synced: false }`
|
||||
|
||||
### 2. Rotate the key
|
||||
|
||||
1. **Generate new key** in OpenRouter UI or via API
|
||||
2. **Update container .env**
|
||||
```bash
|
||||
sudo docker exec agent-zero sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env
|
||||
```
|
||||
3. **Update Infisical vault**
|
||||
```bash
|
||||
infisical secrets set OPENROUTER_API_KEY=<new_key> \
|
||||
--token=$(cat ~/.infisical-token) \
|
||||
--projectId=agents \
|
||||
--env=production \
|
||||
--domain=https://vault.sysloggh.net
|
||||
```
|
||||
4. **Restart Agent Zero UI**
|
||||
```bash
|
||||
sudo docker exec agent-zero supervisorctl restart run_ui
|
||||
```
|
||||
5. **Verify** — Run "verify" action again
|
||||
|
||||
### 3. Update (key changed but no rotation)
|
||||
|
||||
1. **Update container .env** (same as rotate step 2)
|
||||
2. **Sync vault** (same as rotate step 3)
|
||||
3. **Restart run_ui** (same as rotate step 4)
|
||||
|
||||
## Current Key Inventory
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| **Key Prefix** | `«vault: agents/production OPENROUTER_API_KEY»` |
|
||||
| **Full Key** | `«vault: agents/production OPENROUTER_API_KEY»` (in vault + /a0/usr/.env) |
|
||||
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
||||
| **Free Tier** | No |
|
||||
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
||||
| **Last Verified** | 2026-09-01 |
|
||||
| **Vault Sync** | ⏳ Pending (service token not on kagentz) |
|
||||
|
||||
## Key Rotation Log
|
||||
|
||||
| Date | Action | Notes |
|
||||
|------|--------|-------|
|
||||
| 2026-09-01 | fix-401 | Old key `«vault: agents/production OPENROUTER_API_KEY»…` returned 401 "User not found". Replaced with new key `«vault: agents/production OPENROUTER_API_KEY»…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||
|
||||
## Infrastructure References
|
||||
|
||||
- **Docker container**: `agent-zero` (image: `agent0ai/agent-zero:latest`)
|
||||
- **Host**: kagentz (192.168.68.14, Proxmox LXC CT105)
|
||||
- **Volume**: `/var/lib/docker/volumes/agent_zero/_data` → `/a0/usr`
|
||||
- **Config path**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=…`)
|
||||
- **Model preset**: "Cost Efficient" (uses `openrouter/moonshotai/kimi-k3`)
|
||||
- **Model config**: `/a0/usr/plugins/_model_config/config.json`
|
||||
|
||||
## Verification Before Acting
|
||||
|
||||
**Key is a lead, not a fact.** Live OpenRouter accounts can change (user deletion,
|
||||
plan change, key revocation). Before acting on this contract:
|
||||
|
||||
1. Verify the key against OpenRouter's `/auth/key` endpoint
|
||||
2. Check the user ID matches the expected account
|
||||
3. Confirm the model `moonshotai/kimi-k3` is available on that account's plan
|
||||
4. Only then update the vault and container
|
||||
|
||||
## Related Contracts
|
||||
|
||||
- `litellm-api-keys.prose.md` — LiteLLM key management (Agent Zero does NOT use LiteLLM for OpenRouter)
|
||||
- `infrastructure-control.prose.md` — Proxmox topology, container locations
|
||||
- `gpu-fleet.prose.md` — Fleet-wide agent key inventory (add Agent Zero here)
|
||||
+139
-24
@@ -36,6 +36,59 @@ def warn(rule, message):
|
||||
WARNINGS.append(f"[{rule}] {message}")
|
||||
|
||||
|
||||
# Derivation rule: a model name is any scalar under a mapping key named `model` or
|
||||
# `model_name`, at any depth. The top-level `model:` SECTION is the exception where the model
|
||||
# name lives under `default`/`model`/`model_name` inside that section, so it is descended
|
||||
# specially. The only other exception is key `models` (litellm key-generation params carry a
|
||||
# list of model names). EXTEND THE ALLOWLIST for a new exception; do NOT add another field by
|
||||
# hand.
|
||||
MODEL_KEYS = ("model", "model_name")
|
||||
MODEL_SECTION_KEYS = ("default", "model", "model_name")
|
||||
MODEL_LIST_KEYS = ("models",)
|
||||
|
||||
|
||||
def _iter_model_values(node, path=""):
|
||||
"""Yield (path, value) for every model-name-bearing scalar in a config."""
|
||||
if isinstance(node, dict):
|
||||
for key, value in node.items():
|
||||
child = f"{path}.{key}" if path else key
|
||||
if key in MODEL_KEYS:
|
||||
if isinstance(value, dict):
|
||||
for subkey in MODEL_SECTION_KEYS:
|
||||
subvalue = value.get(subkey)
|
||||
if isinstance(subvalue, str):
|
||||
yield (f"{child}.{subkey}", subvalue)
|
||||
for subkey, subvalue in value.items():
|
||||
if isinstance(subvalue, (dict, list)):
|
||||
yield from _iter_model_values(subvalue, f"{child}.{subkey}")
|
||||
elif isinstance(value, list):
|
||||
yield from _iter_model_values(value, child)
|
||||
else:
|
||||
yield (child, value)
|
||||
elif key in MODEL_LIST_KEYS:
|
||||
yield from _iter_model_list(value, child)
|
||||
elif isinstance(value, (dict, list)):
|
||||
yield from _iter_model_values(value, child)
|
||||
elif isinstance(node, list):
|
||||
for i, item in enumerate(node):
|
||||
yield from _iter_model_values(item, f"{path}[{i}]")
|
||||
|
||||
|
||||
def _iter_model_list(node, path):
|
||||
"""Yield scalars under an allowlisted `models` key (list of names or list of dicts)."""
|
||||
if isinstance(node, list):
|
||||
for i, item in enumerate(node):
|
||||
yield from _iter_model_list(item, f"{path}[{i}]")
|
||||
elif isinstance(node, dict):
|
||||
for key, value in node.items():
|
||||
if key in MODEL_KEYS and isinstance(value, str):
|
||||
yield (f"{path}.{key}", value)
|
||||
elif isinstance(value, (dict, list)):
|
||||
yield from _iter_model_list(value, f"{path}.{key}")
|
||||
else:
|
||||
yield (path, node)
|
||||
|
||||
|
||||
def audit(path):
|
||||
with open(path) as f:
|
||||
cfg = yaml.safe_load(f)
|
||||
@@ -61,12 +114,22 @@ def audit(path):
|
||||
)
|
||||
|
||||
# --- Rule 5: Main Config Base URL ---
|
||||
expected_base = "http://192.168.68.116/v1"
|
||||
check(
|
||||
model.get("base_url") == expected_base,
|
||||
"Rule 5",
|
||||
f"model.base_url must be {expected_base} (got {model.get('base_url')!r}) — /v1 not /litellm/v1",
|
||||
)
|
||||
# Canonical internal base (hermes-key-enforcement.prose.md:19/38/57/84/91/96)
|
||||
# and public base (serves /v1 only, per 2026-09-19 probe from CT 116).
|
||||
# The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only.
|
||||
# FAIL anything else (do not widen to accept any path ending in /v1).
|
||||
# Internal /v1 is non-canonical but working (authenticated via nginx), so WARN not FAIL.
|
||||
canonical_internal = "http://192.168.68.116/litellm/v1"
|
||||
public_host = "https://litellm.sysloggh.net/v1"
|
||||
non_canonical_internal = "http://192.168.68.116/v1"
|
||||
allowed_bases = (canonical_internal, public_host)
|
||||
actual_base = model.get("base_url")
|
||||
if actual_base in allowed_bases:
|
||||
check(True, "Rule 5", f"model.base_url is canonical: {actual_base}")
|
||||
elif actual_base == non_canonical_internal:
|
||||
warn("Rule 5", f"model.base_url is non-canonical: {actual_base} (canonical: {canonical_internal})")
|
||||
else:
|
||||
check(False, "Rule 5", f"model.base_url must be one of {allowed_bases} (got {actual_base!r})")
|
||||
|
||||
# --- Rule 6: max_tokens Is Required ---
|
||||
check(
|
||||
@@ -89,15 +152,16 @@ def audit(path):
|
||||
)
|
||||
|
||||
# --- Rule 8: GPU Workload Distribution ---
|
||||
# gpu-light (and gemma-4-12b) were retired 2026-09-12; the RTX 5070 stable alias is gpu-vision.
|
||||
check(
|
||||
aux.get("vision", {}).get("model") == "gpu-light",
|
||||
aux.get("vision", {}).get("model") == "gpu-vision",
|
||||
"Rule 8",
|
||||
f"auxiliary.vision.model must be gpu-light (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
|
||||
f"auxiliary.vision.model must be gpu-vision (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
|
||||
)
|
||||
check(
|
||||
aux.get("web_extract", {}).get("model") == "gpu-light",
|
||||
aux.get("web_extract", {}).get("model") == "gpu-vision",
|
||||
"Rule 8",
|
||||
f"auxiliary.web_extract.model must be gpu-light (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
|
||||
f"auxiliary.web_extract.model must be gpu-vision (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
|
||||
)
|
||||
|
||||
# --- Rule 9: Compression Threshold ---
|
||||
@@ -109,7 +173,7 @@ def audit(path):
|
||||
check(
|
||||
comp.get("max_context_window") == 131072,
|
||||
"Rule 9",
|
||||
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity",
|
||||
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)",
|
||||
)
|
||||
|
||||
# --- Rule 10: Default Model Must Be syslog-auto ---
|
||||
@@ -178,23 +242,74 @@ def audit(path):
|
||||
f"custom_providers[0].base_url must end with /v1 (got {cp.get('base_url')!r})",
|
||||
)
|
||||
|
||||
# --- No raw model names (Rule 7/8 spirit) ---
|
||||
raw_names = {"gemma-4-12b", "qwen3.6-27B-code", "qwen3.6-35B-udq4", "ornith-1.0-35b"}
|
||||
for section_path, section_dict in [
|
||||
("model", model), ("compression", comp),
|
||||
("auxiliary.vision", aux.get("vision", {})),
|
||||
("auxiliary.web_extract", aux.get("web_extract", {})),
|
||||
("auxiliary.compression", aux.get("compression", {})),
|
||||
("delegation", deleg),
|
||||
]:
|
||||
m = section_dict.get("model", "")
|
||||
if m in raw_names:
|
||||
# --- Retired/raw model names (Rule 7/8 spirit) ---
|
||||
# The audit's job is to catch configs that are BROKEN, not to enforce a style preference.
|
||||
# NON-RESOLVING names (removed 2026-09-12, verified 400/403 via live LiteLLM) must hard-FAIL:
|
||||
# gpu-light -> gpu-vision ; gemma-4-12b -> gpu-vision
|
||||
# crew-auto -> syslog-auto (its 64K cap is retired; no cap in force) ; ornith-1.0-35b -> strix-moe
|
||||
# RESOLVING names (verified 200) are discouraged but working, so they only WARN:
|
||||
# qwen3.6-27B-code -> gpu-dense ; qwen3.6-35B-udq4 -> strix-moe
|
||||
# Failing a working alias would reject valid configs - the exact defect this change fixes.
|
||||
non_resolving = {
|
||||
"gpu-light": "gpu-vision",
|
||||
"gemma-4-12b": "gpu-vision",
|
||||
"crew-auto": "syslog-auto (its 64K cap is retired; no cap in force)",
|
||||
"ornith-1.0-35b": "strix-moe",
|
||||
"qwen3.6-27B-code": "gpu-dense",
|
||||
"qwen3.6-35B-udq4": "strix-moe",
|
||||
}
|
||||
raw_but_live = {}
|
||||
for field_path, value in _iter_model_values(cfg):
|
||||
if value in non_resolving:
|
||||
check(
|
||||
False,
|
||||
"Rule 7/8",
|
||||
f"{field_path} = {value!r} is retired and no longer resolves (2026-09-12) — use {non_resolving[value]}",
|
||||
)
|
||||
elif value in raw_but_live:
|
||||
warn(
|
||||
"Rule 7/8",
|
||||
f"{section_path}.model = {m!r} — raw model name, use stable alias instead "
|
||||
f"(gpu-light, gpu-dense, strix-moe, syslog-auto)",
|
||||
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
||||
)
|
||||
|
||||
# --- MCP Server Checks (Rule 15) ---
|
||||
# Valid MCP server endpoints
|
||||
VALID_MCP_ENDPOINTS = {
|
||||
'ra-h-os': 'http://192.168.68.65:3100/mcp',
|
||||
'litellm': 'https://litellm.sysloggh.net/mcp',
|
||||
}
|
||||
|
||||
# Check MCP servers if they exist
|
||||
mcp_servers = cfg.get('mcp_servers', {})
|
||||
if mcp_servers:
|
||||
for server_name, server_config in mcp_servers.items():
|
||||
url = server_config.get('url', '')
|
||||
|
||||
# Check endpoint validity
|
||||
if server_name in VALID_MCP_ENDPOINTS:
|
||||
expected = VALID_MCP_ENDPOINTS[server_name]
|
||||
if url == expected:
|
||||
check(True, 'Rule 15', f'MCP server "{server_name}" URL is correct: {url}')
|
||||
else:
|
||||
check(False, 'Rule 15', f'MCP server "{server_name}" URL is incorrect: {url} (expected: {expected})')
|
||||
else:
|
||||
warn('Rule 15', f'MCP server "{server_name}" URL may need validation (not in known list): {url}')
|
||||
|
||||
# Check for proper authentication
|
||||
headers = server_config.get('headers', {})
|
||||
has_auth = False
|
||||
for key, value in headers.items():
|
||||
if 'key' in key.lower() or 'auth' in key.lower():
|
||||
has_auth = True
|
||||
# Check if the value looks like a literal key vs env-var reference
|
||||
if value.startswith('Bearer ') and value[7:].startswith('sk-'):
|
||||
check(True, 'Rule 15', f'MCP server "{server_name}" has valid auth header: {key}')
|
||||
else:
|
||||
warn('Rule 15', f'MCP server "{server_name}" header may use env-var instead of literal key: {key} = {value}')
|
||||
break
|
||||
if not has_auth:
|
||||
warn('Rule 15', f'MCP server "{server_name}" has no authentication header')
|
||||
|
||||
# --- Report ---
|
||||
print(f"{'=' * 60}")
|
||||
print(f"Hermes Config Audit: {path}")
|
||||
|
||||
@@ -45,9 +45,10 @@ description: >
|
||||
## Status
|
||||
|
||||
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
|
||||
for any Hermes agent that connects to Zulip (Mumuni, Tanko, Koby, Koonimo). The pi
|
||||
for Hermes agents that connect to Zulip (Mumuni, Koby, Koonimo). The pi
|
||||
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
|
||||
plugin system, which is unaffected.
|
||||
plugin system, which is unaffected. **Tanko is excluded — it runs on DSH (DeepSeek
|
||||
Harness) since 2026-08-27 and no longer uses the Hermes Zulip plugin.**
|
||||
|
||||
## Parameters
|
||||
|
||||
|
||||
+133
-9
@@ -1,5 +1,5 @@
|
||||
registry_version: 0.1.0
|
||||
last_updated: '2026-07-13T00:00:00Z'
|
||||
last_updated: '2026-09-11T00:00:00Z'
|
||||
updated_by: mumuni
|
||||
categories:
|
||||
- compliance
|
||||
@@ -38,6 +38,7 @@ index:
|
||||
by_category:
|
||||
compliance:
|
||||
- hermes-key-enforcement
|
||||
- litellm-api-keys
|
||||
- hermes-config-template
|
||||
- hermes-agent-baseline
|
||||
monitoring:
|
||||
@@ -101,6 +102,7 @@ index:
|
||||
proxmox:
|
||||
- proxmox-monitor
|
||||
litellm:
|
||||
- litellm-api-keys
|
||||
- litellm-health
|
||||
- litellm-self-heal
|
||||
memory:
|
||||
@@ -628,18 +630,18 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 3.0.0
|
||||
version: 3.4.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/15 * * * *'
|
||||
description: "Every 15 minutes \u2014 monitors all Zulip-connected agents"
|
||||
description: "Every 15 minutes \u2014 monitors the Zulip-connected agents under this host's control (pi, DSH, Agent Zero)"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 120
|
||||
requires:
|
||||
- Zulip API key for abiba-bot@chat.sysloggh.net
|
||||
- SSH access to all Hermes agents
|
||||
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
|
||||
verification:
|
||||
postconditions:
|
||||
- check: bot registration active
|
||||
@@ -693,8 +695,8 @@ contracts:
|
||||
version: 1.0.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/10 * * * *'
|
||||
description: "Every 10 minutes \u2014 LiteLLM proxy health"
|
||||
cadence: '5 3,7,11,15,19,23 * * *'
|
||||
description: "4-hourly staggered dispatch via fm-send (run contract litellm-health)"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
@@ -750,7 +752,7 @@ contracts:
|
||||
sensitivity: critical
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 1.0.0
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: event_driven
|
||||
description: Triggered by relay message from litellm-health or infrastructure-monitoring
|
||||
@@ -1155,7 +1157,7 @@ contracts:
|
||||
type: scheduled
|
||||
cadence: 0 3 * * *
|
||||
description: Daily at 3am ET
|
||||
cron_job_id: null
|
||||
cron_job_id: b59f3cc21f4c # provisioned on kagentz 2026-09-08 (okyeame-memory-audit, glm-5.3-flash)
|
||||
execution:
|
||||
agent: mumuni
|
||||
timeout: 600
|
||||
@@ -1365,7 +1367,7 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: ops
|
||||
version: 1.0.0
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: 0 2 * * 0
|
||||
@@ -1867,3 +1869,125 @@ contracts:
|
||||
last_run: null
|
||||
last_status: null
|
||||
drift_alerts: []
|
||||
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||
- name: litellm-api-keys
|
||||
file: litellm-api-keys.prose.md
|
||||
kind: function
|
||||
category: compliance
|
||||
sensitivity: critical
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: on_demand
|
||||
cadence: null
|
||||
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 120
|
||||
requires: []
|
||||
protocol:
|
||||
- Load contract from prose-contracts/main
|
||||
- Retrieve master key from Infisical (project=infrastructure env=production)
|
||||
- Read live key-scoped model roster from CT 116 /v1/models
|
||||
- Create/rotate/verify the requested agent key with an EXPLICIT models list
|
||||
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
|
||||
verification:
|
||||
postconditions:
|
||||
- check: standard agent key is local-only
|
||||
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
|
||||
expect: 0 cloud models
|
||||
- check: key exists with correct alias
|
||||
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
|
||||
expect: 200 with matching alias
|
||||
artifact: key creation/rotation report
|
||||
receipt:
|
||||
format: json
|
||||
storage: ~/.hermes/runs/litellm-api-keys/
|
||||
graph_node: true
|
||||
escalation:
|
||||
info:
|
||||
action: log_to_receipt
|
||||
notify: []
|
||||
warning:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
critical:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
- ops
|
||||
|
||||
koby_report_only: true
|
||||
koby_host: "CT 111 (tdunna)"
|
||||
koby_ip: ".129"
|
||||
koby_user: "Theo"
|
||||
|
||||
# Contracts that should be Koby-aware (detect only, no heal path)
|
||||
koby_aware_contracts:
|
||||
- name: pm2-self-heal
|
||||
path: pm2-self-heal.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby PM2 processes reported to Zulip, never auto-restarted on .129"
|
||||
|
||||
- name: zulip-health
|
||||
path: zulip-health.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby Zulip bridge issues reported to Zulip, never repaired on .129"
|
||||
|
||||
- name: hermes-zulip-restore
|
||||
path: hermes-zulip-restore.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby Zulip restoration skipped, only diagnostic alerts"
|
||||
|
||||
- name: abiba-zulip-restore
|
||||
path: abiba-zulip-restore.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Abiba-Zulip restoration not applicable to Koby"
|
||||
|
||||
- name: litellm-self-heal
|
||||
path: litellm-self-heal.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby LiteLLM issues reported, never fixed on .129"
|
||||
|
||||
- name: disk-gc-threat-response
|
||||
path: disk-gc-threat-response.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby disk GC threats reported, never executed on .129"
|
||||
|
||||
- name: memory-fixer
|
||||
path: memory-fixer.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby memory issues reported, never fixed on .129"
|
||||
|
||||
- name: memory-audit-maintenance
|
||||
path: memory-audit-maintenance.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby memory audits reported, never performed on .129"
|
||||
|
||||
- name: gpu-self-heal
|
||||
path: gpu-self-heal.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby GPU issues reported, never fixed on .129"
|
||||
|
||||
- name: gpu-monitor
|
||||
path: gpu-monitor.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby GPU monitoring reports only, never repairs on .129"
|
||||
|
||||
- name: agent-health-check
|
||||
path: agent-health-check.prose.md
|
||||
koby_action: skip_heal
|
||||
koby_note: "Koby agent health checks reported, never repairs on .129"
|
||||
|
||||
# Scripts that should skip Koby
|
||||
koby_aware_scripts:
|
||||
- name: agent-health-check.py
|
||||
path: scripts/agent-health-check.py
|
||||
koby_action: skip_heal
|
||||
koby_note: "Script should only run diagnostics on Koby, not repairs"
|
||||
|
||||
@@ -2,6 +2,10 @@
|
||||
|
||||
Generated: 2026-07-13 20:59:18 ET
|
||||
|
||||
> **Point-in-time snapshot.** Schedules and cadences are authoritative in
|
||||
> `contract-registry.yaml`; any schedule quoted below may be stale. Do not use
|
||||
> this file as the source of truth for a contract's trigger.
|
||||
|
||||
---
|
||||
|
||||
## hermes-key-enforcement
|
||||
@@ -442,7 +446,7 @@ IMPORTANT: If the contract file does not exist in prose-contracts/main, report f
|
||||
|
||||
## litellm-health
|
||||
|
||||
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** */10 * * * *
|
||||
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** see contract-registry.yaml (authoritative)
|
||||
|
||||
```
|
||||
Contract Enforcement: litellm-health
|
||||
@@ -450,7 +454,7 @@ Contract Enforcement: litellm-health
|
||||
Category: monitoring
|
||||
Domain: litellm
|
||||
Owner: abiba
|
||||
Schedule: Every 10 minutes — LiteLLM proxy health
|
||||
Schedule: see contract-registry.yaml (authoritative)
|
||||
|
||||
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
|
||||
|
||||
|
||||
@@ -1,9 +1,14 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
# ⛔ The guest/host-keyed report-only gate the GC executor MUST honour lives in the body
|
||||
# "Hard gate" YAML block below — that block is authoritative and is the only copy.
|
||||
kind: responsibility
|
||||
name: disk-gc-threat-response
|
||||
description: >
|
||||
Recurring disk health scan, garbage collection, and threat response
|
||||
across 15 Proxmox CTs + 3 GPU bare-metal hosts. Triggered by incident
|
||||
across 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts
|
||||
(fleet verified against `pvesh get /cluster/resources` 2026-09-12). Triggered by incident
|
||||
2026-07-04 where CT 105 (kagentz) hit 87% disk (49G/59G) from
|
||||
Docker image bloat — 5 dangling images, 15 build cache layers.
|
||||
Recovered 35.67GB. Second incident 2026-07-09: amdpve (.15) Docker
|
||||
@@ -11,6 +16,7 @@ description: >
|
||||
id: 067NV8KJ03ZG71S44N41F31022
|
||||
version: 1.0.0
|
||||
---
|
||||
---
|
||||
|
||||
# Disk GC & Threat Response
|
||||
|
||||
@@ -22,7 +28,7 @@ logged within 5 minutes of discovery.
|
||||
|
||||
## Scope
|
||||
|
||||
All 15 CTs via `pct-run` + 3 GPU bare-metal hosts via direct SSH.
|
||||
All 20 Proxmox guests (17 LXC via `pct-run` + 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts via direct SSH.
|
||||
Docker hosts get special attention:
|
||||
|
||||
| Host | CT | Disk Risk | GC Strategy |
|
||||
@@ -37,7 +43,7 @@ Docker hosts get special attention:
|
||||
|
||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
|
||||
|
||||
## Threat Levels
|
||||
## Threat Levels (GUEST filesystems)
|
||||
|
||||
| Level | Threshold | Response | Escalation |
|
||||
|-------|-----------|----------|------------|
|
||||
@@ -47,6 +53,39 @@ Docker hosts get special attention:
|
||||
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
|
||||
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
|
||||
|
||||
## Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)
|
||||
|
||||
Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are **report-only** — no automatic deletion of media or datastore content ever.
|
||||
|
||||
| Level | Threshold | Response | Escalation |
|
||||
|-------|-----------|----------|------------|
|
||||
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
||||
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
||||
| **HOST-RED** | 95% | Name the volume + % + absolute free space; **media volumes: "capacity decision — owner to decide"; pbs-datastore/host-root: "immediate owner attention"** | Zulip DM + channel alert (state-change only) |
|
||||
|
||||
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
||||
|
||||
**State lives in a small JSON state file:** the scanner resolves it to an absolute path from the script's location — `$(dirname "$0")/state/host-disk-bands.json` (i.e., the `state/` directory next to the `scripts/` directory in the prose-contracts repo). The state file is gitignored runtime state — the scanner creates it on first run. Keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
|
||||
|
||||
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
|
||||
|
||||
```
|
||||
storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
|
||||
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
|
||||
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
|
||||
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
|
||||
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN
|
||||
```
|
||||
|
||||
**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.
|
||||
|
||||
**Action classes by volume type:**
|
||||
- **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
|
||||
- **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
|
||||
- **pbs-datastore** (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.
|
||||
|
||||
**Justification (measured 2026-09-15, firstmate):** storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
|
||||
@@ -78,6 +117,32 @@ and escalation trail.
|
||||
- May also be invoked manually: `prose run disk-gc-threat-response`
|
||||
- Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
|
||||
|
||||
## Scanner: scripts/disk-gc-scan.py
|
||||
|
||||
The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability
|
||||
verdicts deterministic:
|
||||
|
||||
1. **Retry on failure:** Each probe retries once before declaring a guest unreachable.
|
||||
2. **Named probe target:** Every rendered line names the guest, CT id, node, and
|
||||
access method actually used.
|
||||
3. **Failure kind printed:** An unreachable guest is reported with its failure kind
|
||||
(timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare
|
||||
"unreachable" verdict.
|
||||
4. **Per-guest access method:** The correct access path is selected from a per-guest
|
||||
map so the wrong path cannot be picked by an executor improvising:
|
||||
- CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees
|
||||
loop0/59G instead of the real 99G filesystem)
|
||||
- CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM)
|
||||
- All other CTs = `pct-run <ct_id>` (which uses `pct exec` via SSH to the node)
|
||||
5. **Every figure traces to a named probe:** The scan output prints the exact command
|
||||
that produced each disk figure, so two different guests can never render
|
||||
identical numbers without the probe commands proving it.
|
||||
|
||||
Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output).
|
||||
|
||||
The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate
|
||||
from the `report_only_guests` YAML block above.
|
||||
|
||||
## Shape
|
||||
|
||||
- `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
|
||||
@@ -96,37 +161,88 @@ and escalation trail.
|
||||
|
||||
## Execution
|
||||
|
||||
### Host filesystems: report-only, NEVER auto-delete
|
||||
|
||||
Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.
|
||||
|
||||
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
|
||||
|
||||
**CT 111 / hostname `tdunna` / 192.168.68.129 is DETECT-AND-REPORT-ONLY.** It belongs to Theo.
|
||||
The captain ruled 2026-08-17 and re-confirmed 2026-09-10 that Theo handles CT 111 himself.
|
||||
At **every** threat level — AMBER, RED, or CRITICAL — the executor must:
|
||||
|
||||
- push the threat row and alert the owner, and
|
||||
- **never** call `gc-executor`, and **never** run any GC command against that guest: no
|
||||
`apt-get clean/autoremove`, no `journalctl --vacuum-*`, no `find /var/log -delete`, no
|
||||
`/tmp`/`/var/tmp` deletion, no snap removal, no `docker system prune`.
|
||||
|
||||
This gate is keyed on **guest id / hostname / IP**, not on an agent name. The frontmatter
|
||||
`report_only_agents` marker (e.g. `koby`) names an AGENT while the scan unit is a GUEST, so an
|
||||
agent-name marker can silently miss the guest it lives on — it must never be the only gate.
|
||||
|
||||
**The authoritative machine-readable exclusion list is the YAML block below.** The executor
|
||||
reads it at run time; `scripts/disk-gc-plan.py` turns a fleet scan into the action plan using it.
|
||||
Extend the list here, never by hand-maintaining a second copy. The Execution loop below MUST
|
||||
call that planner and MUST NOT reimplement the gate.
|
||||
|
||||
```yaml
|
||||
# disk-gc report-only guests — authoritative. Keyed on guest/host, not agent.
|
||||
report_only_guests:
|
||||
- guest: 111
|
||||
hostname: tdunna
|
||||
ip: 192.168.68.129
|
||||
node: storepve
|
||||
reason: "Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10"
|
||||
```
|
||||
|
||||
### Loop
|
||||
|
||||
```prose
|
||||
let fleet = call disk-scanner
|
||||
scope: all
|
||||
|
||||
let threats = []
|
||||
for ct in fleet:
|
||||
if ct.usage_pct >= 95:
|
||||
push threats { ct: ct.id, level: "CRITICAL", pct: ct.usage_pct }
|
||||
else if ct.usage_pct >= 85:
|
||||
push threats { ct: ct.id, level: "RED", pct: ct.usage_pct }
|
||||
else if ct.usage_pct >= 75:
|
||||
push threats { ct: ct.id, level: "AMBER", pct: ct.usage_pct }
|
||||
-- The report-only gate is IMPLEMENTED IN scripts/disk-gc-plan.py and MUST NOT be
|
||||
-- reimplemented here. That planner reads the contract's `report_only_guests` YAML block
|
||||
-- and matches on guest id OR hostname OR IP, so the tested gate is the executed gate.
|
||||
let plan = call disk-gc-plan
|
||||
fleet: fleet
|
||||
|
||||
-- sort by severity descending
|
||||
sort threats by pct desc
|
||||
for row in plan:
|
||||
if row.action == "report-only":
|
||||
-- Excluded guest: alert only. No gc-executor call is constructed for it, at any level.
|
||||
call alerter
|
||||
threat: row
|
||||
result: { action: "report-only", reason: row.reason }
|
||||
else:
|
||||
let result = call gc-executor
|
||||
ct: row.target
|
||||
level: row.level
|
||||
strategy: lookup-gc-strategy(row.target)
|
||||
|
||||
for threat in threats:
|
||||
let result = call gc-executor
|
||||
ct: threat.ct
|
||||
level: threat.level
|
||||
strategy: lookup-gc-strategy(threat.ct)
|
||||
|
||||
call alerter
|
||||
threat: threat
|
||||
result: result
|
||||
call alerter
|
||||
threat: row
|
||||
result: result
|
||||
|
||||
call summary-reporter
|
||||
fleet: fleet
|
||||
threats: threats
|
||||
plan: plan
|
||||
```
|
||||
|
||||
## GC SCHEDULE (PBS datastore only)
|
||||
|
||||
The PBS GC schedule is defined in ONE authoritative place: `/etc/cron.d/pbs-gc` on storepve.
|
||||
The schedule is `0 20 * * *` (20:00 LOCAL = 00:00 UTC, since host timezone is America/New_York).
|
||||
This applies **only** to the PBS datastore (`/tank/pbs-backup`), NOT to media volumes.
|
||||
Media volumes (/media/*) are report-only at all threat levels.
|
||||
|
||||
The cron runs `/usr/local/bin/pbs-gc.sh` which executes:
|
||||
```bash
|
||||
proxmox-backup-manager garbage-collection start storepve-datastore
|
||||
```
|
||||
This is NOT a `prune` operation; it is a GC pass that reclaims unreferenced chunks.
|
||||
There is no `--keep-daily` flag; retention is governed by jobs.cfg (keep-daily=35).
|
||||
The GC does not touch media volumes or any other filesystem.
|
||||
|
||||
## GC Strategies by Host Type
|
||||
|
||||
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
||||
@@ -179,6 +295,12 @@ done
|
||||
|
||||
## Alert Templates
|
||||
|
||||
### HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)
|
||||
```
|
||||
⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
|
||||
Action: {volume_type-specific action}
|
||||
```
|
||||
|
||||
### AMBER (75-84%)
|
||||
```
|
||||
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
||||
@@ -236,7 +358,9 @@ dangling images and orphaned build cache. No automated GC was in place.
|
||||
## Incident Log: 2026-07-09 — amdpve docker bloat
|
||||
|
||||
### Discovery
|
||||
Scheduled fleet disk scan via `pct-run` across all 15 CTs + 3 GPU bare-metal hosts.
|
||||
Scheduled fleet disk scan across all 20 Proxmox guests (17 LXC via `pct-run`, 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts.
|
||||
> **Report-only gate applies to this scan:** CT 111 (`tdunna`, 192.168.68.129) is alerted but never
|
||||
> garbage-collected at any level.
|
||||
amdpve (.15) flagged at 78% (AMBER threshold: 75%).
|
||||
|
||||
### Diagnosis
|
||||
@@ -269,26 +393,35 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
- Contract now scans GPU bare-metal hosts alongside CTs
|
||||
- Access via `pct-run` script for all CTs (no hardcoded IPs)
|
||||
|
||||
## Access Matrix (documented 2026-07-09)
|
||||
## Access Matrix (verified against `pvesh get /cluster/resources` 2026-09-12)
|
||||
|
||||
### CT Access (via pct-run)
|
||||
| CT | Name | Node | Status |
|
||||
|----|------|------|--------|
|
||||
| 100 | abiba | hwepve | local |
|
||||
| 102 | adguard | minipve | ✅ reachable |
|
||||
| 104 | authentik | minipve | ✅ reachable |
|
||||
| 105 | kagentz | hwepve | ✅ reachable |
|
||||
| 106 | ra-h-os | storepve | ✅ reachable |
|
||||
| 107 | pbs | storepve | ✅ reachable |
|
||||
| 108 | media | storepve | ✅ reachable |
|
||||
| 110 | gitea | minipve | ✅ reachable |
|
||||
| 111 | tdunna | amdpve | ✅ reachable |
|
||||
| 112 | tanko | amdpve | ✅ reachable |
|
||||
| 113 | baggy | amdpve | ✅ reachable |
|
||||
| 114 | mumuni | hwepve | ✅ reachable |
|
||||
| 115 | scottdenya | amdpve | ✅ reachable |
|
||||
| 116 | syslog-api | minipve | ✅ reachable |
|
||||
| 117 | zulip | storepve | ✅ reachable |
|
||||
### Guest Access (via `pct-run` — CT id only, node resolved by `scripts/pct-run.sh`)
|
||||
| Guest | Name | Node | Type | Status |
|
||||
|------|------|------|------|--------|
|
||||
| 100 | abiba | minipve | lxc | ✅ reachable (probed via pct-run like any other guest; no local shortcut) |
|
||||
| 102 | adguard | minipve | lxc | ✅ reachable |
|
||||
| 104 | authentik | minipve | lxc | ✅ reachable |
|
||||
| 105 | kagentz | **amdpve** | lxc | ✅ reachable (was documented as minipve — corrected) |
|
||||
| 106 | ra-h-os | storepve | lxc | ✅ reachable |
|
||||
| 107 | pbs | storepve | lxc | ✅ reachable |
|
||||
| 108 | media | storepve | lxc | ✅ reachable |
|
||||
| 110 | gitea | minipve | lxc | ✅ reachable |
|
||||
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
|
||||
| 112 | tanko | amdpve | lxc | ✅ reachable |
|
||||
| 113 | baggy | amdpve | lxc | ✅ reachable |
|
||||
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
|
||||
| 116 | syslog-api | minipve | lxc | ✅ reachable |
|
||||
| 117 | zulip | storepve | lxc | ✅ reachable |
|
||||
| 118 | jdownloader | storepve | lxc | ✅ reachable |
|
||||
| 119 | infisical-vault | minipve | lxc | ✅ reachable |
|
||||
| 120 | adguard2 | amdpve | lxc | ✅ reachable |
|
||||
|
||||
### QEMU VMs (via direct SSH)
|
||||
| VM | Name | Node | IP | Status |
|
||||
|----|------|------|-----|--------|
|
||||
| 101 | llm-gpu (workload now bare metal .8) | acerpve | — | ✅ reachable |
|
||||
| 103 | ocu-llm (workload now bare metal .110) | ocupve | — | ✅ reachable |
|
||||
| 109 | docker-vm | storepve | 192.168.68.7 | ✅ reachable |
|
||||
|
||||
### GPU Bare Metal (via direct SSH)
|
||||
| Host | IP | GPU | Status |
|
||||
@@ -297,9 +430,15 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
| ocu-llm | 192.168.68.110 | RTX 5070 | ✅ reachable |
|
||||
| amdpve | 192.168.68.15 | Strix Halo | ✅ reachable |
|
||||
|
||||
### KVM VM (via direct SSH)
|
||||
| Host | IP | Role | Status |
|
||||
|------|-----|------|--------|
|
||||
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
|
||||
|
||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
||||
> **Fleet count:** 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts. Corrected
|
||||
> 2026-09-12: CT 105 → amdpve, CT 111 → storepve, and guests 118/119/120 were missing.
|
||||
>
|
||||
> **CT 100 probe gap (folded in):** CT 100 previously reported "unreachable (not reported)" every
|
||||
> run. Root cause is the same stale access layer: `pct-run` resolves the guest's node from its map,
|
||||
> and the map/contract must reflect `pvesh /cluster/resources`. Verified working from inside CT 100:
|
||||
> `scripts/pct-run.sh 100 "df -P / | tail -1"` → `23% /`. Probe CT 100 through `pct-run` like any
|
||||
> other guest — never through a local-only path, since the scanner itself runs inside CT 100 and a
|
||||
> container has no `pct` binary.
|
||||
>
|
||||
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
|
||||
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.
|
||||
|
||||
+36
-2
@@ -13,7 +13,7 @@ the description:
|
||||
|
||||
1. **What system does this contract touch?** Name the hosts, CTs, containers,
|
||||
and services explicitly. "The inference fleet" is vague. "GPU .8 (RTX 3090,
|
||||
qwen), .110 (RTX 5070, gemma), .15 (Strix Halo, strix-moe), and LiteLLM on CT
|
||||
qwen), .110 (RTX 5070, gpu-vision), .15 (Strix Halo, strix-moe), and LiteLLM on CT
|
||||
116" is specific.
|
||||
|
||||
2. **Who runs this contract, and when?** State the agent, the trigger (cron,
|
||||
@@ -185,7 +185,7 @@ what, and why should I care?
|
||||
|
||||
```
|
||||
❌ "Monitors infrastructure health"
|
||||
✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker
|
||||
✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker
|
||||
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
|
||||
```
|
||||
|
||||
@@ -211,6 +211,40 @@ Errors tell the operator what went wrong and what to do about it. Be specific:
|
||||
Check router /health/unified at http://192.168.68.116/health/unified instead."
|
||||
```
|
||||
|
||||
### Report provenance
|
||||
|
||||
Every report a contract produces must lead with the **absolute path the probe
|
||||
executed from** — `pwd -P`, or the running script's absolute path. A live-state
|
||||
report without provenance is unactionable: a report from a stale copy (a worktree
|
||||
clone, a retired cron entry, a diverged consumer) looks identical to a live
|
||||
fault, and the team burns rounds repairing healthy infrastructure. This is not
|
||||
optional. The 2026-09-09 probe-drift rounds cost three false `DEGRADED` reports
|
||||
because a stale consumer probed the wrong port and nothing in the report said
|
||||
where it ran.
|
||||
|
||||
Pair it with the **scoped any-HTTP-response liveness rule**: for unauthenticated
|
||||
or auth-gated endpoints — where any HTTP answer proves a listener is up (the
|
||||
PVE API's `401`, LiteLLM health's `301` redirect) — a probe is ALIVE on ANY HTTP
|
||||
status, including `301` redirects and `401`/`403` auth challenges. **DOWN =
|
||||
connection refused (`000`) or timeout only.**
|
||||
|
||||
Probes whose success condition is specifically a bare `200` are NOT covered by
|
||||
the any-HTTP rule. On those — authenticated probes such as the Zulip message
|
||||
POST and the router `/health` — an unexpected status (`401`/`403` from a bad or
|
||||
missing credential, `5xx`, or anything other than the expected `200`) is an
|
||||
**ALERT**, not "alive".
|
||||
|
||||
```markdown
|
||||
**Report format**: Begin every report with the absolute execution path
|
||||
(`pwd -P` / script path). On auth-gated endpoints, alive = ANY HTTP status and
|
||||
DOWN = `000`/timeout only; on probes whose expected result is `200`, any other
|
||||
status is an alert.
|
||||
```
|
||||
|
||||
The lint pipeline enforces the provenance clause: any contract with a
|
||||
`**Report format**` line must state an absolute path (`pwd -P`, `absolute path`,
|
||||
or `executed from`).
|
||||
|
||||
### Comments
|
||||
|
||||
Comments in contracts explain WHY, not WHAT. The execution steps say what to
|
||||
|
||||
@@ -0,0 +1,282 @@
|
||||
# Probe-drift round 2 — per-leg before/after evidence
|
||||
|
||||
**Date:** 2026-09-10
|
||||
**Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`
|
||||
**Branch:** `fm/probe-drift-round2-20260909`
|
||||
|
||||
Every command below was run from the absolute path above; output is pasted
|
||||
verbatim. This is the evidence trail for the four scoped corrections; it is not
|
||||
a contract (never `prose run` it).
|
||||
|
||||
---
|
||||
|
||||
## Leg 1 — agent-health-check (item 1)
|
||||
|
||||
**Before** — from `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`,
|
||||
`python3 scripts/agent-health-check.py --no-deploy` (v2, base of this branch):
|
||||
|
||||
```
|
||||
🏥 Agent Health Check v2 — 2026-09-10 01:16 UTC
|
||||
|
||||
🔑 LiteLLM Keys:
|
||||
✅ tanko: key valid → syslog-auto
|
||||
✅ abiba: key valid → syslog-auto
|
||||
✅ koby: key valid → syslog-auto
|
||||
✅ koonimo: key valid → syslog-auto
|
||||
|
||||
🎮 GPU Port Health:
|
||||
✅ gpu-rtx3090 (.8): healthy (pid=472206)
|
||||
✅ gpu-rtx5070 (.110): healthy (pid=207601)
|
||||
✅ gpu-strixhalo (.15): healthy (pid=4098872)
|
||||
|
||||
🤖 Agent Gateways:
|
||||
✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
|
||||
⚠️ abiba: gw=no-state-file zulip=? streaming=no errors_10m=0 pid=?
|
||||
✅ koby: gw=running zulip=connected streaming=no errors_10m=0 pid=360900
|
||||
✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125
|
||||
|
||||
🖥️ CT Liveness:
|
||||
✅ tanko (CT 112 on amdpve): running
|
||||
✅ abiba (CT 100 on minipve): running
|
||||
❌ koby (CT 111 on amdpve): PVE UNREACHABLE
|
||||
✅ koonimo (CT 113 on amdpve): running
|
||||
|
||||
📝 Config Integrity:
|
||||
⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
|
||||
✅ abiba: config.yaml valid YAML
|
||||
✅ koby: config.yaml valid YAML
|
||||
✅ koonimo: config.yaml valid YAML
|
||||
|
||||
🔌 Wrapper/CLI Integrity:
|
||||
⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
|
||||
⚠️ abiba: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
|
||||
❌ abiba: hermes-real NOT FOUND (wrapper broken)
|
||||
⚠️ abiba: .env may be missing LITELLM_API_KEY entry
|
||||
⚠️ koby: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
|
||||
❌ koby: hermes-real NOT FOUND (wrapper broken)
|
||||
✅ koby: wrapper + .env key present
|
||||
⚠️ koonimo: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
|
||||
✅ koonimo: wrapper + .env key present
|
||||
|
||||
🔐 Vault Secrets:
|
||||
✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
|
||||
✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
|
||||
✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ
|
||||
|
||||
❌ 6 FAILURE(S): ct-unreachable:koby:192.168.68.15 | wrapper-infisical-path:abiba | wrapper-no-hermes-real:abiba | wrapper-infisical-path:koby | wrapper-no-hermes-real:koby | wrapper-infisical-path:koonimo
|
||||
```
|
||||
|
||||
Root causes (all stale expectations; no live fault):
|
||||
|
||||
| Failure | Why it was stale |
|
||||
|---|---|
|
||||
| `ct-unreachable:koby:192.168.68.15` | CT 111 (tdunna/koby) runs on **storepve (.6)**, not amdpve (.15). |
|
||||
| `wrapper-*:abiba` | Abiba is pi-only since the harness purge. `/root/.local/bin/hermes` is a dangling symlink; no `hermes-real`, no `~/.hermes/.env`. |
|
||||
| `wrapper-*:koby` | Koby is **report-only** (captain ruling 2026-08-17, Rule 17): detect and report, never repair — its legs must not count as fleet failures. Koby's wrapper is also the genuine no-infisical case: it sources `~/.hermes/.env` rather than `/usr/bin/infisical`, which the check now accepts. |
|
||||
| `wrapper-infisical-path:koonimo` | Koonimo's wrapper **does** reference `/usr/bin/infisical` — but past the old check's `head -20` window, so the check looked for the path in the wrong slice and false-failed. The fix that mattered was reading the full wrapper body (and then verifying any absolute infisical path it finds actually exists). |
|
||||
|
||||
**After** — same absolute path, `python3 scripts/agent-health-check.py --no-deploy` (v4):
|
||||
|
||||
```
|
||||
$ pwd -P
|
||||
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
|
||||
$ python3 scripts/agent-health-check.py --no-deploy
|
||||
🏥 Agent Health Check v4 — 2026-09-10 01:22 UTC
|
||||
📍 executed from: script=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts/scripts/agent-health-check.py cwd=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
|
||||
|
||||
🔑 LiteLLM Keys:
|
||||
✅ tanko: key valid → syslog-auto
|
||||
✅ abiba: key valid → syslog-auto
|
||||
✅ koby: key valid → syslog-auto
|
||||
✅ koonimo: key valid → syslog-auto
|
||||
|
||||
🎮 GPU Port Health:
|
||||
✅ gpu-rtx3090 (.8): healthy (pid=472206)
|
||||
✅ gpu-rtx5070 (.110): healthy (pid=207601)
|
||||
✅ gpu-strixhalo (.15): healthy (pid=4098872)
|
||||
|
||||
🤖 Agent Gateways:
|
||||
✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
|
||||
✅ abiba: pi-only runtime — no Hermes gateway since the harness purge (CT 100, SSH OK)
|
||||
🔍 koby: REPORT-ONLY mode (diagnostic only, no repairs on .129)
|
||||
✅ koby: gateway running (pid=360900, report-only mode)
|
||||
✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125
|
||||
|
||||
🖥️ CT Liveness:
|
||||
✅ tanko (CT 112 on amdpve): running
|
||||
✅ abiba (CT 100 on minipve): running
|
||||
✅ koby (CT 111 on storepve): running
|
||||
✅ koonimo (CT 113 on amdpve): running
|
||||
|
||||
📝 Config Integrity:
|
||||
⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
|
||||
⏭️ abiba: pi-only runtime — no Hermes config.yaml since the harness purge
|
||||
✅ koby: config.yaml valid YAML
|
||||
✅ koonimo: config.yaml valid YAML
|
||||
|
||||
🔌 Wrapper/CLI Integrity:
|
||||
⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
|
||||
⏭️ abiba: pi-only runtime — no hermes CLI wrapper since the harness purge
|
||||
ℹ️ koby: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK
|
||||
❌ koby: hermes-real NOT FOUND (wrapper broken)
|
||||
🔍 report-only (koby): wrapper-no-hermes-real:koby — reported, not counted/repaired
|
||||
✅ koby: wrapper + .env key present
|
||||
✅ koonimo: wrapper infisical path OK
|
||||
✅ koonimo: wrapper + .env key present
|
||||
|
||||
🔐 Vault Secrets:
|
||||
✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
|
||||
✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
|
||||
✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ
|
||||
|
||||
✅ All checks passed
|
||||
exit=0
|
||||
```
|
||||
|
||||
**Live vantage proof** (same worktree):
|
||||
|
||||
```
|
||||
$ ssh root@192.168.68.15 "pct status 111"
|
||||
Configuration file 'nodes/amdpve/lxc/111.conf' does not exist
|
||||
$ ssh root@192.168.68.6 "pct status 111; pct list | grep '^ *111'"
|
||||
status: running
|
||||
111 running tdunna
|
||||
$ ssh root@192.168.68.129 "hostname"
|
||||
tdunna
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Leg 2 — infrastructure-monitoring PVE API (item 2)
|
||||
|
||||
**Before** — the contract's probe, aimed at the monitoring host CT 116:
|
||||
|
||||
```
|
||||
$ curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json
|
||||
000
|
||||
```
|
||||
|
||||
CT 116 runs no `pveproxy`, so it never answers on `:8006`. The probe target was
|
||||
wrong, which is what read as PVE-API `000`.
|
||||
|
||||
**After** — probing the five real cluster nodes (`:8006/api2/json/version`),
|
||||
alive under the any-HTTP-response rule (`401` = up, unauthenticated):
|
||||
|
||||
```
|
||||
$ for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
|
||||
printf '%s:8006 -> %s\n' "$node" "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
|
||||
done
|
||||
192.168.68.9:8006 -> 401
|
||||
192.168.68.5:8006 -> 401
|
||||
192.168.68.15:8006 -> 401
|
||||
192.168.68.6:8006 -> 401
|
||||
192.168.68.12:8006 -> 401
|
||||
```
|
||||
|
||||
`401` on every node = alive by design. `DOWN` is `000`/timeout only. (The
|
||||
contract's LiteLLM probe was the same class: `/litellm/health` answers `301` →
|
||||
`/litellm/health/liveliness`, so it is now specified as any-HTTP too.)
|
||||
|
||||
---
|
||||
|
||||
## Leg 3 — gpu-monitor GPU probes (item 3)
|
||||
|
||||
**Before** — the false alarm came from probing bare port 80 on GPU hosts:
|
||||
|
||||
```
|
||||
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8/health
|
||||
000
|
||||
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110/health
|
||||
000
|
||||
```
|
||||
|
||||
Nothing listens on GPU port 80, so the monitor reported
|
||||
`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` three times on 2026-09-09.
|
||||
|
||||
**After** — the real endpoints answer:
|
||||
|
||||
```
|
||||
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
|
||||
200
|
||||
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
|
||||
200
|
||||
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.15:8080/health
|
||||
200
|
||||
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
|
||||
301 # Location: http://192.168.68.116/gpu/gpu-data — the same payload
|
||||
$ curl -s -o /dev/null -w '%{http_code}' -L http://192.168.68.116/health/unified
|
||||
200
|
||||
```
|
||||
|
||||
`301` is healthy under the any-HTTP-response rule. The contract now requires GPU
|
||||
health on `:8080` (or router `/health/unified`) and forbids bare port 80 on a
|
||||
GPU host.
|
||||
|
||||
---
|
||||
|
||||
## Leg 4 — report provenance (item 4)
|
||||
|
||||
Every contract report must now lead with the absolute path it executed from.
|
||||
`docs/AUTHORING-GUIDE.md` documents the rule and `scripts/prose-lint.sh`
|
||||
enforces it:
|
||||
|
||||
```
|
||||
$ pwd -P
|
||||
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
|
||||
$ bash scripts/prose-lint.sh
|
||||
✅ Report provenance present in all report-format contracts
|
||||
...
|
||||
✅ LINT PASSED (12 warning(s))
|
||||
```
|
||||
|
||||
The health script prints `📍 executed from: script=… cwd=…` and includes
|
||||
`execution_path`/`cwd` in `--json` output.
|
||||
|
||||
---
|
||||
|
||||
## Full suite
|
||||
|
||||
```
|
||||
$ pwd -P
|
||||
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
|
||||
$ python3 -m pytest -q
|
||||
24 passed
|
||||
$ shellcheck scripts/prose-lint.sh
|
||||
(clean)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Follow-up findings (observed, intentionally NOT changed here)
|
||||
|
||||
These are adjacent stale expectations discovered while verifying the four
|
||||
scoped legs. Each touches a CRITICAL/HIGH-sensitivity artifact or an unrelated
|
||||
script, so it is recorded for the captain/verify mate rather than silently
|
||||
repaired.
|
||||
|
||||
1. **`infrastructure-control.prose.md` (CRITICAL) CT 111 node assignment.**
|
||||
Lines ~109 and ~615 place `tdunna` (CT 111, koby) on **amdpve**. Live
|
||||
verification on 2026-09-10 shows `pct status 111` = `running` on
|
||||
**storepve (.6)** and `Configuration file 'nodes/amdpve/lxc/111.conf' does
|
||||
not exist` on .15. `agent-health-check.py` now carries the live-verified
|
||||
`storepve` mapping (the script is not the topology source of truth); the
|
||||
CRITICAL contract itself needs an authorized correction.
|
||||
**✅ Resolved 2026-09-12:** the topology was corrected in its owner,
|
||||
`infrastructure-control.prose.md` (CT 111 → storepve, CT 105 → amdpve), and
|
||||
`scripts/pct-run.sh` now matches. This snapshot is left as observed; treat
|
||||
those owner documents as authoritative.
|
||||
2. **Strix Halo `:8080` firewall claim is stale.** `prose-ai-review.sh`
|
||||
ground-truth rule #4 and `gpu-monitor.prose.md` say `:8080` is firewalled to
|
||||
`.116` only and `.24` cannot probe it. Live on .15:
|
||||
`-A INPUT -s 192.168.68.24/32 -p tcp --dport 8080 -j ACCEPT`, and a probe
|
||||
from .24 returns `200`. The contract keeps routing Strix via the router
|
||||
(safe), but the claim no longer matches iptables.
|
||||
3. **`contract-registry.yaml` references `agent-health-check.prose.md`**, which
|
||||
does not exist in the repo. The registry entry (with `koby_action: skip_heal`)
|
||||
is aspirational/stale.
|
||||
4. **Pre-existing script defects, untouched:** `scripts/pm2-self-heal.sh` has a
|
||||
bash syntax error at lines 19–20 (`bash -n` fails), and `shellcheck` fails on
|
||||
five untouched scripts (`netbird-add-domain.sh`, `pct-run.sh`,
|
||||
`pm2-self-heal.sh`, `prose-ai-review.sh`, `swap-gpu-dense-model.sh`).
|
||||
`scripts/prose-lint.sh` — the one shell file touched here — is now
|
||||
shellcheck-clean.
|
||||
+79
-127
@@ -6,29 +6,26 @@ description: >
|
||||
registration, health checks, LiteLLM sync, agent key management, GPU
|
||||
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
|
||||
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
|
||||
gpu-dense, gpu-light. These never change — only the underlying model does.
|
||||
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
||||
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
|
||||
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
||||
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
|
||||
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||
For larger context needs → fall back to external providers (deepseek).
|
||||
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
|
||||
These never change — only the underlying model does.
|
||||
Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx).
|
||||
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
|
||||
UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability.
|
||||
Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context).
|
||||
Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K.
|
||||
For >128K on NVIDIA hosts → fall back to external providers (deepseek).
|
||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||
UPDATED 2026-07-27: gpu-dense swapped to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled
|
||||
(UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking
|
||||
tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning.
|
||||
VRAM ~22.4/24.6GB (91%).
|
||||
agent: abiba
|
||||
triggers:
|
||||
- on model add/remove
|
||||
- on GPU health degradation
|
||||
- on agent key rotation
|
||||
- on router restart (roster must be loaded)
|
||||
- on harness container restart (LiteLLM reloads its model list)
|
||||
---
|
||||
|
||||
## Maintains
|
||||
|
||||
- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models
|
||||
- gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116 `litellm_config.yaml`
|
||||
- router: { status: "healthy", roster_loaded: bool, models: array }
|
||||
- litellm: { status: "healthy", keys: array, models: array }
|
||||
- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB
|
||||
@@ -40,46 +37,42 @@ triggers:
|
||||
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
|
||||
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
|
||||
|
||||
## Fleet Topology (Current — July 2026)
|
||||
## Fleet Topology (Current — 2026-09-11, router decommissioned)
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────────────┐
|
||||
│ CT 116 (192.168.68.116) — Inference Harness Host │
|
||||
│ │
|
||||
│ nginx:80 (entrypoint) │
|
||||
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
|
||||
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
|
||||
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
|
||||
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
|
||||
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
|
||||
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
|
||||
│ ├─ /health/* → harness-litellm:4000 (health probes) │
|
||||
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
|
||||
│ │
|
||||
│ Containers: │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
|
||||
│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │
|
||||
│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │
|
||||
│ │ fallback │ │not in │ │ UI │ │ data src │ │
|
||||
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
|
||||
│ │ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
|
||||
│ │ :5432 │ │ :6379 │ │ :9090 │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
└────────────────────┼─────────────────────────────────────────────┘
|
||||
│
|
||||
┌───────────────┼───────────────┬──────────────────┐
|
||||
│ │ │ │
|
||||
┌────▼─────┐ ┌──────▼──────┐ ┌────▼──────┐ ┌───────▼──────┐
|
||||
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
|
||||
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
|
||||
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
|
||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
||||
└──────────┘
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ LiteLLM │ │Dashboard │ │ Grafana │ │
|
||||
│ │ :4000 │ │ :3000 │ │ :3001 │ │
|
||||
│ │ keys+sync│ │ harness │ │Prometheus│ │
|
||||
│ │ fallback │ │ UI │ │ data src │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
│ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
|
||||
│ │ :5432 │ │ :6379 │ │ :9090 │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
└───────┼──────────────────────────────────────────────────────────┘
|
||||
│
|
||||
┌─┴─────────────┬───────────────┬───────────────┐
|
||||
│ │ │ │
|
||||
┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐
|
||||
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│
|
||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||
│ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │
|
||||
└────────────┘ └────────────┘ └────────────┘ └────────────┘
|
||||
```
|
||||
|
||||
## Stable Role-Based Aliases (Introduced 2026-07-15)
|
||||
@@ -87,61 +80,27 @@ triggers:
|
||||
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names.
|
||||
When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
|
||||
|
||||
| Alias | GPU | Current Model | Will Route To |
|
||||
|-------|-----|---------------|---------------|
|
||||
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
|
||||
| `gpu-dense` | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q4_K_XL | Whatever runs on RTX 3090 |
|
||||
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
|
||||
| Alias | Serves | Where | Kind |
|
||||
|-------|--------|-------|------|
|
||||
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
|
||||
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
|
||||
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
|
||||
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
|
||||
|
||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
|
||||
but are deprecated for agent configs. Only the stable aliases survive model swaps.
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there.
|
||||
|
||||
## Current Model Assignments (2026-07-15)
|
||||
|
||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy |
|
||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
||||
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
|
||||
**No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
|
||||
and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
|
||||
is retired and returns 400 `Invalid model name`.
|
||||
|
||||
## Routing Configuration (LiteLLM — July 2026)
|
||||
|
||||
### syslog-auto Weighted Pool (Direct GPU — bypasses router)
|
||||
Model, alias, rpm/weight and fallback values are owned by CT 116
|
||||
`/opt/inference-harness/litellm_config.yaml` (see § Stable Role-Based Aliases above).
|
||||
|
||||
| Model | GPU | Weight | RPM Cap | Timeout |
|
||||
|-------|-----|--------|---------|---------|
|
||||
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
|
||||
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
|
||||
|
||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
||||
|
||||
### Direct Model Endpoints
|
||||
|
||||
| Model | RPM Cap | Notes |
|
||||
|-------|---------|-------|
|
||||
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
|
||||
| SmartCode-Fable-5-27B-UD-Q3_K_XL | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) |
|
||||
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
|
||||
|
||||
### Stable Aliases (for agent configs — never change)
|
||||
|
||||
| Alias | RPM Cap | Routes To | Purpose |
|
||||
|-------|---------|-----------|---------|
|
||||
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
|
||||
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
|
||||
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
|
||||
|
||||
### Fallback Chains
|
||||
- gemma → qwen
|
||||
- qwen → gemma
|
||||
- strix-moe → qwen → gemma
|
||||
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
|
||||
|
||||
### Why Strix Halo RPM Is Capped
|
||||
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
|
||||
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
|
||||
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
|
||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||
|
||||
## Operations
|
||||
|
||||
@@ -165,7 +124,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
||||
6. Cleanup model files (optional)
|
||||
|
||||
### heal
|
||||
1. Check all GPUs via router internal `:9000/health/unified`
|
||||
1. Check all GPUs via gpu-monitor `{{gpu_dashboard_url}}/gpu-data`
|
||||
2. Check LiteLLM health via nginx `:80/litellm/health/liveliness`
|
||||
3. Reset stuck circuit breakers if idle (Redis)
|
||||
4. Restart dead llama-server instances via SSH
|
||||
@@ -173,7 +132,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
||||
6. Verify GPU monitor server is running on pi (:9100)
|
||||
7. Verify watchdog is running on pi
|
||||
8. Restart router if roster not loaded (check logs for STARTUP ROSTER)
|
||||
9. Reload roster via `POST :9000/admin/roster/reload` if available
|
||||
9. Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing)
|
||||
|
||||
### sync-keys
|
||||
1. List all agent keys in LiteLLM DB via `GET /key/list`
|
||||
@@ -192,9 +151,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
|
||||
2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts
|
||||
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
||||
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
||||
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (strix-moe does NOT exist — legacy name, do not use)
|
||||
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract.
|
||||
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
|
||||
@@ -208,7 +165,7 @@ Plaintext keys removed from this contract post-vault-migration.
|
||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|
||||
|-------|-----|-----|---------------|------------|--------|
|
||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
|
||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | SSH root |
|
||||
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | SSH root |
|
||||
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
|
||||
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
|
||||
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
|
||||
@@ -233,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
|
||||
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
||||
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||
|
||||
## Prometheus & Grafana
|
||||
|
||||
@@ -250,11 +207,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
|
||||
- **RTX 3090 (2026-07-27)**: Swapped to SmartCode-Fable-5-27B-UD-Q3_K_XL (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`.
|
||||
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
|
||||
- **LiteLLM timeout tuning (verified 2026-07-27)**: SmartCode-Fable-5-27B 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf`, alias `strix-moe`, 256K context (n_ctx 262144), --parallel 2 --kv-unified, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
|
||||
@@ -263,19 +216,19 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
- **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled.
|
||||
- **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
|
||||
- **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
|
||||
- **Alias-retirement sweep (2026-09-12)**: `gemma-4-12b`, `gpu-light` and `crew-auto` are retired and replaced by `gpu-vision` / no cap respectively. The agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, `gpu-self-heal.prose.md`, `hermes-key-enforcement.prose.md`, `inference-optimization.prose.md`, `litellm-client-timeouts.prose.md` and the executable `audit-hermes-config.py` were all updated to the live canonical alias in the same change. **koby's config on .129 still names `gpu-light` (and `gemma-4-E4B`); .129 is report-only, so that is recorded for its owner and NOT edited here.**
|
||||
|
||||
## GPU Inference Benchmarks (Current)
|
||||
|
||||
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|
||||
|-----|-------|-----------|--------------|----------|---------|
|
||||
| RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q3_K_XL | **TBD** | — | — | **128K** |
|
||||
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
|
||||
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
||||
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
|
||||
| Strix Halo (.15) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe) | **65** | 140 | — | **256K** |
|
||||
|
||||
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
||||
Benchmarks from 2026-07-17. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe), n_ctx 262144, --parallel 2 --kv-unified. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||
GPU contexts: RTX 3090 (.8) and RTX 5070 (.110) at 128K; Strix Halo (.15) at 256K (2026-09-12).
|
||||
|
||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
|
||||
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
|
||||
History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
||||
|
||||
@@ -286,47 +239,46 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
||||
### Stable Aliases — CRITICAL
|
||||
|
||||
All agent configs MUST use stable role-based aliases, never model-specific names:
|
||||
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
|
||||
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
|
||||
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
|
||||
- `auxiliary.web_extract.model: gpu-light`
|
||||
- `compression.model: syslog-auto`
|
||||
- `auxiliary.vision.model: gpu-vision`
|
||||
- `delegation.model: gpu-dense`
|
||||
- `auxiliary.web_extract.model: gpu-vision`
|
||||
|
||||
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
|
||||
|
||||
### Context Windows
|
||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
|
||||
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12)
|
||||
- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek)
|
||||
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
|
||||
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||
- Mumuni compression model alias: `syslog-auto`
|
||||
|
||||
### Mumuni Agent Profile
|
||||
|
||||
Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This profile is the reference for all agent configs:
|
||||
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and `audit-hermes-config.py`; whether Mumuni's LIVE config currently complies is a separate operational question.
|
||||
|
||||
| Setting | Value | Notes |
|
||||
|---------|-------|-------|
|
||||
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
|
||||
| `model.default` | `syslog-auto` | Balanced default (pool router) |
|
||||
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
|
||||
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
|
||||
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
|
||||
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
|
||||
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
||||
| `compression.model` | `syslog-auto` | Rule 7: auto-routing, prevents Strix Halo overload |
|
||||
| `aux.compression.model` | `syslog-auto` | Must match `compression.model` (Rule 7) |
|
||||
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
|
||||
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
|
||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||
| `compression.threshold` | 0.65 | Triggers at ~85K |
|
||||
| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) |
|
||||
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
|
||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||
| `personalities` | `creative` | Creative assistant personality |
|
||||
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
||||
| Main model timeout | 300s | LiteLLM global timeout |
|
||||
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
|
||||
|
||||
### Agent Update Status (2026-07-15)
|
||||
|
||||
| Agent | Host | Status |
|
||||
|-------|------|--------|
|
||||
| **Mumuni** | CT100 (.24) | ✅ Updated to stable aliases |
|
||||
| **Mumuni** | CT105 (.14) | ✅ Updated to stable aliases |
|
||||
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
|
||||
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
|
||||
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
|
||||
|
||||
+99
-25
@@ -2,7 +2,7 @@
|
||||
kind: responsibility
|
||||
name: gpu-monitor
|
||||
description: >
|
||||
Comprehensive GPU fleet monitor — polls every subsystem (sidecars, router,
|
||||
Comprehensive GPU fleet monitor — polls every subsystem (sidecars,
|
||||
LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard,
|
||||
checks alert thresholds, and exposes a JSON API for downstream consumers.
|
||||
agent: abiba
|
||||
@@ -27,27 +27,38 @@ agent: abiba
|
||||
┌──────┐ ┌──────┐ ┌────────┐
|
||||
│.8:8080│ │.110 │ │.116:80 │
|
||||
│RTX3090│ │:8080 │ │nginx │
|
||||
│gemma │ │RTX5070│ │router │
|
||||
└──────┘ │qwen27B│ │LiteLLM │
|
||||
│qwen │ │RTX5070│ │router │
|
||||
└──────┘ │vision │ │LiteLLM │
|
||||
└──────┘ │dashboard│
|
||||
└────────┘
|
||||
```
|
||||
|
||||
Note: JSON sidecar exporters at :8090 were never deployed on any
|
||||
GPU host. Router falls back to GPU /health direct probe. Monitor
|
||||
should use router /health/unified as source of truth for GPU status.
|
||||
Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot
|
||||
poll .15:8080 directly; must go through router on .116.
|
||||
```
|
||||
|
||||
**PORT RULE (verified 2026-09-10):** GPU per-host health lives on **:8080**
|
||||
(`http://<gpu-host>:8080/health`); Prometheus GPU exporters live on **:9400**.
|
||||
There is NO listener on bare port 80 for any GPU host — `http://192.168.68.8/health`
|
||||
and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe
|
||||
as a GPU liveness signal: on 2026-09-09 that produced three false
|
||||
`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while
|
||||
`http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health`
|
||||
answered `200`. Port 80 is valid only on the harness host (.116), never on a GPU host.
|
||||
|
||||
### Subsystems Polled
|
||||
|
||||
| Subsystem | Endpoint | Frequency | Metrics |
|
||||
|-----------|----------|-----------|---------|
|
||||
| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) |
|
||||
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
|
||||
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
||||
| GPU Status (all, via fleet API) | `http://192.168.68.116/gpu/gpu-data` | 15s | models, CB, scores, GPU status from gpu-monitor on .24:9100 |
|
||||
| GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
|
||||
| GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
|
||||
| Fleet (unified) | `http://192.168.68.116/health/unified` | 15s | nginx `301` → `/gpu/gpu-data` served by gpu-monitor = alive (router decommissioned 2026-09-11) |
|
||||
| Harness (basic) | `http://192.168.68.116/health` | 15s | nginx → LiteLLM `/health/liveliness` |
|
||||
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
||||
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
||||
| Strix Halo | `http://192.168.68.116/health/unified` (nginx → fleet API) | 15s | Strix Halo status via gpu-monitor — cannot poll .15:8080 directly (firewalled to .116 only) |
|
||||
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
|
||||
|
||||
### Alert Delivery
|
||||
@@ -61,12 +72,28 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
|
||||
|
||||
## Alert Thresholds
|
||||
|
||||
### Liveness rule (scoped)
|
||||
|
||||
The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness
|
||||
endpoints, where any HTTP answer proves a listener is up. Applied here: nginx's
|
||||
`/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data`
|
||||
(the same payload) and LiteLLM's `/litellm/health` answers `301` →
|
||||
`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on
|
||||
**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included —
|
||||
and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as
|
||||
zulip-health (Tanko) and infrastructure-monitoring.
|
||||
|
||||
Probes whose success condition is specifically a bare `200` are NOT covered by
|
||||
the any-HTTP rule. On those — the GPU `:8080/health` endpoints, nginx `/health`,
|
||||
and the dashboard — an unexpected status (`401`/`403`, `5xx`, or
|
||||
anything other than the expected `200`) is an **ALERT**, not "alive".
|
||||
|
||||
| Metric | Warning | Critical |
|
||||
|--------|---------|----------|
|
||||
| GPU Temp | >80°C | >90°C |
|
||||
| VRAM Usage | >90% | >95% |
|
||||
| GPU Util | >95% | >98% |
|
||||
| Sidecar Unreachable | — | info (sidecars not deployed — use router /health/unified) |
|
||||
| Sidecar Unreachable | — | info (sidecars not deployed — use gpu-monitor /gpu-data) |
|
||||
| Model Down | — | critical (circuit breaker open) |
|
||||
|
||||
### JSON API Response Schema (/gpu-data)
|
||||
@@ -97,7 +124,47 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
|
||||
`curl http://localhost:9100/gpu-data | jq` — Full fleet status
|
||||
|
||||
### check-health
|
||||
`curl http://localhost:9100/health` — Monitor self-check
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Provenance — run first; paste the absolute path into the report
|
||||
pwd -P
|
||||
|
||||
# GPU Monitor health
|
||||
curl http://localhost:9100/health | jq
|
||||
# Expected: 200 with {"status": "healthy", "cache_age_seconds": <n>}
|
||||
|
||||
# GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host:
|
||||
# http://192.168.68.8/health has no listener and returns 000 → false DEGRADED.
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
|
||||
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
|
||||
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||
|
||||
# Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
|
||||
# Expected: 301 (or 200 after following the redirect) — any HTTP status = alive
|
||||
|
||||
# Router basic health (via nginx on port 80 — router .116 only, never a GPU host)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
|
||||
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||
|
||||
# LiteLLM health (via nginx on port 80)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
||||
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
|
||||
|
||||
# Dashboard
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/
|
||||
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the probe executed
|
||||
from** (`pwd -P`, or the monitor script's absolute path) so a stale-consumer
|
||||
report is distinguishable from a real fault at read time. Summarize actual
|
||||
results from each probe. Apply the scoped liveness rule above: on auth-gated
|
||||
endpoints only connection-refused (`000`) or timeout is DOWN; on bare-200 probes
|
||||
any other status is an alert. Never probe a GPU host on bare port 80.
|
||||
|
||||
### view-dashboard
|
||||
Open `http://localhost:9100/` in browser — Live HTML dashboard
|
||||
@@ -107,12 +174,14 @@ Open `http://localhost:9100/` in browser — Live HTML dashboard
|
||||
pkill -f gpu-monitor-server.py
|
||||
python3 /root/scripts/gpu-monitor-server.py &
|
||||
```
|
||||
Or via PM2: `pm2 restart gpu-monitor`
|
||||
Managed by systemd (verified 2026-09-11): `systemctl restart gpu-monitor`
|
||||
|
||||
### check-router
|
||||
The router health is accessed through nginx on port 80 (NOT port 9000 directly).
|
||||
`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy
|
||||
`curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only)
|
||||
### check-fleet
|
||||
Fleet health is accessed through nginx on port 80 on the harness host
|
||||
(.116) — NOT port 9000 (router decommissioned 2026-09-11), and NOT bare
|
||||
port 80 on a GPU host.
|
||||
`curl http://192.168.68.116/health/unified` — nginx answers `301` → `/gpu/gpu-data` (fleet monitor payload) = alive
|
||||
`curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED)
|
||||
|
||||
## Configuration Files
|
||||
|
||||
@@ -124,13 +193,18 @@ The router health is accessed through nginx on port 80 (NOT port 9000 directly).
|
||||
|
||||
## Execution
|
||||
|
||||
1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback)
|
||||
2. **Poll router** (every 15s): GET .116/health via nginx:80
|
||||
3. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80
|
||||
4. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only)
|
||||
5. **Poll dashboard** (every 15s): GET .116/dashboard/
|
||||
6. **Check alerts**: Compare metrics against thresholds
|
||||
7. **Compute summary**: Fleet-wide health aggregation
|
||||
8. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html
|
||||
9. **Serve API**: HTTP server on port 9100
|
||||
10. **Repeat** every 15 seconds
|
||||
**Port discipline:** probe GPU hosts on `:8080` (or the router's
|
||||
`/health/unified`); probe port 80 only on the router (.116). Never bare port 80
|
||||
on a GPU host.
|
||||
|
||||
1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback). `301` → `/gpu/gpu-data` counts as alive.
|
||||
2. **Fallback direct GPU probe** (only if router /health/unified is DOWN): GET `http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` — **:8080 only, never bare port 80**.
|
||||
3. **Poll router** (every 15s): GET .116/health via nginx:80
|
||||
4. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80
|
||||
5. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only)
|
||||
6. **Poll dashboard** (every 15s): GET .116/dashboard/
|
||||
7. **Check alerts**: Compare metrics against thresholds
|
||||
8. **Compute summary**: Fleet-wide health aggregation
|
||||
9. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html
|
||||
10. **Serve API**: HTTP server on port 9100
|
||||
11. **Repeat** every 15 seconds
|
||||
|
||||
+25
-19
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: responsibility
|
||||
name: gpu-self-heal
|
||||
description: >
|
||||
@@ -7,15 +9,16 @@ description: >
|
||||
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
||||
VRAM trend analysis, and predictive alerting.
|
||||
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
|
||||
Router (port 9000) references replaced with direct GPU routing.
|
||||
Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing.
|
||||
Benchmark baselines refreshed to live values.
|
||||
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
|
||||
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
|
||||
Stable role-based aliases (strix-moe, gpu-dense, gpu-vision) from gpu-fleet.
|
||||
agent: abiba
|
||||
depends_on:
|
||||
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
||||
---
|
||||
---
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -38,20 +41,19 @@ depends_on:
|
||||
- On fix: verify with benchmark inference test before declaring resolved
|
||||
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Current Fleet Baseline (2026-07-18)
|
||||
|
||||
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||
|-------|-----|------|-------|------|-----|-------|------|
|
||||
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
|
||||
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf | ~10/64GB (16%) | 256K | 62.9 | Compression, summarization, long docs |
|
||||
|
||||
Key notes:
|
||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
|
||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
|
||||
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
|
||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
||||
@@ -62,7 +64,7 @@ Key notes:
|
||||
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
||||
- **Fix**:
|
||||
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
||||
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
|
||||
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gpu-vision → gpu-dense, gpu-dense → gpu-vision)
|
||||
3. If all GPUs hot, alert about cooling infrastructure
|
||||
- **Verify**: Temp drops below 80°C within 5 minutes
|
||||
- **Escalate after**: 3 verification failures → Zulip alert
|
||||
@@ -102,10 +104,10 @@ Key notes:
|
||||
|
||||
### Rule 5: Circuit Breaker Stuck Open
|
||||
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
||||
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
|
||||
- **Note**: Router (port 9000) was decommissioned 2026-09-11. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
|
||||
- **Fix**:
|
||||
1. Verify GPU /health returns 200 on direct port (:8080)
|
||||
2. If GPU healthy, alert but do NOT reset via router API (deprecated)
|
||||
2. If GPU healthy, alert but do NOT reset via router API (decommissioned 2026-09-11)
|
||||
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
|
||||
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
|
||||
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
|
||||
@@ -140,10 +142,10 @@ Key notes:
|
||||
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
||||
|
||||
### Rule 9: Context Window Optimization
|
||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at its current context
|
||||
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
||||
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
||||
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||
- Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||
- **Fix**:
|
||||
- If tok/s > baseline → context has headroom, consider increasing
|
||||
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||
@@ -155,17 +157,19 @@ Key notes:
|
||||
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
|
||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||
- **Target distribution**:
|
||||
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
||||
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
|
||||
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
||||
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity).
|
||||
- RTX 5070 (gpu-vision, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks.
|
||||
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model).
|
||||
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
||||
- **Weights are not restated here** — the live `syslog-auto` pool weights and rpm caps live in CT 116 `/opt/inference-harness/litellm_config.yaml`, the single source of truth.
|
||||
- **Fix**:
|
||||
- Alert if any GPU is handling workload outside its designated role
|
||||
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
|
||||
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-vision, strix-moe)
|
||||
- Track per-GPU request distribution via LiteLLM spend logs
|
||||
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Execution
|
||||
@@ -247,6 +251,7 @@ call update-gpu-health
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Reporting
|
||||
@@ -263,6 +268,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
||||
- Per-GPU tok/s trend over 7 days
|
||||
- Regression alerts if any GPU degrades >10% week-over-week
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
||||
@@ -271,7 +277,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
||||
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
||||
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
||||
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
|
||||
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
|
||||
5. **CB auto-reset**: ✅ Router decommissioned 2026-09-11 — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
|
||||
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
|
||||
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
||||
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
|
||||
@@ -306,6 +312,6 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
||||
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
|
||||
|
||||
### L6: Stable Aliases Replace Model Names
|
||||
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
|
||||
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15; `gpu-light` was superseded by `gpu-vision` on 2026-09-12.
|
||||
- Self-heal must use aliases for reporting and alerting, not model-specific names.
|
||||
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
|
||||
|
||||
@@ -5,12 +5,30 @@ version: 1.0.0
|
||||
description: >
|
||||
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
||||
configuration state, keys, workarounds, and audit procedure. When an agent's
|
||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||
author: Abiba (pi agent)
|
||||
---
|
||||
|
||||
# Hermes Agent Baseline — Canonical Good State
|
||||
|
||||
## Reachability Detection
|
||||
|
||||
Before checking agent baseline, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
|
||||
|
||||
```bash
|
||||
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
|
||||
scripts/hermes-reachability-check.sh <host> "api_key:" "/root/.hermes/config.yaml"
|
||||
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key:" "/root/.hermes/config.yaml"
|
||||
|
||||
# Expected outcomes:
|
||||
# - UNREACHABLE: SSH connection failed (host is down)
|
||||
# - VIOLATION: SSH succeeded and found matches (report the finding)
|
||||
# - COMPLIANT: SSH succeeded and found no matches (no api_key in config)
|
||||
#
|
||||
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
|
||||
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
|
||||
```
|
||||
|
||||
## Quick Restore
|
||||
|
||||
```bash
|
||||
@@ -24,13 +42,12 @@ done
|
||||
|
||||
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|
||||
|-------|-----|------|-----|---------------|------------|----------|
|
||||
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
|
||||
| Mumuni | 100 | hwepve | .24 | `mumuni` | Infisical vault | Hermes |
|
||||
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
|
||||
| Koby | 111 | storepve | .129 | `koby` | Infisical vault | **Hermes** |
|
||||
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
|
||||
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
|
||||
|
||||
> **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
|
||||
> CT 111 (tdunna, 192.168.68.129, storepve) is report-only — Theo's box; alert only, never garbage-collect.
|
||||
|
||||
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
|
||||
Keys are stored in Infisical vault (project=agents, env=production) and injected at
|
||||
@@ -41,7 +58,7 @@ runtime via `infisical run --` wrapper. Plaintext keys removed from this baselin
|
||||
```
|
||||
Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime)
|
||||
↓
|
||||
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server)
|
||||
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
|
||||
└── Key DB (Postgres)
|
||||
```
|
||||
|
||||
@@ -53,9 +70,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) →
|
||||
|
||||
## Config Pattern — Mandatory Fields
|
||||
|
||||
### For Hermes Agents (Tanko, Mumuni, Koonimo)
|
||||
### For Hermes Agents (Mumuni, Koonimo)
|
||||
|
||||
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
|
||||
|
||||
### 1. Main Model
|
||||
```yaml
|
||||
@@ -71,7 +89,7 @@ model:
|
||||
custom_providers:
|
||||
- name: harness
|
||||
model: syslog-auto
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_mode: chat_completions
|
||||
```
|
||||
@@ -81,8 +99,8 @@ custom_providers:
|
||||
auxiliary:
|
||||
vision:
|
||||
provider: harness
|
||||
model: gemma-4-12b # or syslog-auto
|
||||
base_url: http://192.168.68.116/v1
|
||||
model: gpu-vision # RTX 5070 stable alias (Rule 8; do not use syslog-auto for aux)
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||
timeout: 60
|
||||
@@ -96,13 +114,33 @@ auxiliary:
|
||||
threshold: 0.65
|
||||
target_ratio: 0.3
|
||||
provider: harness
|
||||
model: syslog-auto # or gemma-4-12b
|
||||
base_url: http://192.168.68.116/v1
|
||||
model: syslog-auto # Rule 7: compression must be syslog-auto
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||
timeout: 120
|
||||
```
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
|
||||
|
||||
### FAULT (requires request-level evidence)
|
||||
- Agent's calls are failing with auth errors (401/403 in logs)
|
||||
- Agent's config has no valid API key AND calls are failing
|
||||
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
|
||||
|
||||
### Rules
|
||||
1. Do NOT infer the runtime's credential resolution from config text alone.
|
||||
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
|
||||
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
|
||||
4. State what you OBSERVED, not what the field implies.
|
||||
|
||||
## Known Bug: `api_key_env` Ignored by Auxiliary Client
|
||||
|
||||
**Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478)
|
||||
@@ -169,11 +207,15 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
|
||||
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
||||
```
|
||||
|
||||
### For Koby (CT 111 / tdunna)
|
||||
### For Koby (CT 111 / tdunna) — **REPORT-ONLY MODE**
|
||||
|
||||
Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
|
||||
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
|
||||
|
||||
**⛔ KOBY IS NEVER REPAIRED (2026-08-17, Captain)**: Diagnostic only — detect and report, never fix on .129.
|
||||
No heal step, no restart, no key rotation, no config edit, no memory rewrite, no disk GC, no service touch, no process kill — ever.
|
||||
If a health check shows Koby degraded, **DO NOT** execute any repair action. Instead, report to Zulip and let Theo fix it.
|
||||
|
||||
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
|
||||
|
||||
### For pi Agents (Abiba)
|
||||
@@ -182,7 +224,9 @@ Abiba (CT100) runs pi via PM2 with the Zulip extension.
|
||||
Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`.
|
||||
|
||||
**models.json** — Must only list models authorized for the agent's LiteLLM key.
|
||||
Key is injected via `infisical run --` wrapper at PM2 startup:
|
||||
`/v1/models` is key-scoped and the live registry is CT 116
|
||||
`/opt/inference-harness/litellm_config.yaml`; treat the list below as a snapshot and re-read
|
||||
the registry before applying. Key is injected via `infisical run --` wrapper at PM2 startup:
|
||||
```json
|
||||
{
|
||||
"providers": {
|
||||
@@ -194,9 +238,7 @@ Key is injected via `infisical run --` wrapper at PM2 startup:
|
||||
{ "id": "syslog-auto" },
|
||||
{ "id": "strix-moe" },
|
||||
{ "id": "gpu-dense" },
|
||||
{ "id": "gpu-light" },
|
||||
{ "id": "qwen3.6-27B-code" },
|
||||
{ "id": "gemma-4-12b" }
|
||||
{ "id": "gpu-vision" }
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
+167
-56
@@ -5,7 +5,8 @@ description: >
|
||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
|
||||
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
|
||||
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
|
||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
||||
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
||||
@@ -15,10 +16,28 @@ description: >
|
||||
|
||||
- template_version: "2.1.0"
|
||||
- last_applied: timestamp
|
||||
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
|
||||
- agents_configured: ["mumuni", "abiba", "koby", "koonimo", "kagenz0"] # tanko removed 2026-08-27 (now on DSH/DeepSeek Harness)
|
||||
- agent_keys: map (see Agent Keys section)
|
||||
- infra_endpoints_verified: array
|
||||
|
||||
## Reachability Detection
|
||||
|
||||
Before auditing the config template, verify the host is reachable and can be checked. Use the shared reachability helper from the clone root:
|
||||
|
||||
```bash
|
||||
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
|
||||
scripts/hermes-reachability-check.sh <host> "base_url:" "/root/.hermes/config.yaml"
|
||||
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "base_url:" "/root/.hermes/config.yaml"
|
||||
|
||||
# Expected outcomes:
|
||||
# - UNREACHABLE: SSH connection failed (host is down)
|
||||
# - VIOLATION: SSH succeeded and found matches (report the finding)
|
||||
# - COMPLIANT: SSH succeeded and found no matches (no base_url in config)
|
||||
#
|
||||
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
|
||||
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
|
||||
```
|
||||
|
||||
## Agent Keys (LiteLLM — Current 2026-07-11)
|
||||
|
||||
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
|
||||
@@ -30,8 +49,6 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
|
||||
|
||||
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
||||
|-------|-----------|------|-----|-----------|
|
||||
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
||||
| Mumuni | `mumuni` | 192.168.68.24 | root@.24 | 6 profiles ✱ |
|
||||
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
||||
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
||||
| Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — |
|
||||
@@ -94,17 +111,17 @@ work immediately after restart.
|
||||
```yaml
|
||||
# ─── Model Selection ───
|
||||
model:
|
||||
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
|
||||
default: <agent_model> # e.g., strix-moe, gpu-dense, syslog-auto
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
|
||||
context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K).
|
||||
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
|
||||
# and falls back to 256K when /v1/models lacks a context
|
||||
# field (llama-server does). Without this override, agents
|
||||
# silently run syslog-auto at 256K (verified 2026-08-09).
|
||||
# Set 65536 if using gemma-4-12b directly (tight VRAM).
|
||||
# Set 65536 if pinning a single model directly (tight VRAM).
|
||||
|
||||
fallback_providers:
|
||||
provider: deepseek
|
||||
@@ -129,13 +146,19 @@ mcp_servers:
|
||||
url: http://192.168.68.65:3100/mcp
|
||||
timeout: 120
|
||||
connect_timeout: 60
|
||||
litellm:
|
||||
url: https://litellm.sysloggh.net/mcp
|
||||
headers:
|
||||
x-litellm-api-key: "Bearer <AGENT_KEY>" # Rule 15: must be a REAL key (sk-...), not an env-var name
|
||||
# Note: MCP endpoint requires Accept: application/json, text/event-stream header
|
||||
# This is handled by the MCP client library; don't add to config
|
||||
|
||||
# ─── Compression ───
|
||||
compression:
|
||||
enabled: true
|
||||
model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
|
||||
provider: harness
|
||||
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
|
||||
max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12).
|
||||
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
||||
target_ratio: 0.30
|
||||
protect_last_n: 40
|
||||
@@ -145,49 +168,49 @@ compression:
|
||||
|
||||
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
||||
# All auxiliary services MUST use identical model, base_url, and api_key_env:
|
||||
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
|
||||
# base_url: http://192.168.68.116/v1
|
||||
# model: gpu-vision # stable alias (NOT a raw model name)
|
||||
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
# api_key_env: LITELLM_API_KEY
|
||||
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
||||
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
||||
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
||||
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
|
||||
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
|
||||
# in agent configs — use the stable aliases so model swaps don't break agents.
|
||||
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
||||
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
|
||||
auxiliary:
|
||||
vision:
|
||||
provider: harness
|
||||
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
|
||||
base_url: http://192.168.68.116/v1
|
||||
model: gpu-vision # stable alias for RTX 5070
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
timeout: 60
|
||||
download_timeout: 30
|
||||
web_extract:
|
||||
provider: harness
|
||||
model: gpu-light # stable alias for RTX 5070
|
||||
base_url: http://192.168.68.116/v1
|
||||
model: gpu-vision # stable alias for RTX 5070
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
timeout: 30
|
||||
compression:
|
||||
provider: harness
|
||||
model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
|
||||
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
|
||||
|
||||
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
|
||||
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
|
||||
# delegation.model and x_search.model use gpu-dense (NOT retired raw name).
|
||||
|
||||
delegation:
|
||||
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
|
||||
model: gpu-dense # stable alias for RTX 3090
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
|
||||
# ─── Custom Provider ───
|
||||
custom_providers:
|
||||
- name: harness
|
||||
model: syslog-auto # weighted pool (default)
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_mode: chat_completions
|
||||
```
|
||||
@@ -201,6 +224,60 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
||||
3. **After update**: Restart Hermes on the agent host
|
||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||
|
||||
## MCP Server Configuration
|
||||
|
||||
MCP server entries in `mcp_servers:` must follow the format shown in the Template section.
|
||||
|
||||
**Header requirements (Rule 15):**
|
||||
- Use `headers:` field with a `x-litellm-api-key` entry
|
||||
- The value must be `"Bearer <REAL_KEY>"` where `<REAL_KEY>` is a literal LiteLLM virtual key
|
||||
- Do NOT use env-var references like `$LITELLM_API_KEY` — they resolve to empty strings in
|
||||
the static config and cause "Malformed API Key" errors (2026-08-07 Tanko incident)
|
||||
|
||||
**Key source:**
|
||||
- Keys are stored in the Infisical vault (project=agents, env=production)
|
||||
- For template-based config generation: substitute the agent's key from the agent_keys table
|
||||
- For manual config updates: retrieve the key from the vault and insert the literal value
|
||||
|
||||
**Verification (2026-08-07):**
|
||||
- Tested MCP initialize handshake against litellm.sysloggh.net/mcp with agent virtual key
|
||||
- Confirmed: 200 response with `serverInfo.name: "litellm-mcp-server"`
|
||||
- Confirmed: tools/list returns 200 (MCP endpoint accessible with virtual keys)
|
||||
- Note: Per-key MCP grants are now supported (verified 2026-09-18), resolving the earlier
|
||||
contradiction with infrastructure-update.prose.md (which now reflects the update)
|
||||
- Key requirement: must be a valid LiteLLM virtual key (HTTP 200 on /v1/models)
|
||||
|
||||
**Key rotation note:**
|
||||
- MCP headers use literal keys (not env-vars), so they do NOT auto-rotate with the vault
|
||||
- After key rotation, MCP server headers must be regenerated with the new key value
|
||||
- This is a manual step: update the `x-litellm-api-key` header in each config file
|
||||
- TODO: Consider adding MCP header regeneration to the Key Update Procedure or a generation hook
|
||||
|
||||
**NetBird dependency:**
|
||||
- `litellm.sysloggh.net` is a NetBird endpoint (see Rule 5 and Infrastructure Stack table)
|
||||
- NetBird outages cause 502 errors on MCP requests, not auth failures
|
||||
- Diagnose: if MCP requests fail with 502, check NetBird status before investigating keys
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
|
||||
|
||||
### FAULT (requires request-level evidence)
|
||||
- Agent's calls are failing with auth errors (401/403 in logs)
|
||||
- Agent's config has no valid API key AND calls are failing
|
||||
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
|
||||
|
||||
### Rules
|
||||
1. Do NOT infer the runtime's credential resolution from config text alone.
|
||||
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
|
||||
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
|
||||
4. State what you OBSERVED, not what the field implies.
|
||||
|
||||
## Configuration Rules
|
||||
|
||||
### Rule 1: Shared Infra Is Locked
|
||||
@@ -237,10 +314,13 @@ The following MUST be identical across ALL profiles:
|
||||
- When main config uses `api_key_env`, sub-agents automatically use it
|
||||
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
|
||||
|
||||
### Rule 5: Main Config Base URL
|
||||
|- Use direct IP: `http://192.168.68.116/v1`
|
||||
### Rule 5: Main Config Base URL (UPDATED 2026-08-09)
|
||||
|- Use the authenticated LiteLLM path: `http://192.168.68.116/litellm/v1` (canonical, captain-approved migration)
|
||||
|- Legacy `http://192.68.68.116/v1` also works — nginx fronts BOTH paths with key auth
|
||||
(verified 2026-08-09: 401 without key, 200 with key, on both /v1 and /litellm/v1)
|
||||
|- Both locations have `proxy_read_timeout 600s` (verified in harness-nginx nginx.conf) —
|
||||
the old "60s timeout on /litellm/" claim was stale and is retracted
|
||||
|- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down
|
||||
|- NOT the old path (`/litellm/v1`) — nginx now routes `/v1` directly
|
||||
|
||||
### Rule 6: max_tokens Is Required (Thermal Safety)
|
||||
- **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable
|
||||
@@ -251,52 +331,54 @@ The following MUST be identical across ALL profiles:
|
||||
- For agents needing longer outputs: raise to 8192, but never omit
|
||||
|
||||
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
|
||||
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
|
||||
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
|
||||
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
|
||||
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
||||
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
|
||||
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
|
||||
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
|
||||
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
|
||||
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
||||
- **OPERATIONAL DECISION (2026-07-23): Use `syslog-auto` for compression across all agents.**
|
||||
The `syslog-auto` alias routes to the Strix Halo, but uses the weighted pool instead of pinning
|
||||
to `strix-moe` directly. This prevents sustained Strix Halo thermal load because the pool can
|
||||
fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model`
|
||||
MUST be `syslog-auto`.
|
||||
- All auxiliary services MUST use identical routing:
|
||||
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
|
||||
- `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK)
|
||||
- `api_key_env: LITELLM_API_KEY`
|
||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
||||
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
|
||||
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
||||
(64GB UMA, 128K context) — the designated compression GPU. This frees the
|
||||
(64GB UMA, 256K context) — the designated compression GPU. This frees the
|
||||
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
||||
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
|
||||
- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
|
||||
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
|
||||
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
|
||||
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
|
||||
- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
|
||||
- **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations
|
||||
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
|
||||
- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs
|
||||
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
|
||||
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
|
||||
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
|
||||
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
|
||||
- `auxiliary.compression.model: syslog-auto` (Strix Halo)
|
||||
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
|
||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
||||
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
|
||||
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 9: Compression Threshold for 128K Models
|
||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
||||
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
|
||||
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
|
||||
- **Koby Exception**: Per captain ruling 2026-08-11, Koby is a DeepSeek-primary external agent; its primary model remains `deepseek-v4-flash` (via api.deepseek.com to preserve DeepSeek-specific reasoning, while other sections follow Rule 10.
|
||||
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
||||
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
||||
- `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe
|
||||
and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
|
||||
- `syslog-auto` is the LiteLLM routing model — it load-balances across the live pool
|
||||
(see CT 116 `/opt/inference-harness/litellm_config.yaml` for the current members and weights). Using it protects against:
|
||||
- Model name typos that cause 403 errors and silent worker failures
|
||||
- Single GPU downtime (routing falls back automatically)
|
||||
- Key/model authorization mismatches
|
||||
@@ -319,12 +401,16 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
|
||||
verify ALL FOUR of these against the live config. They are the only root causes found in production:
|
||||
|
||||
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
|
||||
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
|
||||
MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K).
|
||||
A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used.
|
||||
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
|
||||
2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`,
|
||||
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
|
||||
has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s.
|
||||
Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING)
|
||||
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
|
||||
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
|
||||
or `http://192.168.68.116/v1` (legacy, still authenticated via nginx). BOTH verified 200 with
|
||||
key + 600s proxy_read_timeout on 2026-08-09. Never bare `:4000` direct.
|
||||
Check: `grep -nE 'base_url: http://192.168.68.116(:4000)?/v1' ~/.hermes/config.yaml` — the
|
||||
ONLY paths allowed are `/v1` or `/litellm/v1` (both via nginx :80).
|
||||
`:4000` or missing `litellm/v1`/`v1` prefix = violation.
|
||||
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
|
||||
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
|
||||
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
|
||||
@@ -341,6 +427,17 @@ curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $
|
||||
|
||||
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
|
||||
|
||||
### Rule 14: Hermes Context Detection Uses `max_model_tokens`, NOT `max_input_tokens`
|
||||
|
||||
**CRITICAL**: Hermes context detection reads `max_model_tokens` (128K), NOT `max_input_tokens` (64K cap).
|
||||
|
||||
- **Abiba and Hermes agents**: `max_model_tokens: 131072` (128K) — unlimited context
|
||||
- **Crewmates (ops, tune, verify, auth-keys, build)**: `max_input_tokens: 64000` (64K) — capped
|
||||
- If you see `max_input_tokens: 64000` in an Abiba/Hermes config, that's a mistake
|
||||
- Using `max_input_tokens` for Hermes agents causes premature context loss
|
||||
- Check: `grep -n 'max_model_tokens\|max_input_tokens' ~/.hermes/config.yaml`
|
||||
- Expected output: `max_model_tokens: 131072` (not max_input_tokens)
|
||||
|
||||
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
|
||||
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
|
||||
|
||||
@@ -399,14 +496,28 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
||||
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
||||
before and after any config change to catch this and all other rule violations.
|
||||
|
||||
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
|
||||
- Every MCP server entry must point at the correct endpoint:
|
||||
- ra-h-os = http://192.168.68.65:3100/mcp
|
||||
- litellm = https://litellm.sysloggh.net/mcp
|
||||
- MCP entries must carry a REAL key value in the header.
|
||||
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
|
||||
endpoints and result in "Malformed API Key" floods.
|
||||
- Ensure the header value is the actual key (e.g., `sk-...`).
|
||||
### Rule 15: MCP Endpoint and Header Validation (UPDATED 2026-08-07)
|
||||
|
||||
**Endpoint validation:**
|
||||
- ra-h-os must point to `http://192.168.68.65:3100/mcp`
|
||||
- litellm must point to `https://litellm.sysloggh.net/mcp`
|
||||
- Mismatched endpoints cause silent failures (e.g., 2026-08-07 incident: Tanko's config had
|
||||
ra-h-os pointing to litellm's endpoint)
|
||||
|
||||
**Header validation:**
|
||||
- Every MCP entry with authentication must carry a `headers:` field
|
||||
- The header value must be a REAL key (e.g., `Bearer sk-abc123...`), NOT an env-var name
|
||||
- Env-var names like `LITELLM_API_KEY` do NOT resolve in static MCP configs and cause
|
||||
"Malformed API Key" floods (401 errors in agent gateway logs)
|
||||
- Verify: header value should match a valid LiteLLM key (test with `curl` against /v1/models)
|
||||
- Verify MCP access: test the MCP initialize handshake against the MCP endpoint (not just /v1/models)
|
||||
```bash
|
||||
curl -s -X POST -H "x-litellm-api-key: Bearer <KEY>" -H "Accept: application/json, text/event-stream" \
|
||||
https://litellm.sysloggh.net/mcp -d '{"jsonrpc":"2.0","id":1,"method":"initialize",...}' \
|
||||
| jq '.data.result.serverInfo' # should show serverInfo.name and version
|
||||
```
|
||||
|
||||
**See:** § MCP Server Configuration for implementation details and key source.
|
||||
|
||||
## Execution
|
||||
|
||||
|
||||
+100
-26
@@ -16,7 +16,26 @@ author: Abiba (pi agent)
|
||||
|
||||
## Rule (One Sentence)
|
||||
|
||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with authenticated path `http://192.168.68.116/litellm/v1/responses` — hardcoded keys AND unauthenticated `/v1` direct access are both forbidden.**
|
||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
|
||||
|
||||
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
|
||||
|
||||
## Model Access Tiers (2026-09-20)
|
||||
|
||||
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
|
||||
|
||||
| Tier | Models | Who gets it |
|
||||
|------|--------|-------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
|
||||
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
|
||||
|
||||
**Rules:**
|
||||
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
|
||||
2. Agent keys MUST carry an **explicit local-only** `models` list.
|
||||
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
|
||||
4. The master key always bypasses scoping — it is admin-only, never for inference.
|
||||
|
||||
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
|
||||
|
||||
## Scope
|
||||
|
||||
@@ -34,11 +53,12 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
|
||||
|
||||
| Path | Auth | Status |
|
||||
|------|------|--------|
|
||||
| `http://192.168.68.116/v1` | None (direct) | ❌ **DEPRECATED** — being phased out |
|
||||
| `http://192.168.68.116/litellm/v1/responses` | Bearer `sk-*` key | ✅ **CURRENT** — authenticated LiteLLM proxy |
|
||||
| `http://192.168.68.116/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **VALID** — authenticated via nginx :80 (verified 2026-08-09: 401 without key, 200 with) |
|
||||
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
|
||||
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
|
||||
|
||||
All harness/litellm providers MUST use the authenticated `/litellm/v1/responses` path.
|
||||
Any `base_url` pointing to bare `/v1` on 192.168.68.116 is a **migration violation**.
|
||||
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` non-canonical but working).
|
||||
The public host `https://litellm.sysloggh.net` serves `/v1` ONLY (404 on `/litellm/v1`).
|
||||
|
||||
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
|
||||
|
||||
@@ -100,7 +120,7 @@ auxiliary:
|
||||
fallback_providers:
|
||||
- provider: deepseek
|
||||
base_url: https://api.deepseek.com
|
||||
api_key: sk-b7d9... # ← hardcoded OK (external)
|
||||
api_key: sk-synthetic-external-example # ← hardcoded OK (external, synthetic example)
|
||||
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
|
||||
```
|
||||
|
||||
@@ -108,14 +128,62 @@ fallback_providers:
|
||||
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
|
||||
model:
|
||||
provider: harness
|
||||
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION: hardcoded key
|
||||
api_key: sk-synthetic-example-12345 # ← RULE VIOLATION: hardcoded key (synthetic example)
|
||||
|
||||
model:
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/v1 # ← RULE VIOLATION: unauthenticated path
|
||||
base_url: http://192.168.68.116/v1 # ← NON-CANONICAL but WORKING (authenticated via nginx, WARN not FAIL)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
```
|
||||
|
||||
## Reachability Detection
|
||||
|
||||
Before checking for hardcoded keys, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
|
||||
|
||||
```bash
|
||||
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
|
||||
scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
|
||||
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key: sk-" "/root/.hermes/"
|
||||
|
||||
# Expected outcomes:
|
||||
# - UNREACHABLE: SSH connection failed (host is down)
|
||||
# - VIOLATION: SSH succeeded and found matches (report the finding)
|
||||
# - COMPLIANT: SSH succeeded and found no matches (no hardcoded keys in config)
|
||||
#
|
||||
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
|
||||
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
|
||||
```
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### ACCEPTABLE PATTERN
|
||||
Agent keys live in `.env` or `.env.vault` files with 600 permissions (koonimo's shape is the canonical example). A plaintext key inside a `config.yaml` or any `config.yaml.bak-*` file is a violation — the backup files are not part of the runtime credential path and are not watched by the scanner, so a key in them is stale clutter that a future reader can mistake for a working key.
|
||||
|
||||
**Fix procedure** (when a backup file is found with a plaintext key):
|
||||
1. Move the file out of the scanned tree (e.g., `mv /root/.hermes/config.yaml.bak-* /root/hermes-config-backups/`) — do NOT delete the file, just move it so the scanner pattern no longer matches.
|
||||
2. Re-run the reachability check to confirm COMPLIANT.
|
||||
3. Report the before/after check output and the commands you ran.
|
||||
|
||||
**Rationale**: Moving the file preserves history without leaving a credential where a scanner trips over it. Deleting the file loses the historical context. Keeping it in place means the next scan will report it as a finding and waste time re-deciding.
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
|
||||
|
||||
### FAULT (requires request-level evidence)
|
||||
- Agent's calls are failing with auth errors (401/403 in logs)
|
||||
- Agent's config has no valid API key AND calls are failing
|
||||
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
|
||||
|
||||
### Rules
|
||||
1. Do NOT infer the runtime's credential resolution from config text alone.
|
||||
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
|
||||
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
|
||||
4. State what you OBSERVED, not what the field implies.
|
||||
|
||||
## Detection Query
|
||||
|
||||
Run on any Hermes host to detect violations:
|
||||
@@ -133,7 +201,7 @@ grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
|
||||
|
||||
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
||||
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /root/.config/systemd/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-…' /root/.config/systemd/ 2>/dev/null
|
||||
|
||||
# 3. Verify running process env matches dedicated key
|
||||
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
||||
@@ -168,16 +236,22 @@ The agent picks up the new key via `infisical run --` at gateway startup.
|
||||
|
||||
**Keys are permanent and use bare agent name aliases.**
|
||||
|
||||
- **Duration**: `null` — keys never expire. This is enforced by `default_key_generate_params` in `litellm_config.yaml`.
|
||||
- **Duration**: `null` — keys never expire by default. **Expiry must be set EXPLICITLY at creation** with the `duration` parameter (e.g., `90d` for 90 days). The 90-day default is the standard; however, the config default is **NOT honoured** by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns `expires=null`). This has been recorded in `/opt/inference-harness/litellm_config.yaml` to prevent re-filing as a bug.
|
||||
- **Daily Audit**: A daily audit job runs at 00:00 UTC (`/usr/local/bin/litellm-key-renewal-ct116.sh`, cron 00:00). It is **AUDIT-ONLY** and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES `abiba-pi` and `koby` (report-only, and .129 must never be touched), and logs `RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED` when renewal is skipped. **Renewal is NOT implemented** — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end.
|
||||
- **Exclusions**: `abiba-pi` and every firstmate/secondmate/crewmate key stay **WITHOUT an expiry** until a proven renewal path exists. `koby` is **report-only** (never touched). These exclusions are enforced by the audit job.
|
||||
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
|
||||
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
|
||||
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use.
|
||||
- **Max budget**: $100 per key (config default).
|
||||
|
||||
```yaml
|
||||
# In litellm_config.yaml — ensures all future keys inherit these defaults:
|
||||
# NOT currently set in the authority; recommended value. CT 116 litellm_config.yaml has no
|
||||
# default_key_generate_params block today, and a key generated with no explicit models comes back
|
||||
# with an EMPTY models list. `models` is a literal key-generation parameter, so this is a value to
|
||||
# ADD — re-read the live registry at CT 116 /opt/inference-harness/litellm_config.yaml and
|
||||
# re-verify before applying.
|
||||
litellm_settings:
|
||||
default_key_generate_params:
|
||||
models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b"]
|
||||
models: ["syslog-auto", "gpu-dense", "gpu-vision", "strix-moe"]
|
||||
duration: null # ← permanent
|
||||
max_budget: 100
|
||||
metadata:
|
||||
@@ -189,9 +263,9 @@ litellm_settings:
|
||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|
||||
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
||||
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
|
||||
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
|
||||
| Koonimo | 113 | .113 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
|
||||
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | ✅ Fixed | systemd Hermes gateway | 2026-08-29 |
|
||||
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed (DeepSeek-primary) | `infisical run` | 23:30 UTC Jul 5 |
|
||||
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
|
||||
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
||||
| Kagenz0 | 105 | .14 | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
|
||||
|
||||
@@ -200,12 +274,12 @@ litellm_settings:
|
||||
|
||||
### Migration Status: Authenticated Path
|
||||
|
||||
| Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status |
|
||||
|-------|--------------------------|--------------------|--------|
|
||||
| Mumuni | ✅ 5 sections | 0 | ✅ Authenticated |
|
||||
| Tanko | ⚠️ No SSH access | — | Needs check |
|
||||
| Koby | ⚠️ No route to host | — | Needs check |
|
||||
| Koonimo | ⚠️ Connection timed out | — | Needs check |
|
||||
| Agent | `/litellm/v1` | Legacy `/v1` | Status |
|
||||
|-------|--------------|-------------|--------|
|
||||
| Mumuni | ✅ harness provider | ✅ auxiliary on /v1 (valid) | ✅ Authenticated (verified 2026-08-09) |
|
||||
| Tanko | ✅ 5 sections | 0 | ✅ Migrated 2026-08-08, keys 200 |
|
||||
| Koby | ✅ custom provider (harness name) | — | ✅ External DeepSeek primary (intentional, captain ruling 2026-08-11) |
|
||||
| Koonimo | ✅ .114 (baggy) | — | ✅ 128K context applied 2026-08-09 |
|
||||
|
||||
### Systemd Service Pattern (2026-07-11 — vault migration)
|
||||
|
||||
@@ -284,14 +358,14 @@ auxiliary:
|
||||
vision:
|
||||
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
model: gemma-4-12b
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
compression:
|
||||
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
model: gemma-4-12b
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
```
|
||||
|
||||
|
||||
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
|
||||
| Param | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) |
|
||||
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
||||
|
||||
## Maintains
|
||||
@@ -47,7 +47,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to target host (direct or via amdpve for CTs)
|
||||
- SSH access to target host (direct, or via the guest's Proxmox node for CTs)
|
||||
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
|
||||
- Python 3 with `httpx` installed on target
|
||||
|
||||
@@ -55,9 +55,8 @@ connectivity recovery including end-to-end DM validation.
|
||||
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Mumuni | CT100 | — | 192.168.68.24 | /root/.hermes | root |
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
|
||||
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||
|
||||
| Field | Value | Trust |
|
||||
@@ -73,7 +72,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
### Step 1: Resolve Target
|
||||
|
||||
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
|
||||
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
|
||||
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||
|
||||
### Step 2: Pull Latest Plugin Source
|
||||
|
||||
@@ -121,7 +120,8 @@ cp plugins/platforms/zulip/adapter.py \
|
||||
plugins/platforms/zulip/plugin.yaml \
|
||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
# Fix ownership (Tanko only — runs as jerome user)
|
||||
# Fix ownership (was Tanko-only, runs as jerome user)
|
||||
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
|
||||
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: function
|
||||
name: hermes-zulip-restore
|
||||
description: >
|
||||
Restores Zulip connectivity for any Hermes agent (Mumuni CT100, Tanko CT112,
|
||||
Koby CT111, Shumba on Lucky's mini PC). Deploys the zulip-platform adapter to the correct bundled plugin
|
||||
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
||||
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
||||
a fresh agent deployment.
|
||||
@@ -12,6 +12,7 @@ version: 1.0.0
|
||||
status: active
|
||||
runtime_contract: 2
|
||||
---
|
||||
---
|
||||
|
||||
# Hermes Zulip Restore — Bring Any Agent Back to Good State
|
||||
|
||||
@@ -23,7 +24,7 @@ gateway restart, and connection validation.
|
||||
|
||||
| Param | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) |
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -36,13 +37,13 @@ gateway restart, and connection validation.
|
||||
|
||||
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
|
||||
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
|
||||
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko)
|
||||
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko, historical — DSH since 2026-08-27)
|
||||
- Gateway restarted and zulip platform reports state `connected`
|
||||
- HTML stripping enabled for `/approve` and `/deny` slash command support
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to target host (direct or via amdpve for CTs)
|
||||
- SSH access to target host (direct, or via the guest's Proxmox node for CTs)
|
||||
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
|
||||
- Python 3 with `httpx` installed on target
|
||||
- Zulip server accessible at `https://chat.sysloggh.net`
|
||||
@@ -51,9 +52,7 @@ gateway restart, and connection validation.
|
||||
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Mumuni | CT100 (abiba) | hwepve | 192.168.68.24 | /root/.hermes | root |
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||
|
||||
| Field | Value | Trust |
|
||||
@@ -68,7 +67,7 @@ gateway restart, and connection validation.
|
||||
### Step 1: Locate Target
|
||||
|
||||
Map `target` to connectivity parameters from the live-state table above.
|
||||
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
|
||||
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||
|
||||
### Step 2: Deploy Zulip Adapter
|
||||
|
||||
@@ -94,8 +93,8 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
|
||||
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
||||
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
# Fix ownership (Tanko only)
|
||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only
|
||||
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin)
|
||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
|
||||
|
||||
# Clean up
|
||||
rm -rf /tmp/zulip-deploy
|
||||
@@ -184,6 +183,7 @@ https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins/src/branch/feat/z
|
||||
Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching`
|
||||
Pull request #33 is the primary integration branch.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
**Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied.
|
||||
|
||||
@@ -4,7 +4,7 @@ kind: responsibility
|
||||
description: >
|
||||
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
|
||||
assignments, agent context management, and prompt caching — to reduce response
|
||||
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
|
||||
times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12).
|
||||
id: 067NC6KP02RG60S50M40E30928
|
||||
---
|
||||
|
||||
@@ -23,7 +23,7 @@ management, and prompt caching — without sacrificing agent capability.
|
||||
.123, any others on .129/.122) including compression, model, context_window,
|
||||
prompt_caching, memory settings
|
||||
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
|
||||
qwen .8:8080, gemma .110:8080)
|
||||
gpu-dense .8:8080, gpu-vision .110:8080)
|
||||
|
||||
### Maintains
|
||||
|
||||
@@ -56,14 +56,14 @@ duration.
|
||||
**Context is the root cause.** Every ~46K prompt token costs ~87s of
|
||||
prefill time at 532 tok/s. Fix context first, routing second.
|
||||
|
||||
- **Route by task**: qwen for code/standard queries; gemma for
|
||||
compression/auxiliary; strix-moe for compression tasks.
|
||||
- **Route by task**: gpu-dense for code/standard queries; gpu-vision for
|
||||
vision/web-auxiliary; syslog-auto for compression.
|
||||
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
|
||||
compact at 51K, not 85K. Target 15% tail (not 30%).
|
||||
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
||||
never change between turns. Single-digit cache hit rate is unacceptable.
|
||||
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
||||
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||
GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||
|
||||
### Shape
|
||||
|
||||
@@ -87,8 +87,8 @@ call apply-liteLLM-routing
|
||||
|
||||
call apply-agent-compression
|
||||
agent: mumuni
|
||||
host: 192.168.68.24
|
||||
config_path: /root/.hermes/config.yaml
|
||||
host: 192.168.68.14
|
||||
config_path: /home/hermes/.hermes/config.yaml
|
||||
|
||||
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
|
||||
|
||||
@@ -96,8 +96,10 @@ call enable-prompt-caching
|
||||
hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110]
|
||||
|
||||
-- Phase 5: Verify end-to-end latency
|
||||
-- `models` is a literal verification parameter (a snapshot only): the authoritative registry is
|
||||
-- CT 116 /opt/inference-harness/litellm_config.yaml; re-read it before use.
|
||||
|
||||
call verify-latency
|
||||
host: 192.168.68.116
|
||||
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
|
||||
models: [syslog-auto, gpu-dense, gpu-vision, strix-moe]
|
||||
```
|
||||
|
||||
+129
-49
@@ -3,7 +3,7 @@ kind: pattern
|
||||
name: infrastructure-control
|
||||
description: >
|
||||
Full infrastructure monitoring and control pattern covering the
|
||||
6-node Proxmox cluster, 3 Docker ecosystems (22 containers),
|
||||
5-node Proxmox cluster, 3 Docker ecosystems (22 containers),
|
||||
NFS storage, and network services. Defines monitors, remediations,
|
||||
and the access matrix for all environments.
|
||||
|
||||
@@ -13,10 +13,11 @@ description: >
|
||||
against the live system. Policy fields are authoritative. See the
|
||||
`verify-before-mutate` skill.
|
||||
|
||||
**Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110),
|
||||
AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve),
|
||||
Abiba placement (hwepve not amdpve), added hwepve as 6th node,
|
||||
added dns.sysloggh.net route.
|
||||
**Last verified:** 2026-08-15 — hwepve removed from Tabiri cluster
|
||||
(now 5 nodes: minipve, amdpve, storepve, acerpve, ocupve). hwepve
|
||||
(192.168.68.4) is a standalone PVE node + NetBird routing peer;
|
||||
London relocation pending. CT 100 (abiba) is on minipve; CT 105
|
||||
(kagentz) is on amdpve.
|
||||
---
|
||||
|
||||
# Infrastructure Control Pattern
|
||||
@@ -34,21 +35,21 @@ description: >
|
||||
┌─────────────┐ ┌──────────────┐ ┌──────────────┐
|
||||
│ Abiba │ │ Tanko │ │ Mumuni │
|
||||
│ (pi) │ │ (Hermes) │ │ (Hermes) │
|
||||
│ CT 100 │ │ CT 112 │ │ CT 114 │
|
||||
│ CT 100 │ │ CT 112 │ │ CT 100 │
|
||||
└──────┬──────┘ └──────┬───────┘ └──────┬───────┘
|
||||
│ │ │
|
||||
└──────────────────┼────────────────────┘
|
||||
▼
|
||||
┌──────────────────────────────────────┐
|
||||
│ Proxmox Cluster API │
|
||||
│ minipve.sysloggh.net:443 │
|
||||
│ (monitoring@pve!mumuni token) │
|
||||
└────┬──────┬──────┬──────┬──────┬─────┘
|
||||
│ │ │ │ │
|
||||
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
|
||||
▼ ▼ ▼ ▼ ▼
|
||||
minipve amdpve storepve acerpve ocupve hwepve
|
||||
(.12) (.15) (.6) (.9) (.5) (.4)
|
||||
┌────────────────────────────────────┐
|
||||
│ Proxmox Cluster API │
|
||||
│ minipve.sysloggh.net:443 │
|
||||
│ (monitoring@pve!mumuni token) │
|
||||
└────┬──────┬──────┬──────┬──────────┘
|
||||
│ │ │ │
|
||||
┌────┘ ┌────┘ ┌────┘ ┌────┘
|
||||
▼ ▼ ▼ ▼
|
||||
minipve amdpve storepve acerpve ocupve
|
||||
(.12) (.15) (.6) (.9) (.5)
|
||||
|
||||
▼
|
||||
┌─────────────────────────────────────────────┐
|
||||
@@ -100,21 +101,29 @@ description: >
|
||||
|
||||
## Section 2: Proxmox Cluster — Monitoring
|
||||
|
||||
### Nodes (6)
|
||||
### Nodes (5)
|
||||
|
||||
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
||||
|------|----|-----|-----|---------|------|
|
||||
| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
|
||||
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
|
||||
| minipve | .12 | 16C | 30GB | abiba, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||
| amdpve | .15 | 32C | 62GB | kagentz, tanko, baggy, scottdenya, adguard2 | Agents, compute |
|
||||
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
|
||||
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
|
||||
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
||||
| hwepve | .4 | 12C | 15GB | abiba, kagentz, (mumuni CT 114 stopped) | Agents (new node) |
|
||||
|
||||
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
|
||||
> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni
|
||||
> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a
|
||||
> second instance on minipve at .123 — distinguish by CT ID, not hostname.
|
||||
> **Note:** CTs on storepve include jdownloader (CT 118) and tdunna (CT 111).
|
||||
> AdGuard (CT 102) is on minipve at .10, not acerpve. Abiba (CT 100) is on
|
||||
> minipve (moved from hwepve 2026-08-15); kagentz (CT 105) is on amdpve.
|
||||
> Mumuni runs inside Abiba CT100 (.24); CT 114 (mumuni) no longer exists in the cluster.
|
||||
>
|
||||
> **hwepve (192.168.68.4) — STANDALONE (removed from Tabiri 2026-08-15):**
|
||||
> Huawei MateBook 16 (KLVL-WXX9), pve-manager/9.2.10, kernel 7.0.14-8-pve.
|
||||
> Zero VMs/CTs. Being relocated to London as a standalone PVE node + NetBird
|
||||
> routing peer (relocation pending). Localizations applied: timezone
|
||||
> Europe/London, lid-switch ignore, sleep/suspend/hibernate targets masked,
|
||||
> cluster-shared storage removed (remaining: local, local-lvm, storage,
|
||||
> mediastore). prometheus-node-exporter active on :9100; net.ipv4.ip_forward=1;
|
||||
> NetBird client not yet installed (enrollment pending setup key).
|
||||
|
||||
### Checks (every 5 min)
|
||||
|
||||
@@ -172,8 +181,8 @@ description: >
|
||||
**Stirling-PDF** (deployed 2026-07-03, Authentik SSO 2026-07-03):
|
||||
- URL: `https://pdf.sysloggh.net` (public) / `http://192.168.68.7:8989` (direct)
|
||||
- Swagger: `http://192.168.68.7:8989/swagger-ui.html`
|
||||
- Admin credentials: `admin` / `kakashi20stirling`
|
||||
- API key: `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88`
|
||||
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_USER»` / `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||
- API key: `«vault: infrastructure/production STIRLING_API_KEY»`
|
||||
- Authentik OAuth2: configured but disabled (requires paid Server license). Ready to enable: set `SECURITY_OAUTH2_ENABLED=true` + `SECURITY_LOGINMETHOD=all`
|
||||
- Compose: `/opt/home_stack/docker-compose.yml`
|
||||
- Control script: `/opt/home_stack/infra-control.sh`
|
||||
@@ -190,18 +199,18 @@ description: >
|
||||
|
||||
### Ecosystem B: CT 116 syslog-api (192.168.68.116)
|
||||
|
||||
8 containers in inference-harness stack:
|
||||
12 containers on CT 116 — 11 in the inference-harness stack + trove-agent-docker (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1; trove-agent-docker added 2026-09-11):
|
||||
|
||||
| Container | Image | Port | Role |
|
||||
|-----------|-------|------|------|
|
||||
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | API proxy, key mgmt, fallbacks |
|
||||
| harness-router | inference-harness-router | :9000 (127.0.0.1) | GPU routing, slot booking, CB |
|
||||
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | API proxy, key mgmt, fallbacks |
|
||||
| harness-nginx | nginx:alpine | :80 | Entrypoint, /v1→LiteLLM, /dashboard/ |
|
||||
| harness-postgres | postgres:16-alpine | :5432 | LiteLLM DB (keys, spend, config) |
|
||||
| harness-redis | redis:7-alpine | :6379 | Router slots, circuit breakers |
|
||||
| harness-redis | redis:7-alpine | :6379 | LiteLLM cache + rate-limit state |
|
||||
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
|
||||
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
|
||||
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
|
||||
| trove-agent-docker | ghcr.io/techdox/trove-agent-docker:latest | outbound agent (no port) | Trove host agent — service inventory + metrics, added 2026-09-11 |
|
||||
|
||||
**Nginx routing**:
|
||||
- `/v1/*` → harness-litellm:4000 (API)
|
||||
@@ -213,9 +222,8 @@ description: >
|
||||
|
||||
**Prometheus targets**:
|
||||
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
||||
- 192.168.68.110:9400 (RTX 5070 — gemma)
|
||||
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
||||
- 192.168.68.24:9401 (Router metrics exporter)
|
||||
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
|
||||
- 192.168.68.15:9400 (Strix Halo — strix-moe)
|
||||
- harness-litellm:4000 (LiteLLM health)
|
||||
|
||||
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
|
||||
@@ -356,6 +364,79 @@ For docker-vm specifically:
|
||||
- No PBS backup in 48h → fail
|
||||
```
|
||||
|
||||
### Backup Safety Preconditions (2026-09-15)
|
||||
|
||||
#### Background & Rationale
|
||||
|
||||
Two incidents from 2026-09-13/14 demonstrate that backup operations can catastrophically fail when storage conditions are not verified first:
|
||||
|
||||
1. **acerpve thin-pool VM 101** (acerpve, 192.168.68.9, 2026-09-13): A snapshot-mode vzdump of VM 101 on acerpve filled the LVM thin pool. `dmsetup status pve-data-tpool` showed `thin-pool Error` (then `Fail`), the host root remounted `emergency_ro`, ordinary commands failed with I/O errors, LVM tools returned nothing and VM 101 (the RTX 3090 host) went unreachable while the host still answered ping and ssh. It happened TWICE in one day with different modes: snapshot at 13:39Z and a `--mode stop` cold run at 18:52Z. Both times a reboot rolled the failed transaction back and the pool returned rw (~30% data, ~1.2% metadata). Pool capacity was NOT the obvious explanation - ~816G with ~572G free - which is why the metadata/snapshot-pressure hypothesis stands unproven. A full or errored thin pool fails EVERY volume on the VG at once, including the host root.
|
||||
|
||||
2. **amdpve 0700 tmpdir** (amdpve, 192.168.68.15, 2026-09-14): A custom vzdump `tmpdir` created with mode 0700 broke a whole night of container backups: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES`, because the archive step runs through an unprivileged user namespace and could not traverse a root-owned 0700 directory. Fixed with `chmod 1777` (match /var/tmp) and proved with a real backup.
|
||||
|
||||
3. **acerpve GPU-host fact**: VM 101 (llm-gpu) and VM 103 (ocu-llm) are in NO scheduled job, so their only coverage is one-off runs - and for VM 101 that is deliberate until the thin-pool is understood.
|
||||
|
||||
> ⚠️ **Hostname Resolution Warning (2026-09-15)**: The PVE node hostnames (acerpve, amdpve, minipve, storepve, ocupve) all resolve to the VPS (72.61.0.17, the Netbird VPS at srv1079750.hstgr.cloud) via the wildcard `*.dns.sysloggh.net` record, NOT to the actual nodes. So `ssh acerpve` lands on the VPS. **Nodes must be addressed by IP**: acerpve 192.168.68.9, amdpve 192.168.68.15, storepve 192.168.68.6, minipve 192.168.68.12, ocupve 192.168.68.5. Guest CTs are reached through their node (`pct exec`). Guest hostnames that resolve on the LAN (e.g. kagentz = 192.168.68.14) are fine. (The DNS address records are a separate decision — row: dag-daemon-node-hostnames-resolve-to-the-vps-20260915.)
|
||||
|
||||
#### PREFLIGHT Preconditions (Before ANY snapshot-mode backup on thin-pool hosts)
|
||||
|
||||
Before starting ANY snapshot-mode vzdump on a host whose storage is an LVM thin pool, the following checks MUST pass:
|
||||
|
||||
```bash
|
||||
# Check 1: Pool headroom (PRIMARY - yields percentages directly)
|
||||
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname, see warning above
|
||||
lvs -o lv_name,data_percent,metadata_percent,lv_size pve/data
|
||||
# Example output (acerpve, 192.168.68.9):
|
||||
# LV Data% Meta% LSize
|
||||
# data 29.95 1.22 <816.21g
|
||||
# Required thresholds (documented minimum):
|
||||
# data_percent < 90% (80% recommended for safety margin)
|
||||
# metadata_percent < 70% (metadata fills faster than data)
|
||||
|
||||
# Check 2: Verify pool is not in error state (dmsetup shows the raw DM device)
|
||||
# Run on the NODE (e.g. ssh root@192.168.68.9 for acerpve) — NOT by bare hostname
|
||||
dmsetup status pve-data-tpool | grep -q "Error\|Fail" && exit 1
|
||||
# dmsetup status pve-data-tpool field order (verified on 192.168.68.9):
|
||||
# $1=start $2=length $3="thin-pool" $4=transaction-id
|
||||
# $5=metadata_used/metadata_total (blocks) $6=data_used/data_total (sectors)
|
||||
# remaining fields are flags ("-", "rw", "discard_passdown", "queue_if_no_space", ...)
|
||||
# This is only used for the ERROR-STATE check; use the lvs command above for percentages.
|
||||
# metadata_percent = $5 / ($5 split by /) [second number in pair]
|
||||
```
|
||||
|
||||
**Minimum thresholds**: If either `data_percent >= 90%` or `metadata_percent >= 70%`, the backup MUST NOT start. State explicitly that these are hard stops, not warnings.
|
||||
|
||||
**Why this is a precondition**: A full or errored thin pool fails EVERY volume on the VG at once, including the host root. This is not a soft failure - it takes down the entire Proxmox host.
|
||||
|
||||
#### Staging Directory Requirement (2026-09-14 incident)
|
||||
|
||||
Any custom vzdump `tmpdir` MUST be world-traversable and writable exactly like `/var/tmp` (mode 1777). The archive step of vzdump runs in an unprivileged user namespace and cannot traverse a root-owned 0700 directory.
|
||||
|
||||
**Symptom to recognize**: `fstat "<dir>/vzdumptmp<n>_<ct>//." failed - EACCES` on every container in the backup run.
|
||||
|
||||
**Fix**: `chmod 1777 <custom-tmpdir>` before starting vzdump.
|
||||
|
||||
#### Task Start Rule for Truncating Shells
|
||||
|
||||
When starting a backup task from a shell that may truncate output (e.g., pipes, `head`), always use:
|
||||
|
||||
```bash
|
||||
pvesh create /storage/backup --output-format json -- ... | head -2
|
||||
# ❌ Can kill the backup task ("broken pipe" status)
|
||||
```
|
||||
|
||||
Instead, capture JSON output without piping to truncating commands:
|
||||
|
||||
```bash
|
||||
# Use --output-format json and capture to variable
|
||||
result=$(pvesh create /storage/backup --output-format json -- ...)
|
||||
# Then parse result if needed
|
||||
```
|
||||
|
||||
#### GPU Host Backup Status (acerpve VM 101)
|
||||
|
||||
VM 101 (llm-gpu) and VM 103 (ocu-llm) have NO scheduled backup job. Coverage is manual one-off runs only. This is intentional for VM 101 until the thin-pool failure mechanism is understood and documented.
|
||||
|
||||
## Section 5: Network Services — Monitoring
|
||||
|
||||
### 5.1 Service Inventory
|
||||
@@ -555,7 +636,7 @@ monitor, or integration breaks.
|
||||
```bash
|
||||
# Full cluster status
|
||||
PVE="https://minipve.sysloggh.net"
|
||||
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
AUTH="Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
|
||||
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
|
||||
|
||||
# Docker health from Abiba
|
||||
@@ -579,9 +660,6 @@ curl -s -H "Authorization: Bearer $(infisical secrets get LITELLM_MASTER_KEY --p
|
||||
# Storage check
|
||||
ssh root@192.168.68.7 "df -h /media/storage /media/mediastore"
|
||||
|
||||
# Router roster reload (if needed)
|
||||
curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \
|
||||
-H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814"
|
||||
|
||||
# Restart stuck GPU (saturation watchdog alternative)
|
||||
ssh root@192.168.68.8 "systemctl restart llama-server"
|
||||
@@ -592,26 +670,26 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
||||
|
||||
| CT | Name | Node | IP | Role | Agent |
|
||||
|----|------|------|----|------|-------|
|
||||
| 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi |
|
||||
| 100 | abiba | minipve | .24 | Pi agent | ✅ pi |
|
||||
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
|
||||
| 102 | adguard | **minipve** | **.10** | DNS | ❌ |
|
||||
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
|
||||
| 104 | authentik | minipve | .11 | OIDC | ❌ |
|
||||
| 105 | kagentz | **hwepve** | — | Agent Zero | ✅ |
|
||||
| 105 | kagentz | amdpve | — | Agent Zero | ✅ |
|
||||
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
|
||||
| 107 | pbs | storepve | — | Backups | ❌ |
|
||||
| 108 | media | storepve | — | Media | ❌ |
|
||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
||||
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
||||
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
||||
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
|
||||
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
|
||||
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
||||
| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
|
||||
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
||||
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
||||
| 117 | zulip | storepve | .19 | Chat | ❌ |
|
||||
| 118 | jdownloader | storepve | .20 | JDownloader LXC (dedicated, migrated from docker-vm 2026-08-01) | ✅ |
|
||||
| 119 | infisical-vault | minipve | — | Vault | ❌ |
|
||||
| 120 | adguard2 | amdpve | — | DNS (secondary AdGuard) | ❌ |
|
||||
|
||||
## Appendix C: Docker Compose Files Location
|
||||
|
||||
@@ -631,20 +709,22 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
|
||||
|
||||
| CT | Name | Node | pct-run |
|
||||
|-----|------|------|---------|
|
||||
| 100 | abiba | hwepve | `pct-run 100` |
|
||||
| 105 | kagentz | hwepve | `pct-run 105` |
|
||||
| 111 | tdunna | amdpve | `pct-run 111` |
|
||||
| 100 | abiba | minipve | `pct-run 100` |
|
||||
| 105 | kagentz | amdpve | `pct-run 105` |
|
||||
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
|
||||
| 112 | tanko | amdpve | `pct-run 112` |
|
||||
| 113 | baggy | amdpve | `pct-run 113` |
|
||||
| 115 | scottdenya | amdpve | `pct-run 115` |
|
||||
| 104 | authentik | minipve | `pct-run 104` |
|
||||
| 110 | gitea | minipve | `pct-run 110` |
|
||||
| 114 | mumuni | hwepve | `pct-run 114` |
|
||||
| 116 | syslog-api | minipve | `pct-run 116` |
|
||||
| 106 | ra-h-os | storepve | `pct-run 106` |
|
||||
| 107 | proxmox-backup | storepve | `pct-run 107` |
|
||||
| 108 | media | storepve | `pct-run 108` |
|
||||
| 117 | zulip | storepve | `pct-run 117` |
|
||||
| 118 | jdownloader | storepve | `pct-run 118` |
|
||||
| 119 | infisical-vault | minipve | `pct-run 119` |
|
||||
| 120 | adguard2 | amdpve | `pct-run 120` |
|
||||
| 102 | adguard | **minipve** | `pct-run 102` |
|
||||
|
||||
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
|
||||
@@ -652,7 +732,7 @@ GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use S
|
||||
ssh root@192.168.68.8 # RTX 3090
|
||||
ssh root@192.168.68.110 # RTX 5070
|
||||
ssh root@192.168.68.15 # Strix Halo
|
||||
ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni)
|
||||
ssh root@192.168.68.4 # hwepve — standalone London node + NetBird routing peer (relocation pending)
|
||||
```
|
||||
|
||||
## Section 7: Agent Health Check (consolidated — 2026-07-05)
|
||||
@@ -679,4 +759,4 @@ which was kill+nohup outside systemd) are banned by policy.
|
||||
| Script | Why Disabled |
|
||||
|--------|-------------|
|
||||
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
|
||||
| `zulip-monitor.sh` (Abiba) | Replaced by agent-health-check.py + PM2 auto-restart |
|
||||
| `zulip-monitor.sh` (Abiba) | Standalone cron replaced by agent-health-check.py + PM2 auto-restart; the script itself remains active as the execution step of the `zulip-health` responsibility contract (see `zulip-health.prose.md`) |
|
||||
|
||||
@@ -140,7 +140,6 @@ After ALL updates (apt + images + restarts), verify every critical service is ba
|
||||
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
|
||||
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
|
||||
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
|
||||
| Hermes gateways | SSH to Mumuni CT 100, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
|
||||
|
||||
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
|
||||
|
||||
|
||||
@@ -17,7 +17,7 @@ description: >
|
||||
⚠️ This contract is target-state aspirational — but GPU export + alerting
|
||||
are now as-built (verified 2026-08-09).
|
||||
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
||||
version: 1.0.0
|
||||
version: 1.0.1
|
||||
---
|
||||
|
||||
## Architecture
|
||||
@@ -103,58 +103,202 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
||||
|
||||
## Execution
|
||||
|
||||
### Liveness rule (scoped)
|
||||
|
||||
The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints,
|
||||
where any HTTP answer proves a listener is up: the PVE API
|
||||
(`https://<node>:8006/api2/json/version`) and LiteLLM health
|
||||
(`/litellm/health`, `301` → `/litellm/health/liveliness`). For those endpoints a
|
||||
probe is **ALIVE** on **ANY** HTTP status — `401`/`403` auth challenges and `3xx`
|
||||
redirects included — and **DOWN = connection refused (`000`) or timeout only**.
|
||||
The PVE API legitimately answers `401` to an unauthenticated probe — that is the
|
||||
healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and
|
||||
gpu-monitor.
|
||||
|
||||
Probes whose success condition is specifically a bare `200` are NOT covered by
|
||||
the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
||||
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
|
||||
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
|
||||
|
||||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||
service answered — report the code, never "down". A redirect is not a failure.
|
||||
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe,
|
||||
and that is a statement about YOUR PROBE, not about the service.
|
||||
2. **A failed probe is never a service verdict.** Print
|
||||
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
||||
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
||||
then report. Apply the same shape as scripts/disk-gc-scan.py.
|
||||
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
||||
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
||||
(retried at 25s: also timeout)" is actionable.
|
||||
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
||||
tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
**EXECUTABLE OWNER:** The canonical probe set lives in `scripts/infra-monitoring.sh`.
|
||||
A check run is a single command: `bash scripts/infra-monitoring.sh` (from the
|
||||
repository root). Paste its raw output verbatim into the report. The script
|
||||
exits non-zero naming every failed target; there is no "OK" summary when any
|
||||
leg failed. Port drift is caught by `scripts/test_infra_monitoring.sh` which
|
||||
asserts every probed port matches the documented value.
|
||||
|
||||
**PROBE SHAPE (per standing rules above):**
|
||||
- Every probe prints: `✅ <name>: alive` on success, or `🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)` on failure
|
||||
- PVE API failures include `(any-HTTP liveness, -k for self-signed)` to distinguish TLS vs connection
|
||||
- Retry once on connection failure at longer timeout (25s connect, 30s max)
|
||||
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
||||
- Report the actual probe output, not a summary verdict
|
||||
|
||||
```bash
|
||||
# Provenance — run first; paste the absolute path into the report
|
||||
pwd -P
|
||||
|
||||
# ============================================================
|
||||
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
|
||||
# ============================================================
|
||||
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
|
||||
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
|
||||
if [ -z "$zulip_key" ]; then
|
||||
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
|
||||
else
|
||||
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
|
||||
echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
|
||||
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||
fi
|
||||
|
||||
# ============================================================
|
||||
# 2. PM2 PROCESS HEALTH
|
||||
# ============================================================
|
||||
pm2 jlist
|
||||
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
|
||||
# spoton-service removed 2026-09-14 (not in live set)
|
||||
|
||||
# ============================================================
|
||||
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
|
||||
# ============================================================
|
||||
# Probe /metrics (the Prometheus scrape target), not bare /
|
||||
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
|
||||
if [ "$code" == "000" ]; then
|
||||
# Retry with longer timeout
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
|
||||
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "GPU exporter http://$host:9400/metrics -> $code"
|
||||
fi
|
||||
done
|
||||
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
|
||||
|
||||
# ============================================================
|
||||
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
|
||||
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Router http://192.168.68.116/health -> $code"
|
||||
fi
|
||||
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||
|
||||
# ============================================================
|
||||
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
|
||||
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
|
||||
fi
|
||||
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
|
||||
|
||||
# ============================================================
|
||||
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
|
||||
# ============================================================
|
||||
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
|
||||
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
|
||||
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
|
||||
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "PVE API https://$node:8006/api2/json/version -> $code"
|
||||
fi
|
||||
done
|
||||
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
|
||||
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
|
||||
|
||||
# ============================================================
|
||||
# 7. PROMETHEUS TARGETS — bare-200 probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
|
||||
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
|
||||
fi
|
||||
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||
|
||||
# ============================================================
|
||||
# 8. GRAFANA HEALTH — any-HTTP probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
|
||||
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
|
||||
fi
|
||||
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
|
||||
|
||||
# ============================================================
|
||||
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
|
||||
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
|
||||
fi
|
||||
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the probe executed
|
||||
from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault
|
||||
at read time. For each probe, print the target name, the full URL, and the HTTP
|
||||
code (or failure kind with retry details). Apply the standing probe rules: any
|
||||
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
||||
a failure.
|
||||
|
||||
### Docker Stats and PVE Exporter Ports
|
||||
|
||||
These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:
|
||||
|
||||
| Exporter | Port | Container | Metrics |
|
||||
|----------|------|-----------|---------|
|
||||
| **Docker Stats** | **9324** | harness-docker-stats | `docker_container_*` (per-container CPU/mem/network) |
|
||||
| **PVE Exporter** | **9221** | harness-pve-exporter | `pve_*` (5 cluster-level metrics) |
|
||||
|
||||
**IMPORTANT**: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (`builder_builds_*`, `containerd_build_info_*`).
|
||||
|
||||
```bash
|
||||
# Docker Stats (harness-docker-stats)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
|
||||
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||
|
||||
# PVE Exporter (harness-pve-exporter)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
|
||||
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||
```
|
||||
|
||||
### Phase 1: GPU Exporters
|
||||
|
||||
**NVIDIA (.8 and .110)**:
|
||||
1. Download `nvidia_gpu_exporter` binary
|
||||
2. Create systemd service `nvidia-gpu-exporter.service`
|
||||
3. Start and enable
|
||||
|
||||
**AMD (.15)**:
|
||||
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
|
||||
2. Parses `amdgpu_top --json -d 1000` output
|
||||
3. Exposes key metrics at `:9400/metrics` via Python http.server
|
||||
4. Create systemd service
|
||||
5. Start and enable
|
||||
|
||||
### Phase 2: Prometheus
|
||||
|
||||
1. Create `/opt/monitoring/` directory on CT 116
|
||||
2. Write `prometheus.yml` with scrape configs for all targets
|
||||
3. Add to docker-compose (or separate compose file)
|
||||
4. Start container
|
||||
|
||||
### Phase 3: Grafana
|
||||
|
||||
1. Create `/opt/monitoring/grafana/` directories
|
||||
2. Provision Prometheus datasource
|
||||
3. Provision GPU fleet dashboard JSON
|
||||
4. Provision LiteLLM dashboard JSON
|
||||
5. Add to docker-compose
|
||||
6. Start container
|
||||
|
||||
### Phase 4: Verification
|
||||
|
||||
1. Verify all 3 GPU exporters return 200 at :9400/metrics
|
||||
2. Verify Prometheus targets all UP at :9090/targets
|
||||
3. Verify Grafana accessible at :3001 with dashboards
|
||||
4. Verify LiteLLM metrics flowing to Prometheus
|
||||
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
|
||||
|
||||
## Verification Commands
|
||||
|
||||
```bash
|
||||
# GPU exporters
|
||||
curl -s http://192.168.68.8:9400/metrics | grep nvidia
|
||||
curl -s http://192.168.68.110:9400/metrics | grep nvidia
|
||||
curl -s http://192.168.68.15:9400/metrics | grep amdgpu
|
||||
|
||||
# Prometheus
|
||||
curl -s http://192.168.68.116:9090/api/v1/targets
|
||||
|
||||
# Grafana
|
||||
curl -s http://192.168.68.116:3001/api/health
|
||||
|
||||
# LiteLLM metrics (already live)
|
||||
curl -s http://192.168.68.116:4001/metrics | head -20
|
||||
```
|
||||
|
||||
@@ -2,16 +2,17 @@
|
||||
kind: responsibility
|
||||
name: infrastructure-update
|
||||
description: >
|
||||
Autonomous system-wide update contract covering all 6 Proxmox nodes,
|
||||
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
|
||||
Autonomous system-wide update contract covering all 5 Proxmox nodes,
|
||||
15+ containers/VMs, and 5 Docker ecosystems (docker-vm .7, CT 116 .116,
|
||||
CT 117, hwpve .11, NetBird VPS 72.61.0.17). Updates apt packages,
|
||||
Docker images, and container stacks in safe waves with health checks
|
||||
and automatic rollback on failure.
|
||||
agent: abiba
|
||||
triggers:
|
||||
- on "infra update" command
|
||||
- weekly (Sunday 03:00 EDT) via cron
|
||||
- weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
|
||||
- on security advisory relay from Mumuni
|
||||
version: 1.2.0
|
||||
version: 1.4.0
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -56,11 +57,10 @@ Before ANY update wave:
|
||||
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 100 (mumuni/abiba, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
||||
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
||||
|
||||
@@ -81,7 +81,14 @@ Before ANY update wave:
|
||||
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
|
||||
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
|
||||
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
|
||||
| CT 117 (zulip, storepve) | Zulip | `docker pull zulip/docker-zulip:latest && docker restart zulip-zulip-1` |
|
||||
| CT 116 (.116, via minipve) | Trove docker agent (trove-agent-docker) | `pct exec 116 -- bash -c 'cd /opt/trove-agent && docker compose pull && docker compose up -d'` |
|
||||
| CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) |
|
||||
| CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) |
|
||||
| hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` |
|
||||
| NetBird VPS (72.61.0.17) | NetBird (server, dashboard, proxy, traefik, crowdsec) | `ssh root@72.61.0.17 'cd /root && docker compose pull && docker compose up -d'` |
|
||||
| VM 109 (.7) | Trove test | `cd /opt/trove-test && docker compose pull && docker compose up -d` |
|
||||
| VM 109 (.7) | docker-stats | `cd /opt/docker-stats && docker compose pull && docker compose up -d` |
|
||||
| CT 116 (.116) | Monitoring (Grafana, Prometheus, Alertmanager, PVE exporter) | `cd /opt/monitoring && docker compose pull && docker compose up -d` |
|
||||
|
||||
**Verify after Wave 3:**
|
||||
- All containers healthy: `docker ps` on each host
|
||||
@@ -89,8 +96,13 @@ Before ANY update wave:
|
||||
- MCP integration test: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools (23 RA-H OS + 67 GitHub)
|
||||
- Zulip test: send test message to #agent-hub
|
||||
- Dashboard loading: `curl localhost:3001/` (via CT 116)
|
||||
- Firecrawl test: `curl :3002/`
|
||||
- Firecrawl test: `curl -X POST http://192.168.68.7:3002/v1/search -H 'Content-Type: application/json' -d '{"query":"health","limit":1}'` → `"success":true` (GET `/` returns 200)
|
||||
- Authentik test: `curl http://192.168.68.11:9000/` → 302 redirect to login
|
||||
- NetBird test: `curl -s -o /dev/null -w '%{http_code}' https://netbird.sysloggh.net/` → 200
|
||||
- harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
|
||||
- SearXNG test: `curl :8888`
|
||||
- Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
|
||||
- Version-pin awareness: a fixed version tag (e.g. `image: ...litellm:1.99.1`) is a no-op for `docker compose pull` just like a digest pin, so the stack silently stops advancing. CT 116 `harness-litellm` is INTENTIONALLY pinned to `1.99.1` (registry `main-stable`/`latest` currently resolve to `1.100.1`, sha256:a3715fa7 — a bleeding-edge jump explicitly declined 2026-09-11). Every run must look up the newest STABLE release tag for any version-pinned image, bump the pin deliberately with user approval, recreate, and re-verify. Never silently revert a pin to a floating tag.
|
||||
|
||||
## Wave 4: Proxmox Kernel Reboot
|
||||
|
||||
@@ -159,13 +171,15 @@ Before Wave 1, snapshot these files:
|
||||
/opt/search-stack/searxng/docker-compose.yml (VM 109 .7)
|
||||
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
||||
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
||||
/root/compose.yml (hwpve .11 — Authentik server/worker/postgres)
|
||||
/root/docker-compose.yml (NetBird VPS — netbird server/dashboard/proxy, traefik, crowdsec)
|
||||
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
||||
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
|
||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||
# Hermes agent configs (key enforcement — 2026-07-10)
|
||||
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
|
||||
/root/.config/systemd/user/hermes-gateway.service (Mumuni CT 114 — EnvironmentFile fixed)
|
||||
/etc/environment (Mumuni CT 114 — LITELLM_API_KEY)
|
||||
/home/hermes/.hermes/config.yaml (Mumuni kagentz CT105; Tanko CT112 uses /home/jerome/.hermes)
|
||||
/etc/systemd/system/hermes-gateway.service (Mumuni kagentz CT105 — system unit, User=hermes)
|
||||
/etc/environment (LITELLM_API_KEY — legacy path, Mumuni now keys via Infisical)
|
||||
```
|
||||
|
||||
## MCP Gateway (2026-07-10)
|
||||
@@ -194,19 +208,19 @@ mcp_servers:
|
||||
| Key | MCP Access |
|
||||
|-----|-----------|
|
||||
| Master key | ✅ Full — 90 tools (vault-injected) |
|
||||
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.90.0-rc.1 |
|
||||
| Agent keys (mumuni, tanko, etc.) | ✅ Per-key grants supported (as of 2026-09-18 verification) |
|
||||
|
||||
### Known Limitations
|
||||
- Per-key MCP server grants not functional — only master key has access
|
||||
- ~~Per-key MCP server grants not functional — only master key has access~~ (resolved 2026-09-18: per-key grants now work)
|
||||
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
|
||||
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
|
||||
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
|
||||
|
||||
### Migration Path
|
||||
When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
||||
1. Grant agent keys `mcp_servers: ["ra_h_os"]`
|
||||
2. Update Hermes `mcp_servers.ra-h-os.url` from `http://192.168.68.65:3100/mcp` → `http://192.168.68.116:4000/mcp/`
|
||||
3. Add `headers: {x-litellm-api-key: "Bearer $LITELLM_API_KEY"}` to MCP config
|
||||
### Migration Path (COMPLETED 2026-09-18)
|
||||
Per-key MCP grants are now supported:
|
||||
1. ✅ Agent keys granted MCP access via `allowed_mcp_servers` field
|
||||
2. ✅ Hermes `mcp_servers.litellm.url` set to `https://litellm.sysloggh.net/mcp`
|
||||
3. ✅ `headers: {x-litellm-api-key: "Bearer <literal_key>"}` added to MCP config
|
||||
|
||||
## Security-Specific Updates
|
||||
|
||||
@@ -219,9 +233,9 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] All 6 PVE nodes updated, no reboot-loop
|
||||
- [ ] All 5 PVE nodes updated, no reboot-loop
|
||||
- [ ] All VMs/CTs running post-update
|
||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
|
||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117 + hwpve .11 + NetBird VPS)
|
||||
- [ ] LiteLLM inference passing (syslog-auto test)
|
||||
- [ ] Zulip server + all 3 agents connected
|
||||
- [ ] GPU fleet at full capacity (3/3)
|
||||
@@ -235,7 +249,7 @@ After completion, send Zulip DM:
|
||||
```
|
||||
📋 Infrastructure Update — YYYY-MM-DD
|
||||
|
||||
Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers
|
||||
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers
|
||||
Security fixes: N CVEs patched
|
||||
Downtime: <service> <duration>
|
||||
Failures: none / <details>
|
||||
|
||||
+126
-10
@@ -67,8 +67,14 @@ description: >
|
||||
4. **If action == "create"**:
|
||||
- Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date)
|
||||
- Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" }
|
||||
- Duration is null (permanent) — inherited from litellm default_key_generate_params
|
||||
- Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"]
|
||||
- Duration is whatever the caller passes; NO default enforcement exists today (CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key with no explicit models returns an empty models list). Agent keys are permanent by policy, not by that block. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
|
||||
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
||||
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
||||
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
||||
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
|
||||
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
|
||||
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
|
||||
In LiteLLM Community both silently grant access to EVERY model, including cloud.
|
||||
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
||||
- Return the new key
|
||||
5. **If action == "rotate"**:
|
||||
@@ -85,6 +91,66 @@ description: >
|
||||
- Confirm key alias matches agent_name in LiteLLM key list
|
||||
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
||||
|
||||
## Cloud Provider Consolidation (2026-09-20)
|
||||
|
||||
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
|
||||
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
|
||||
|
||||
### Provider map (per-account namespacing)
|
||||
|
||||
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
|
||||
stay separate:
|
||||
|
||||
| Prefix | Upstream | Auth | Vault secret |
|
||||
|--------|----------|------|--------------|
|
||||
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
|
||||
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
|
||||
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
|
||||
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
|
||||
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
|
||||
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
|
||||
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
|
||||
|
||||
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
|
||||
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
|
||||
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
|
||||
|
||||
### Access tiers (MUST be enforced per key)
|
||||
|
||||
| Tier | Model names | Granted to |
|
||||
|------|-------------|------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
|
||||
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
|
||||
|
||||
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
|
||||
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
|
||||
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
|
||||
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
|
||||
|
||||
### Creating the cloud-enabled key
|
||||
|
||||
```bash
|
||||
# ALWAYS read the live roster first (key-scoped):
|
||||
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
|
||||
| jq -r '.data[].id'
|
||||
|
||||
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
|
||||
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
|
||||
```
|
||||
|
||||
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
|
||||
any `<prefix>/` cloud model.
|
||||
|
||||
### Adding a new cloud provider
|
||||
|
||||
1. Add the upstream key to Infisical `infrastructure/production/root`.
|
||||
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
|
||||
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
|
||||
3. Restart the `harness-litellm` container.
|
||||
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
|
||||
captain approval.
|
||||
5. Update this table and the access-tier section.
|
||||
|
||||
## Production Vault Access Process (canonical, 2026-07-17)
|
||||
|
||||
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
||||
@@ -142,8 +208,8 @@ through its agent wrapper.
|
||||
safety net for vault outage or token revocation. Must be kept in sync on rotation.
|
||||
Example:
|
||||
```bash
|
||||
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
|
||||
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
|
||||
MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY»
|
||||
MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY»
|
||||
```
|
||||
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
|
||||
```ini
|
||||
@@ -178,7 +244,7 @@ through its agent wrapper.
|
||||
| Agent | Host | Pattern | Keys | Status |
|
||||
|-------|------|---------|------|--------|
|
||||
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
|
||||
| mumuni | .24 (CT100 abiba) | Pi Hermes gateway (no systemd) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||
| mumuni | .14 (kagentz CT105) | systemd unit hermes-gateway.service (user hermes) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
|
||||
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
|
||||
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
|
||||
@@ -189,7 +255,7 @@ through its agent wrapper.
|
||||
### Tanko migration (COMPLETED 2026-07-17)
|
||||
|
||||
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
|
||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
|
||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-synthetic-tanko-example…`)
|
||||
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
|
||||
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
|
||||
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
|
||||
@@ -257,9 +323,51 @@ reads use per-agent identities. This eliminates the single shared token risk.
|
||||
| Agent | .env Keys |
|
||||
|-------|-----------|
|
||||
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
|
||||
| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|
||||
| Koby | (wrapper injects from vault — .env has Telegram token) |
|
||||
| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|
||||
|| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|
||||
|| Koby | (wrapper injects from vault — .env has Telegram token) |
|
||||
|| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|
||||
|| Agent Zero (kagentz .14) | OPENROUTER_API_KEY (direct OpenRouter access) |
|
||||
|
||||
### Agent Zero (kagentz .14) — OpenRouter Integration (2026-09-01)
|
||||
|
||||
Agent Zero runs in Docker on kagentz (CT105) and uses **direct OpenRouter API access**,
|
||||
not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update manager,
|
||||
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
||||
|
||||
**Key Storage:**
|
||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»…`)
|
||||
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
||||
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
||||
(unlike fleet agents which require vault injection)
|
||||
|
||||
**Current Key (2026-09-01):**
|
||||
- **Prefix**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
- **Plan**: Paid (not free tier)
|
||||
- **Usage**: 0 (as of 2026-09-01)
|
||||
|
||||
**Model Configuration:**
|
||||
- **Preset**: "Cost Efficient" (`/a0/usr/plugins/_model_config/presets.yaml`)
|
||||
- **Model**: `openrouter/moonshotai/kimi-k3`
|
||||
- **API Base**: (empty — uses OpenRouter default)
|
||||
|
||||
**Why not LiteLLM proxy?**
|
||||
Agent Zero's architecture was designed before the fleet adopted the LiteLLM proxy
|
||||
standard. The container runs `/exe/self_update_manager.py` and `/a0/run_ui.py` which
|
||||
directly call OpenRouter via Python's requests library. Converting would require:
|
||||
1. Refactoring all LLM calls to use `litellm` library
|
||||
2. Adding vault wrapper injection
|
||||
3. Updating self_update_manager to use proxy-aware key handling
|
||||
|
||||
**Rotation Procedure:**
|
||||
1. Generate new key in OpenRouter UI
|
||||
2. Update container: `sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env`
|
||||
3. Update vault: `infisical secrets set OPENROUTER_API_KEY=<new_key> --projectId=agents --env=production`
|
||||
4. Restart container: `sudo docker exec agent-zero supervisorctl restart run_ui`
|
||||
5. Verify: `curl -s https://openrouter.ai/api/v1/auth/key -H "Authorization: Bearer <new_key>"`
|
||||
|
||||
**Related Contract:**
|
||||
- `agent-zero-openrouter-key.prose.md` — Full agent-zero key management contract
|
||||
|
||||
## Key Rotation Log
|
||||
|
||||
@@ -276,7 +384,15 @@ reads use per-agent identities. This eliminates the single shared token risk.
|
||||
|
||||
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
|
||||
|
||||
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
|
||||
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
|
||||
```bash
|
||||
# PRIMARY (proven, runs on CT 116 with no extra tooling):
|
||||
docker exec harness-litellm printenv LITELLM_MASTER_KEY
|
||||
# Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching)
|
||||
# The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it)
|
||||
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
|
||||
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
|
||||
```
|
||||
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
|
||||
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
|
||||
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
|
||||
|
||||
@@ -0,0 +1,115 @@
|
||||
---
|
||||
kind: pattern
|
||||
name: litellm-client-timeouts
|
||||
description: >
|
||||
Standard client timeout and retry policy for ALL agents calling LiteLLM
|
||||
(CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident:
|
||||
a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures
|
||||
(82% from abiba-pi, 13 from mumuni) against a backend that was actually
|
||||
succeeding at 20-70s per call once clients stopped giving up. Grounded in
|
||||
measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes
|
||||
to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h;
|
||||
nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely
|
||||
client-side. Blast radius if wrong: agents fall back to DeepSeek silently
|
||||
(key/timeout failures present as model degradation, not errors) or abandon
|
||||
healthy-but-slow reasoning calls, fragmenting long tasks.
|
||||
---
|
||||
|
||||
## Maintains
|
||||
|
||||
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
|
||||
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
|
||||
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
|
||||
|
||||
## The measured numbers these values come from
|
||||
|
||||
| Model | avg latency | avg TTFT | p-profile (24h) |
|
||||
|---|---|---|---|
|
||||
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
||||
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
||||
| gpu-vision (retired gemma-4-12b, RTX 5070) | 2.6s | — | RTX 5070, healthy |
|
||||
|
||||
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
||||
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
||||
tolerated 20-70s — 392 successful slow calls in the three hours after recovery.
|
||||
LiteLLM's internal queue time is ~0s; the latency is model inference, not
|
||||
proxy queuing.
|
||||
|
||||
## Parameters
|
||||
|
||||
### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
|
||||
|
||||
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy
|
||||
28.8s average + 120-300s tail. Every 408 in the incident was a client
|
||||
abandoning a request the backend would have answered.
|
||||
- If the transport exposes a timeout setting for the main model, set it to
|
||||
**300s or more**. If it does not (current Hermes custom-provider path has no
|
||||
timeout knob), that is acceptable ONLY because nginx holds the request for
|
||||
600s — but any wrapper, script, or direct API call you write MUST set its own
|
||||
timeout >= 300s for syslog-auto/qwen-class calls.
|
||||
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast
|
||||
here is what caused the incident.
|
||||
|
||||
### 2. Auxiliary tasks — keep template timeouts, one correction
|
||||
|
||||
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on `gemma-4-12b` (retired 2026-09-12); the live RTX 5070 alias is `gpu-vision`.
|
||||
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
||||
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
||||
(gpu-dense backend) is the same speed class as syslog-auto; delegation
|
||||
defaults that assume fast responses will 408 the same way.
|
||||
|
||||
### 3. Retry policy — backoff, not repetition
|
||||
|
||||
- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff
|
||||
(**15s, 45s**) before giving up.
|
||||
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour
|
||||
from one key) was a batch job retrying without backoff while the backend was
|
||||
down — it multiplied load during recovery.
|
||||
- On 401/403: do NOT retry — that is a key/permission problem (see
|
||||
litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
|
||||
- On 429: honor the retry-after header if present, else back off 60s.
|
||||
|
||||
### 4. Health probes — identify yourself and time out sanely
|
||||
|
||||
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12
|
||||
such orphans appeared in the incident window and cost investigation time).
|
||||
Send a real model name and use a real (probe-designated) key.
|
||||
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
|
||||
report "backend slow (>30s)" rather than hanging.
|
||||
- Probe cadence: at most hourly. The litellm-health cron cadence (authoritative
|
||||
trigger in contract-registry.yaml) is the standard; sub-hourly synthetic
|
||||
traffic distorts latency baselines.
|
||||
|
||||
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
|
||||
|
||||
- The incident window showed bulk clients amplifying a backend stall 5:1.
|
||||
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
|
||||
- Any batch loop over N requests MUST sleep >= 5s between requests and honor
|
||||
the retry policy in section 3.
|
||||
|
||||
## Returns
|
||||
|
||||
- A single standard any agent or script can cite: timeouts >= 300s on the
|
||||
syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x,
|
||||
15s/45s), identified probes at 30s/hourly, batch jobs chunked and
|
||||
day-scheduled.
|
||||
- Failure signature recognition: bulk 408s from multiple keys in one window =
|
||||
backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated);
|
||||
single-key 408s = that client's timeout is too short.
|
||||
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s,
|
||||
auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent,
|
||||
stable aliases), litellm-api-keys.prose.md (key/permission failures).
|
||||
|
||||
## Intentionally NOT changed
|
||||
|
||||
- No server-side LiteLLM timeout/cooldown changes proposed — the incident
|
||||
self-recovered and the server is healthy (0.56s live probe); changing
|
||||
server behavior without process-level root cause (CT116 requires root;
|
||||
not reachable from kagentz) would be guessing.
|
||||
- No change to the template's vision/web_extract/compression timeouts —
|
||||
measured data says they are correct.
|
||||
- No per-agent key permission changes — those are litellm-api-keys.prose.md
|
||||
territory (and the open gpu-vision/gemma 403 items are already filed with
|
||||
the key owners).
|
||||
- No model routing changes — syslog-auto's weighted pool behaved correctly
|
||||
throughout the incident.
|
||||
+100
-47
@@ -1,19 +1,12 @@
|
||||
---
|
||||
kind: function
|
||||
name: litellm-health
|
||||
status: deprecated
|
||||
deprecated_on: 2026-07-09
|
||||
replaced_by: litellm-self-heal.prose.md
|
||||
note: >
|
||||
Consolidated into litellm-self-heal.prose.md to eliminate duplication
|
||||
of architecture diagrams, GPU topology, timeout tables, and container
|
||||
lists. Health check is now § Health Check within litellm-self-heal.
|
||||
This file is retained for reference only — use litellm-self-heal instead.
|
||||
status: active
|
||||
description: >
|
||||
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
|
||||
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
|
||||
Router (harness-router :9000) is DEPRECATED — container still runs but
|
||||
is not in the request path. GPU monitoring via Prometheus/Grafana and
|
||||
Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container,
|
||||
image and config removed). GPU monitoring via Prometheus/Grafana and
|
||||
fleet dashboard (gpu-monitor :9100).
|
||||
Designed as a reusable contract for any Syslog agent.
|
||||
|
||||
@@ -26,7 +19,7 @@ description: >
|
||||
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
|
||||
- Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).
|
||||
---
|
||||
|
||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||
@@ -42,8 +35,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
│
|
||||
Grafana :3001
|
||||
|
||||
harness-router :9000 — DEPRECATED, container still runs but
|
||||
NOT in request path. nginx routes /v1 → LiteLLM directly.
|
||||
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
|
||||
and config removed). nginx routes /v1 → LiteLLM directly.
|
||||
Router slot booking + circuit breakers replaced by
|
||||
LiteLLM native fallbacks + timeouts.
|
||||
```
|
||||
@@ -51,9 +44,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
||||
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
||||
- All GPUs at parallel 2 (was parallel 1)
|
||||
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
|
||||
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
|
||||
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
|
||||
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
|
||||
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
|
||||
|
||||
## Parameters
|
||||
|
||||
@@ -75,33 +67,42 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
- SSH key access to backend_host for container checks
|
||||
- Network access to public_url, auth_host, and gpu_dashboard_url
|
||||
- LiteLLM master key for key management endpoints
|
||||
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
|
||||
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
|
||||
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
|
||||
`strix-moe`, `syslog-auto`). Retrieve from the executor's host via:
|
||||
|
||||
```
|
||||
monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
|
||||
master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
|
||||
```
|
||||
|
||||
If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys").
|
||||
The master key must never be used for inference.
|
||||
|
||||
## GPU Fleet Topology
|
||||
|
||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||
|------|-----|----------|---------------|--------|---------|----------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||
| Host | IP | Hardware | Role |
|
||||
|------|-----|----------|------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
||||
|
||||
## Model Fallback Chains (LiteLLM)
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
||||
`crew-auto`).
|
||||
|
||||
| Primary | Timeout | Fallback | Timeout |
|
||||
|---------|---------|----------|---------|
|
||||
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
|
||||
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
|
||||
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
|
||||
| syslog-auto (balanced) | 300s | qwen → gemma | — |
|
||||
|
||||
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
|
||||
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
|
||||
> per-model timeouts in this contract is intentionally superseded — that state is config,
|
||||
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
|
||||
> the dedicated `monitor` key, not the master key, which is admin-only.
|
||||
|
||||
## Containers on CT 116
|
||||
|
||||
| Container | Image | Port | Health Check |
|
||||
|-----------|-------|------|-------------|
|
||||
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness |
|
||||
| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) |
|
||||
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
|
||||
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
||||
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
||||
| harness-redis | redis:7-alpine | :6379 | PING |
|
||||
@@ -113,38 +114,90 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
1. **Read parameters** — Use provided values or defaults
|
||||
|
||||
2. **Check public endpoints**:
|
||||
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
|
||||
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
|
||||
the surface it targets. Never point a check at a path that only resolves on the other
|
||||
surface.
|
||||
|
||||
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
|
||||
ROOT; the `/litellm/` prefix does not exist there and 404s:
|
||||
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
|
||||
|
||||
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
|
||||
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
|
||||
added 2026-09-11)
|
||||
|
||||
3. **Check LiteLLM health (no-auth)**:
|
||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||
|
||||
4. **Check backend container health**:
|
||||
- SSH to {{backend_host}} → `docker ps` → verify 8 containers healthy
|
||||
- Critical: harness-litellm, harness-router, harness-nginx, harness-postgres
|
||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus
|
||||
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
|
||||
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
|
||||
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
|
||||
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
|
||||
|
||||
5. **Check router roster loaded**:
|
||||
- GET http://{{backend_host}}:9000/health → expect 200
|
||||
- GET http://{{backend_host}}:9000/health/unified → expect 3 models
|
||||
- If router returns "all GPUs saturated" but GPUs idle: roster not loaded → reload
|
||||
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
|
||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||
- nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload)
|
||||
|
||||
6. **Check GPU fleet health** (via fleet dashboard):
|
||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||
- Verify GPUs reporting status "healthy"
|
||||
- Check alerts array for active warnings/critical
|
||||
|
||||
7. **Check model inference via LiteLLM** — Test each model:
|
||||
- POST /v1/chat/completions model=gemma-4-12b → expect 200
|
||||
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
|
||||
- POST /v1/chat/completions model=strix-moe → expect 200
|
||||
- Use master key for auth
|
||||
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
|
||||
check runs on the **backend edge**, not the public edge, so these paths carry the
|
||||
`/litellm/` prefix:
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
|
||||
- Auth uses the dedicated `monitor` agent key. Retrieve via:
|
||||
`ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"`
|
||||
Do NOT use the master key for inference — the master key is for admin endpoints only
|
||||
(`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via:
|
||||
`ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"`
|
||||
- KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`,
|
||||
`gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe
|
||||
returns 403 and the host is not covered.
|
||||
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
|
||||
alias) and re-run — never drop the host from the probe to make the check pass.
|
||||
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
|
||||
this list. The RTX 5070 host now serves `gpu-vision`.
|
||||
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
|
||||
the set depends on the key. Always state which key a model list was read with — a
|
||||
snapshot without its key is not evidence. This probe uses the `monitor` key on the
|
||||
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
|
||||
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
|
||||
than freezing a list here.
|
||||
|
||||
8. **Check agent keys**:
|
||||
- GET /key/list with master key → verify all 6 agents have keys
|
||||
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
|
||||
- **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use:
|
||||
`ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"`
|
||||
- If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys")
|
||||
|
||||
9. **Check Grafana**:
|
||||
- GET {{grafana_url}}/api/health → expect 200
|
||||
|
||||
10. **Compile and report** — Determine overall_status from individual check results
|
||||
|
||||
## Executor Script (2026-09-13)
|
||||
|
||||
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
|
||||
checks defined above and reports results in a standardized format. Paste its output in
|
||||
the status line.
|
||||
|
||||
- Hand-rolled probes are **not** an acceptable substitute for the script.
|
||||
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
|
||||
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
|
||||
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
|
||||
because the `harness-docker-stats` container binds to localhost on CT 116.
|
||||
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
|
||||
in the remote curl command with proper quoting.
|
||||
|
||||
Expected output on a healthy fleet: 11/11 passing checks.
|
||||
|
||||
+57
-75
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: responsibility
|
||||
name: litellm-self-heal
|
||||
status: deployed
|
||||
@@ -9,21 +11,22 @@ note: >
|
||||
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
|
||||
Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
|
||||
GPU monitoring integrated from gpu-monitor on .24:9100.
|
||||
|
||||
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
|
||||
duplication of architecture diagrams, GPU topology, timeout tables, and container
|
||||
lists. Health check is now § Health Check within this contract.
|
||||
|
||||
|
||||
Health probes are owned by litellm-health.prose.md (dispatched as
|
||||
`run contract: litellm-health`). This contract owns remediation only — it does not
|
||||
re-specify the probes.
|
||||
|
||||
Source of truth for GPU topology and keys: gpu-fleet.prose.md
|
||||
Last verified: 2026-07-12
|
||||
description: >
|
||||
LiteLLM inference stack health monitoring + self-healing. Verifies the full
|
||||
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
|
||||
inference, and agent keys. Applies remediation rules for common failures.
|
||||
LiteLLM inference stack remediation. Applies remediation rules for failures detected
|
||||
by litellm-health.prose.md (nginx → LiteLLM → GPU chain, CT 116 containers, GPU hosts,
|
||||
model inference, and agent keys).
|
||||
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
||||
---
|
||||
---
|
||||
|
||||
# LiteLLM Operations — Health Check + Self-Heal
|
||||
# LiteLLM Operations — Self-Heal (Remediation)
|
||||
|
||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||
|
||||
@@ -38,8 +41,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
│
|
||||
Grafana :3001
|
||||
|
||||
harness-router :9000 — DEPRECATED, container still runs but
|
||||
NOT in request path. nginx routes /v1 → LiteLLM directly.
|
||||
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
|
||||
and config removed). nginx routes /v1 → LiteLLM directly.
|
||||
Router slot booking + circuit breakers replaced by
|
||||
LiteLLM native fallbacks + timeouts.
|
||||
```
|
||||
@@ -55,40 +58,50 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
## GPU Fleet Topology
|
||||
|
||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||
|------|-----|----------|---------------|--------|---------|----------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||
| Host | IP | Hardware | Role |
|
||||
|------|-----|----------|------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
||||
|
||||
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
||||
> Verified on the ground 2026-07-16. The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
||||
|
||||
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
||||
## LiteLLM Model Surface
|
||||
|
||||
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`.
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
||||
`crew-auto`).
|
||||
|
||||
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
||||
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
||||
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them.
|
||||
### Alias Surface
|
||||
|
||||
| Alias | Serves | Where | Kind |
|
||||
|-------|--------|-------|------|
|
||||
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
|
||||
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
|
||||
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
|
||||
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
|
||||
|
||||
### Context Cap Split (2026-08-20; crew cap RETIRED)
|
||||
|
||||
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
|
||||
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
|
||||
- **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force.
|
||||
|
||||
> The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change.
|
||||
|
||||
- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
|
||||
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
|
||||
|
||||
## Model Fallback Chains (LiteLLM)
|
||||
|
||||
| Primary | Timeout | Fallback | Timeout |
|
||||
|---------|---------|----------|---------|
|
||||
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
|
||||
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
|
||||
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
|
||||
| syslog-auto (balanced) | 300s | qwen → gemma | — |
|
||||
|
||||
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
|
||||
> Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed
|
||||
> by the single-source-of-truth re-scope; read those values from the CT 116 config named
|
||||
> above rather than from this contract.
|
||||
|
||||
## Containers on CT 116
|
||||
|
||||
| Container | Image | Port | Health Check |
|
||||
|-----------|-------|------|-------------|
|
||||
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness |
|
||||
| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) |
|
||||
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
|
||||
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
||||
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
||||
| harness-redis | redis:7-alpine | :6379 | PING |
|
||||
@@ -102,8 +115,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (hardcoded key → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM (deprecated key, no live usage).
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -120,48 +133,14 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
- Also wakes on user request
|
||||
- On failure: re-check after 30s, escalate after 3 consecutive failures
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Health Check
|
||||
|
||||
Run this first on every cycle. Results feed into remediation rules below.
|
||||
|
||||
### 1. Check public endpoints
|
||||
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
|
||||
### 2. Check LiteLLM health (no-auth)
|
||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||
|
||||
### 3. Check backend container health
|
||||
- SSH to {{backend_host}} → `docker ps` → verify 10 containers healthy
|
||||
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-docker-stats, harness-pve-exporter
|
||||
- Deprecated but running: harness-router (not in path, reference only)
|
||||
|
||||
### 4. Check GPU fleet health (via fleet dashboard)
|
||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||
- Verify GPUs reporting status "healthy"
|
||||
- Check alerts array for active warnings/critical
|
||||
|
||||
### 5. Check model inference via LiteLLM — test each model
|
||||
- POST /v1/chat/completions model=gemma-4-12b → expect 200
|
||||
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
|
||||
- POST /v1/chat/completions model=strix-moe → expect 200
|
||||
- Use master key for auth
|
||||
|
||||
### 6. Check agent keys
|
||||
- GET /key/list with master key → verify all 6 agents have keys
|
||||
|
||||
### 7. Check Grafana
|
||||
- GET {{grafana_url}}/api/health → expect 200
|
||||
|
||||
### 8. Compile overall status
|
||||
Determine overall_status from individual check results:
|
||||
- "healthy" — all checks pass
|
||||
- "degraded" — 1-2 non-critical checks fail
|
||||
- "down" — critical checks fail
|
||||
Health probes are owned by `litellm-health.prose.md` (dispatched as `run contract: litellm-health`). This contract owns remediation only — it does not re-specify the probes.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Remediation Rules
|
||||
@@ -202,9 +181,11 @@ Fix → generate keys in LiteLLM via /key/generate → update /etc/environment o
|
||||
Escalate → if SSH access unavailable, send Zulip DM
|
||||
|
||||
### Rule 9: Stale Active Counter in Redis — DEPRECATED
|
||||
Router no longer in path so Redis active counters are unused. Rule retained
|
||||
for reference but inactive. If Redis issues occur, check harness-redis container.
|
||||
Router no longer in path so router active-slot counters are unused. Rule retained
|
||||
for reference but inactive. `harness-redis` now serves only LiteLLM cache and
|
||||
rate-limit state; check the container if cache errors appear.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Reporting
|
||||
@@ -227,6 +208,7 @@ top actions, uptime.
|
||||
If a fix requires another agent (e.g., Authentik restart), relay sent
|
||||
to responsible agent with full context.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Execution
|
||||
|
||||
@@ -1,9 +1,12 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
name: memory-audit-maintenance
|
||||
kind: responsibility
|
||||
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
||||
description: Shared memory audit and maintenance contract for Hermes agents (Mumuni, Koby, Koonimo). Tanko is no longer a Hermes agent (now on DSH/DeepSeek Harness since 2026-08-27) and uses DSH-native memory, so it is excluded from this Hermes roster. Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
|
||||
id: 067NC4KG01RG50R40M30E20918
|
||||
---
|
||||
---
|
||||
|
||||
### Goal
|
||||
|
||||
@@ -15,11 +18,10 @@ Autonomously audit and reorganize an agent's native memory (MEMORY.md, USER.md,
|
||||
|
||||
### Scope
|
||||
|
||||
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
||||
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across Hermes agents (Mumuni, Koby, Koonimo). **Tanko is excluded — it migrated to DSH (DeepSeek Harness) on 2026-08-27 and now uses DSH-native memory (mnemon), not `~/.hermes/memories/`.** Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
|
||||
|
||||
**Agent Roster:**
|
||||
**Agent Roster (Hermes):**
|
||||
- Mumuni
|
||||
- Tanko
|
||||
- Koby (CT 111 / tdunna)
|
||||
- Koonimo (CT 113 / baggy)
|
||||
|
||||
@@ -343,4 +345,3 @@ return {
|
||||
|
||||
### Per-Agent Notes
|
||||
|
||||
Each Hermes agent (Mumuni, Tanko, Tdunna, Baggy) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory.
|
||||
+31
-6
@@ -1,10 +1,13 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: pattern
|
||||
name: memory-fixer
|
||||
description: >
|
||||
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
||||
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||
version: 2.0.0
|
||||
version: 2.1.0
|
||||
---
|
||||
---
|
||||
|
||||
# Memory Fixer
|
||||
@@ -66,13 +69,15 @@ FROM nodes
|
||||
WHERE json_extract(metadata, '$.namespace') IS NULL;
|
||||
```
|
||||
|
||||
### 3. Staleness Review Tagging
|
||||
### 3. Staleness Review Tagging (refresh-suggested nodes only)
|
||||
|
||||
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
|
||||
|
||||
**Archive-suggested nodes are NO LONGER tagged — they are archived outright (see Level 1 fix 4).** Tagging with `[REVIEW: refresh]` applies only to living nodes (infrastructure, deployment, system, system-health, business, philosophy, research, learning, investigation, analysis, project, agent, registry, policy).
|
||||
|
||||
**Exclusion Rules:**
|
||||
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
|
||||
- Nodes whose `description` already starts with `[REVIEW:` are NOT re-processed
|
||||
- Nodes whose `description` already starts with `[REVIEW:` or `[ARCHIVED]` are NOT re-processed
|
||||
|
||||
```sql
|
||||
SELECT id, title, json_extract(metadata, '$.type') as node_type,
|
||||
@@ -101,9 +106,28 @@ LIMIT 10;
|
||||
|
||||
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`.
|
||||
|
||||
### 4. Stale-Node Archiving (Level 1 — standing Kwame directive, 2026-09-11)
|
||||
|
||||
**Kwame's standing directive: stale nodes CAN be archived by the fixer. No per-batch escalation, no `[REVIEW: archive]` tagging — archive them.**
|
||||
|
||||
For every node whose suggested action is `archive` (i.e. its type is NOT one of the living types in fix 3), archive it in a **single** `updateNode` call:
|
||||
|
||||
```python
|
||||
updateNode(id, {
|
||||
"description": "[ARCHIVED] " + originalDescriptionWithoutReviewTag,
|
||||
"metadata": {"state": "archived"}
|
||||
})
|
||||
```
|
||||
|
||||
- `state` transitions **DO work through `updateNode`** (`archived`, and back to `active`). The former "state only accepts processed/not_processed, use SSH" claim was wrong — verified 2026-09-11 by archiving 7 nodes (#61, #373, #388, #465, #475, #526, #1476) over the bridge with `updated_at` auto-bumping. **SSH to the bridge host is a fallback, not a requirement**, and it is blocked from kagentz anyway.
|
||||
- Pass `description` and `metadata` in the **same** call, and always keep the `updates` object nested: `{"id": N, "updates": {…}}`.
|
||||
- Archiving is non-destructive: the node stays in the graph, marked `state: archived` + `[ARCHIVED] ` prefix. **Living nodes (refresh-suggested) are NEVER archived** without a specific Kwame decision — they are the cluster/agent/business canon.
|
||||
|
||||
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
|
||||
|
||||
## Level 2 Escalations (Kwame Decision Required)
|
||||
|
||||
1. **Stale nodes** flagged with `[REVIEW: …]` — Archive, refresh, or keep?
|
||||
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
|
||||
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
|
||||
3. **Orphan Nodes >90 days old** — Archive or connect?
|
||||
|
||||
@@ -141,7 +165,7 @@ Reply with:
|
||||
|
||||
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
|
||||
|
||||
> ⚠️ `updateNode` cannot set `state` to non-standard values (restricted to `processed`/`not_processed`) and cannot add metadata keys. For state transitions and `updated_at` bumps, use **direct SSH + SQLite** on the bridge host:
|
||||
> ⚠️ **Corrected 2026-09-11:** `updateNode` DOES accept `state` changes — `{"updates": {"description": …, "metadata": {"state": "archived"}}}` works over the bridge, and `updated_at` bumps automatically. The old "use direct SSH + SQLite for state transitions" instruction was based on a wrong assumption; SSH is a fallback only (and is blocked from kagentz). Use one `updateNode` call for both the tag and the state.
|
||||
> ```bash
|
||||
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
|
||||
> ```
|
||||
@@ -169,8 +193,9 @@ The result must be 0 rows when all decisions are executed. Report what was done.
|
||||
## Checks
|
||||
|
||||
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
||||
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
|
||||
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
||||
- **Timestamps:** every executed decision bumps `updated_at`, so the node exits the stale window on the next run.
|
||||
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
|
||||
|
||||
## Logging
|
||||
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
|
||||
|
||||
@@ -6,7 +6,7 @@ description: >
|
||||
delegation, verification, and delivery. Defines when to delegate, which
|
||||
worker to use for what, how to handle failures, and the kanban board
|
||||
protocol. Enforces context-window discipline and separation of concerns.
|
||||
Runs on Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent (Pi + Hermes Zulip gateway).
|
||||
Runs on Mumuni (kagentz CT105, minipve, .14) via Hermes agent (Zulip gateway via systemd).
|
||||
version: 1.0.0
|
||||
---
|
||||
|
||||
@@ -19,8 +19,8 @@ version: 1.0.0
|
||||
|
||||
## Topology
|
||||
|
||||
**Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve)
|
||||
**Manager:** Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent
|
||||
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
|
||||
**Manager:** Mumuni (kagentz CT105, minipve, .14) via Hermes agent
|
||||
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
|
||||
|
||||
This contract is infrastructure-agnostic in terms of which nodes are used.
|
||||
@@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6,
|
||||
## Why This Matters
|
||||
|
||||
Without enforced delegation, the manager consumes the full iteration budget
|
||||
(60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs —
|
||||
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
|
||||
leaving no capacity for actual coordination. The result: context overflow
|
||||
(59K tokens in system prompt), iteration exhaustion, and degraded response
|
||||
quality. This contract exists because I blew through my budget checking
|
||||
@@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own.
|
||||
|
||||
**This is a hard rule, not a recommendation.** Violating it produces the exact
|
||||
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
|
||||
"6 nodes present" in the raw data but "6/6 online" in the report — even though
|
||||
"5 nodes present" in the raw data but "5/5 online" in the report — even though
|
||||
one of those nodes was unreachable. The report lied because it used data the
|
||||
raw data never provided.
|
||||
|
||||
@@ -90,8 +90,8 @@ raw data never provided.
|
||||
|
||||
| Worker | Model | Toolsets | Role | Use When |
|
||||
|--------|-------|----------|------|----------|
|
||||
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||
| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||
| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
||||
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
||||
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
||||
@@ -137,7 +137,7 @@ delegate_task(
|
||||
```
|
||||
delegate_task(
|
||||
tasks=[
|
||||
{"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
||||
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
||||
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
|
||||
]
|
||||
)
|
||||
@@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel:
|
||||
{
|
||||
"lane_id": "devops-check",
|
||||
"worker": "syslog-devops",
|
||||
"goal": "Check all 6 Proxmox nodes",
|
||||
"goal": "Check all 5 Proxmox nodes",
|
||||
"status": "dispatched|completed|failed",
|
||||
"output_file": "/tmp/node-report.md"
|
||||
}
|
||||
|
||||
@@ -54,7 +54,7 @@ garbled input. When a user types these commands to Abiba:
|
||||
|
||||
- **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately."
|
||||
- **`/approve session`** → Same response
|
||||
- **`/deny`** → Same response + "For Hermes agents (Tanko, Mumuni), these work with their built-in approval system."
|
||||
- **`/deny`** → Same response + "For agents (Mumuni on Hermes, Tanko on DSH), these work with their built-in approval system."
|
||||
|
||||
This keeps the UX consistent across agents — users can type `/approve` anywhere
|
||||
without getting confused by LLM responses.
|
||||
|
||||
+32
-24
@@ -2,16 +2,10 @@
|
||||
kind: responsibility
|
||||
name: pm2-self-heal
|
||||
description: >
|
||||
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
||||
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
||||
errored. Logs every action to the knowledge graph and alerts the owner via
|
||||
Zulip DM on failures.
|
||||
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
||||
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
||||
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
||||
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
||||
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
||||
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
||||
Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
|
||||
and auto-restarts any that are stopped or errored. Logs every action to
|
||||
Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary).
|
||||
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -19,13 +13,11 @@ description: >
|
||||
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
||||
- abiba-zulip: { status: "online", uptime: string, restarts: number }
|
||||
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
||||
- spoton-service: { status: "online", uptime: string, restarts: number }
|
||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||
- last_check: timestamp
|
||||
|
||||
> **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and
|
||||
> is monitored — the decommission note was stale (process re-added; do not treat
|
||||
> it as removed).
|
||||
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed.
|
||||
|
||||
|
||||
## Continuity
|
||||
|
||||
@@ -40,10 +32,18 @@ description: >
|
||||
- **Verify**: Re-check status after 5 seconds
|
||||
- **Escalate**: If still failed after 2 retries, send Zulip DM to owner
|
||||
|
||||
### Rule 2: Process Restarting Too Often
|
||||
- **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter)
|
||||
### Rule 2: Process Restarting Too Often (crash-loop guard)
|
||||
- **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter) **or** a
|
||||
process reporting a restart count > 1000 while showing "online" (a crash-loop mask)
|
||||
- **Note**: PM2 counter never decrements; only full delete+re-add resets it
|
||||
- **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>`
|
||||
- **Script guard (2026-08-16)**: `scripts/pm2-self-heal.sh` now restarts `abiba-telegram`
|
||||
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
||||
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
||||
spoton incident). Alerts include the restart count.
|
||||
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
|
||||
four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton
|
||||
reference above is historical context for the crash-loop guard, not a live process.
|
||||
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
||||
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
||||
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
||||
@@ -53,17 +53,25 @@ description: >
|
||||
## Execution
|
||||
|
||||
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
||||
2. **Check abiba-telegram**:
|
||||
2. **Check abiba-telegram** (safe to auto-restart):
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
|
||||
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
||||
- If status is "online" → pass, log restarts count
|
||||
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||
4. **Check gitea-runner**:
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
3. **Check abiba-zulip** (self-process, read-only):
|
||||
- If status is "online" → pass, log restarts count
|
||||
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
|
||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
6. **Wait 5 min** → repeat from step 1
|
||||
5. **Check zulip-watchdog**:
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
6. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||
7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
8. **Wait 5 min** → repeat from step 1
|
||||
|
||||
## Example Output (when healthy)
|
||||
|
||||
|
||||
@@ -5,7 +5,7 @@ description: >
|
||||
Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack
|
||||
on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three
|
||||
exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single
|
||||
instance), node_exporter (all 6 PVE nodes), and a custom docker-stats-exporter
|
||||
instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter
|
||||
(Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve
|
||||
the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx).
|
||||
agent: abiba
|
||||
@@ -33,8 +33,8 @@ agent: abiba
|
||||
|
||||
| Exporter | Host:Port | Scope | Notes |
|
||||
|----------|-----------|-------|-------|
|
||||
| prometheus-pve-exporter | .116:9221 (container) | All 6 nodes + guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
|
||||
| node_exporter | .5/.6/.9/.12/.15/.4:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 6 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. hwepve (.4) added 2026-07-19. |
|
||||
| prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
|
||||
| node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. |
|
||||
| docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. |
|
||||
|
||||
## PVE API Token
|
||||
@@ -48,7 +48,7 @@ agent: abiba
|
||||
|
||||
| UID | Title | Panels | Source |
|
||||
|-----|-------|--------|--------|
|
||||
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 6-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
|
||||
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
|
||||
| proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) |
|
||||
| docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) |
|
||||
| gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) |
|
||||
@@ -77,9 +77,9 @@ agent: abiba
|
||||
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
|
||||
| `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions |
|
||||
| `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning |
|
||||
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15/.4 | node_exporter collector config |
|
||||
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config |
|
||||
|
||||
## Cluster "Tabiri" — 6 Nodes
|
||||
## Cluster "Tabiri" — 5 Nodes
|
||||
|
||||
| Node | IP | Role |
|
||||
|------|----|----|
|
||||
@@ -87,8 +87,49 @@ agent: abiba
|
||||
| storepve | 192.168.68.6 | PVE |
|
||||
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
||||
| minipve | 192.168.68.12 | PVE |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
|
||||
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts abiba (lxc/100), kagentz (lxc/105), mumuni (lxc/114). CTs 100/105 migrated from amdpve, CT 114 from minipve 2026-07-20 |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
||||
|
||||
## PBS GC (Proxmox Backup Server)
|
||||
|
||||
### Schedule
|
||||
Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**.
|
||||
|
||||
**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.
|
||||
|
||||
### What Actually Runs
|
||||
Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`.
|
||||
|
||||
**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC.
|
||||
|
||||
### Datastore Location
|
||||
- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free)
|
||||
- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume)
|
||||
|
||||
### Liveness Check
|
||||
The new `proxmox-monitor.sh` leg checks storepve-datastore GC health:
|
||||
- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json`
|
||||
- **FAILS** if `last-run-endtime` is older than 48 hours
|
||||
- Reports age in hours and pending-bytes
|
||||
|
||||
**All six verdict shapes** (exactly as emitted by the script):
|
||||
|
||||
1. **Healthy** (fresh GC, 0 B pending):
|
||||
`✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)`
|
||||
|
||||
2. **Stale** (GC ran >48h ago):
|
||||
`🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)`
|
||||
|
||||
3. **Probe-failed: empty read** (000/timeout/unreadable):
|
||||
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)`
|
||||
|
||||
4. **Probe-failed: unparseable** (non-empty but invalid JSON — the "command not found" case):
|
||||
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)`
|
||||
|
||||
5. **Never-run: datastore absent** (valid JSON but storepve-datastore not in list):
|
||||
`🔴 PBS GC: never-run (storepve-datastore not found in GC list)`
|
||||
|
||||
6. **Never-run: no endtime** (valid JSON with datastore present but last-run-endtime is null/0):
|
||||
`🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)`
|
||||
|
||||
## Operations
|
||||
|
||||
@@ -101,6 +142,35 @@ Edit `build-dashboards.py`, run it, `docker restart harness-grafana`
|
||||
### check-targets
|
||||
`curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'`
|
||||
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Prometheus health (bound to 0.0.0.0:9090 on .116)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy
|
||||
# Expected: 200 (Prometheus is up and healthy)
|
||||
|
||||
# Grafana health (bound to 0.0.0.0:3001 on .116)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health
|
||||
# Expected: 200 (Grafana is up and healthy)
|
||||
|
||||
# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics"
|
||||
# Expected: 200 (docker-stats-exporter is up and responding)
|
||||
|
||||
# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics"
|
||||
# Expected: 200 (pve-exporter is up and responding)
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the probe executed
|
||||
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
|
||||
distinguishable from a real fault at read time. Summarize actual results from
|
||||
each probe. If any probe returns non-200, flag as alert.
|
||||
|
||||
**Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN.
|
||||
|
||||
### restart-exporter
|
||||
`cd /opt/monitoring && docker compose restart pve-exporter docker-stats`
|
||||
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
disk-gc report-only verification — CT 111 / tdunna / 192.168.68.129
|
||||
date: 2026-09-12T19:11:36Z
|
||||
host: abiba (this scanner runs INSIDE CT 100 / abiba)
|
||||
branch head: 6e612ce37b1b9a9688b04e7f816848e7a4185fca
|
||||
command: scripts/disk-gc-plan.py --scan <live fleet scan>
|
||||
|
||||
PURPOSE: prove that on a REAL fleet scan, CT 111 is alerted and NO gc-executor action
|
||||
is emitted for it at any level. No GC command was executed against .129.
|
||||
|
||||
--- live fleet scan (df -P / via scripts/pct-run.sh for LXC, direct SSH for hosts) ---
|
||||
tdunna 84%
|
||||
acerpve 192.168.68.9 77%
|
||||
amdpve 192.168.68.15 76%
|
||||
ocu-llm 192.168.68.110 69%
|
||||
storepve 192.168.68.6 65%
|
||||
kagentz 61%
|
||||
tanko 56%
|
||||
minipve 192.168.68.12 49%
|
||||
authentik 45%
|
||||
infisical-vault 39%
|
||||
adguard 38%
|
||||
ocupve 192.168.68.5 38%
|
||||
scottdenya 35%
|
||||
syslog-api 34%
|
||||
llm-gpu 192.168.68.8 25%
|
||||
abiba 23%
|
||||
baggy 22%
|
||||
jdownloader 21%
|
||||
gitea 16%
|
||||
ra-h-os 13%
|
||||
zulip 12%
|
||||
docker-vm 192.168.68.7 11%
|
||||
adguard2 10%
|
||||
media 9%
|
||||
proxmox-backup-server 4%
|
||||
|
||||
--- planner output (action plan) ---
|
||||
111 AMBER 84.0% -> REPORT-ONLY (no GC) — Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10
|
||||
acerpve AMBER 77.0% -> gc-executor
|
||||
amdpve AMBER 76.0% -> gc-executor
|
||||
|
||||
--- verdict ---
|
||||
CT 111 (tdunna) 84% AMBER -> report-only; no gc-executor row emitted; no GC run on .129.
|
||||
Owned hosts acerpve .9 (77%) and amdpve .15 (76%) -> gc-executor (ours).
|
||||
+331
-71
@@ -1,6 +1,6 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v2
|
||||
/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v4
|
||||
|
||||
Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming,
|
||||
gateway liveness, gateway log health, CT liveness, config YAML integrity,
|
||||
@@ -17,11 +17,42 @@ Changelog:
|
||||
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
|
||||
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
|
||||
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
|
||||
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
|
||||
abiba (.24).
|
||||
Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24).
|
||||
(v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she
|
||||
moved to her own container, kagentz CT 105 / .14, and is monitored there.)
|
||||
v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service,
|
||||
.110 llama-server.service, .15 strix-server.service) — .8 was probing a stale
|
||||
llama-server unit that reads inactive, producing false UNREACHABLE legs.
|
||||
systemctl is-active no longer swallows non-zero exit as SSH failure.
|
||||
Fixed UnboundLocalError on the abiba/koonimo gateway leg (pid unbound in the
|
||||
summary f-string). Abiba's LiteLLM key now comes from /root/.pi/agent/env.sh
|
||||
(#735 agent separation; creds moved out of shared /root/.bashrc).
|
||||
v4 (2026-09-10): probe-drift round 2 (prose-contracts follow-up to #65/#66/#68).
|
||||
abiba declared pi-only runtime — Hermes-era config/wrapper/gateway checks are
|
||||
skipped (harness purge). koby declared report_only per the captain's
|
||||
2026-08-17 ruling: every koby leg is detected and reported, never counted as a
|
||||
fleet failure and never repaired. koby's PVE mapping corrected to storepve
|
||||
(CT 111 tdunna lives on .6 — the old amdpve mapping produced a false
|
||||
ct-unreachable). The wrapper infisical-path check had two stale-expectation
|
||||
bugs: it read only the first 20 lines of the wrapper, so koonimo (whose
|
||||
wrapper does reference /usr/bin/infisical, just past line 20) was falsely
|
||||
FAILed as "path may be wrong"; and it treated the absence of any infisical
|
||||
reference as a fault, though koby's wrapper sources the key from
|
||||
~/.hermes/.env and never invokes infisical. The check now reads the full
|
||||
wrapper body, accepts a no-infisical wrapper, and verifies that any absolute
|
||||
infisical path the wrapper references actually exists. Report-only findings
|
||||
are surfaced in a machine-readable `report_only` array in --json output,
|
||||
separate from `failures`. Every run prints absolute execution provenance
|
||||
(script + cwd) in the header, in the cron ALERT line, and in --json output so
|
||||
a stale-consumer report is distinguishable from a fault at read time.
|
||||
v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed
|
||||
from the AGENTS dict when she moved off this host onto her own container
|
||||
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
|
||||
from her side. This script must not probe mumuni or .24 — the v2 changelog
|
||||
roster line was the last reference still placing her at .24 / CT100.
|
||||
"""
|
||||
|
||||
import subprocess, json, sys, os, time
|
||||
import subprocess, json, sys, os, time, re, io, contextlib
|
||||
from datetime import datetime
|
||||
|
||||
LITELLM = "http://192.168.68.116:80"
|
||||
@@ -30,7 +61,6 @@ INFISICAL_ENV = "prod"
|
||||
|
||||
# PVE node IPs for CT liveness checks
|
||||
PVE_NODES = {
|
||||
"hwepve": "192.168.68.4",
|
||||
"amdpve": "192.168.68.15",
|
||||
"minipve": "192.168.68.12",
|
||||
"storepve": "192.168.68.6",
|
||||
@@ -40,25 +70,60 @@ PVE_NODES = {
|
||||
|
||||
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
||||
AGENTS = {
|
||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"},
|
||||
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
|
||||
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
|
||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
|
||||
# abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its
|
||||
# local env file (key_env below), not from the shared vault or .bashrc.
|
||||
# runtime=pi: abiba has run pi-only since the harness purge. There is no
|
||||
# Hermes gateway, no ~/.hermes/config.yaml and no hermes CLI wrapper on .24
|
||||
# (the /root/.local/bin/hermes symlink is dangling), so the Hermes-era
|
||||
# config/wrapper/gateway legs are skipped rather than reported as faults.
|
||||
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve",
|
||||
"vault_key": None, "runtime": "pi",
|
||||
"key_env": {"file": "/root/.pi/agent/env.sh", "var": "LITELLM_API_KEY"}},
|
||||
# koby = report-only (captain's 2026-08-17 ruling, Rule 17): detect and
|
||||
# report, NEVER repair, and never count against fleet failures. CT 111
|
||||
# (tdunna) lives on storepve (.6) — verified live 2026-09-10; the previous
|
||||
# amdpve mapping made `pct status 111` fail and read as ct-unreachable.
|
||||
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "storepve", "vault_key": "KOBY_LITELLM_API_KEY", "report_only": True},
|
||||
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
|
||||
}
|
||||
|
||||
# Systemd units verified live 2026-09-08 (systemctl list-units on each host):
|
||||
# .8 rtx3090 (gpu-dense) -> llama-chat-api.service (active; the old
|
||||
# llama-server.service unit file is stale/inactive — probing it read as
|
||||
# UNREACHABLE for a healthy process)
|
||||
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
|
||||
# .15 strixhalo (amdpve) -> strix-server.service (active)
|
||||
GPU_HOSTS = {
|
||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
|
||||
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"},
|
||||
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server"},
|
||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"},
|
||||
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
|
||||
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
|
||||
}
|
||||
|
||||
FAIL = []
|
||||
REPORT_ONLY = []
|
||||
|
||||
|
||||
def _fail(key, agent_name=None):
|
||||
"""Record a failure, except for report-only agents.
|
||||
|
||||
Koby is report-only per the captain's 2026-08-17 ruling (Rule 17): its legs
|
||||
are detected and reported, never repaired and never counted as fleet
|
||||
failures. A red fleet alert on a known report-only leg is a false alarm.
|
||||
Report-only findings are tracked separately so --json consumers can still
|
||||
see them without them counting as fleet failures. Any non-report-only agent
|
||||
(or a leg with no agent, e.g. GPU hosts) records normally.
|
||||
"""
|
||||
if agent_name and AGENTS.get(agent_name, {}).get("report_only"):
|
||||
REPORT_ONLY.append(key)
|
||||
print(f" 🔍 report-only ({agent_name}): {key} — reported, not counted/repaired")
|
||||
return
|
||||
FAIL.append(key)
|
||||
|
||||
|
||||
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
|
||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||
|
||||
# Fallback: if no env token, read the shared vault token file
|
||||
if not INFISICAL_TOKEN:
|
||||
# Fallback: read the shared vault token file
|
||||
_token_path = os.path.expanduser("~/.infisical-token")
|
||||
if os.path.isfile(_token_path):
|
||||
try:
|
||||
@@ -66,6 +131,7 @@ if not INFISICAL_TOKEN:
|
||||
INFISICAL_TOKEN = _f.read().strip()
|
||||
except (OSError, UnicodeDecodeError):
|
||||
pass
|
||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||
|
||||
# ── Helpers ──────────────────────────────────────────────────────────
|
||||
|
||||
@@ -162,11 +228,46 @@ def _get_agent_key(agent_name, vault_key_name):
|
||||
|
||||
return None
|
||||
|
||||
# Inject keys from vault for each agent
|
||||
for agent_name in AGENTS:
|
||||
info = AGENTS[agent_name]
|
||||
key = _get_agent_key(agent_name, info.get("vault_key"))
|
||||
AGENTS[agent_name]["key"] = key
|
||||
|
||||
def _read_env_export(path, var):
|
||||
"""Parse `export VAR=value` (or `VAR=value`) out of a local env file.
|
||||
|
||||
#735 agent separation (2026-09-06): agent creds moved out of the shared
|
||||
/root/.bashrc into per-agent env files under /root/.pi/agent/ (bashrc's
|
||||
source line keeps abiba shells resolving them, but the file of record is
|
||||
env.sh). Do NOT fall back to /root/.bashrc here: desktop (.200) SSH
|
||||
sessions override LITELLM_API_KEY with mumuni's key, so sourcing bashrc
|
||||
would validate the wrong identity.
|
||||
"""
|
||||
try:
|
||||
with open(os.path.expanduser(path)) as _f:
|
||||
for line in _f:
|
||||
line = line.strip()
|
||||
if not (line.startswith("export " + var + "=") or line.startswith(var + "=")):
|
||||
continue
|
||||
value = line.split("=", 1)[1].strip().strip('"').strip("'")
|
||||
if value:
|
||||
return value
|
||||
except (OSError, UnicodeDecodeError):
|
||||
pass
|
||||
return None
|
||||
|
||||
|
||||
def load_agent_keys():
|
||||
"""Populate AGENTS[*]["key"] from the vault or the agent's local env file.
|
||||
|
||||
Called from main(), not at import: keeping this out of module scope lets the
|
||||
module be imported (and unit tested) without live vault/SSH access. Vault
|
||||
format is {NAME}_LITELLM_API_KEY (project 322fceab-39da-4854-a55a-568e76c0f13f,
|
||||
env prod); abiba has no vault key and reads LITELLM_API_KEY from its local
|
||||
/root/.pi/agent/env.sh (moved there from /root/.bashrc in #735).
|
||||
"""
|
||||
for agent_name in AGENTS:
|
||||
info = AGENTS[agent_name]
|
||||
key = _get_agent_key(agent_name, info.get("vault_key"))
|
||||
if not key and info.get("key_env"):
|
||||
key = _read_env_export(info["key_env"]["file"], info["key_env"]["var"])
|
||||
AGENTS[agent_name]["key"] = key
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
@@ -177,8 +278,8 @@ def check_keys():
|
||||
for name, agent in AGENTS.items():
|
||||
key = agent.get("key")
|
||||
if not key:
|
||||
print(f" ❌ {name}: NO KEY FOUND (vault empty or unreachable)")
|
||||
FAIL.append(f"key:{name}:no-key")
|
||||
print(f" ❌ {name}: NO KEY FOUND (vault/env empty or unreachable)")
|
||||
_fail(f"key:{name}:no-key", name)
|
||||
continue
|
||||
data = http_json(f"{LITELLM}/v1/models",
|
||||
headers={"Authorization": f"Bearer {key}"})
|
||||
@@ -187,11 +288,11 @@ def check_keys():
|
||||
print(f" ✅ {name}: key valid → {model}")
|
||||
else:
|
||||
print(f" ❌ {name}: KEY FAILURE — auth rejected or unreachable")
|
||||
FAIL.append(f"key:{name}")
|
||||
_fail(f"key:{name}", name)
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
# CHECK 2: GPU Port Conflict Detection (unchanged)
|
||||
# CHECK 2: GPU Port Conflict Detection (unit names verified live 2026-09-08)
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
|
||||
def check_gpu_ports():
|
||||
@@ -200,7 +301,11 @@ def check_gpu_ports():
|
||||
port = gpu["port"]
|
||||
svc = gpu["service"]
|
||||
|
||||
svc_status = ssh(host, f"systemctl is-active {svc}")
|
||||
# `systemctl is-active` exits non-zero when the unit is inactive or
|
||||
# missing, which the ssh() helper would swallow as an SSH failure and
|
||||
# report as UNREACHABLE. `|| true` keeps the real state word so we can
|
||||
# tell "unit inactive" from "host unreachable".
|
||||
svc_status = ssh(host, f"systemctl is-active {svc} || true")
|
||||
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
|
||||
|
||||
if not svc_status:
|
||||
@@ -234,28 +339,94 @@ def check_gpu_ports():
|
||||
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
|
||||
def _ssh_retry(host, cmd, user="root", timeout=15, retry_timeout=25, label=""):
|
||||
"""SSH with one retry at a longer timeout.
|
||||
|
||||
Returns (stdout_or_None, probe_failed_bool, fail_kind).
|
||||
When probe_failed is True, fail_kind is one of: timeout, ssh-failed.
|
||||
"""
|
||||
import subprocess as _sp
|
||||
def _attempt(tmo, conn_tmo):
|
||||
try:
|
||||
r = _sp.run(
|
||||
["ssh", "-o", "StrictHostKeyChecking=no", "-o", f"ConnectTimeout={conn_tmo}",
|
||||
f"{user}@{host}", cmd],
|
||||
capture_output=True, text=True, timeout=tmo)
|
||||
return r.stdout.strip() if r.returncode == 0 else None
|
||||
except _sp.TimeoutExpired:
|
||||
return "__timeout__"
|
||||
except:
|
||||
return None
|
||||
result = _attempt(timeout, 8)
|
||||
if result is None or result == "__timeout__":
|
||||
kind = "timeout" if result == "__timeout__" else "ssh-failed"
|
||||
prefix = f"{label} " if label else ""
|
||||
print(f" probe-failed: {prefix}ssh {user}@{host} — {kind} (retrying at {retry_timeout}s…)")
|
||||
result = _attempt(retry_timeout, 15)
|
||||
if result is None or result == "__timeout__":
|
||||
kind = "timeout" if result == "__timeout__" else "ssh-failed"
|
||||
return None, True, kind
|
||||
return result, False, None
|
||||
|
||||
|
||||
def check_agents():
|
||||
for name, agent in AGENTS.items():
|
||||
host = agent.get("host")
|
||||
user = agent.get("user")
|
||||
ct = agent["ct"]
|
||||
report_only = agent.get("report_only", False)
|
||||
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
|
||||
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
|
||||
# Non-Hermes runtimes have no gateway to probe. dsh = Tanko since
|
||||
# 2026-08-27; pi = abiba since the harness purge (.24 is pi-only).
|
||||
if agent.get("runtime") in ("dsh", "pi"):
|
||||
is_dsh = agent.get("runtime") == "dsh"
|
||||
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
|
||||
since = "since 2026-08-27" if is_dsh else "since the harness purge"
|
||||
live, probe_failed, fail_kind = _ssh_retry(host, "true", user=user)
|
||||
if probe_failed:
|
||||
print(f" ❌ {name}: {label} — probe-failed: ssh {user}@{host} {fail_kind} "
|
||||
f"(retried at 25s: also {fail_kind}) [CT {ct}]")
|
||||
_fail(f"probe-failed:{name}:{fail_kind}", name)
|
||||
else:
|
||||
print(f" ✅ {name}: {label} — no Hermes gateway {since} "
|
||||
f"(ssh {user}@{host} OK, CT {ct})")
|
||||
continue
|
||||
|
||||
if not host or not user:
|
||||
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
||||
continue
|
||||
|
||||
# Gateway process
|
||||
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||
if not pid:
|
||||
# Try alternate binary name
|
||||
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||
if not pid:
|
||||
print(f" ❌ {name}: GATEWAY NOT RUNNING")
|
||||
FAIL.append(f"gateway-down:{name}")
|
||||
# Resolve the Hermes gateway PID with retry. The probe target is
|
||||
# explicit: ssh {user}@{host} pgrep -f hermes gateway.
|
||||
pid, probe_failed, fail_kind = _ssh_retry(
|
||||
host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||
if not pid and not probe_failed:
|
||||
pid, probe_failed, fail_kind = _ssh_retry(
|
||||
host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||
if not pid and not probe_failed:
|
||||
pid = "?"
|
||||
|
||||
if probe_failed:
|
||||
print(f" ❌ {name}: probe-failed: ssh {user}@{host} {fail_kind} "
|
||||
f"(retried at 25s: also {fail_kind}) [CT {ct}] — gateway status UNDETERMINED")
|
||||
_fail(f"probe-failed:{name}:{fail_kind}", name)
|
||||
continue
|
||||
|
||||
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only (captain's 2026-08-17 ruling)
|
||||
if report_only:
|
||||
if pid == "?":
|
||||
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
|
||||
f"-> no process found (reported only, NOT counted) [CT {ct}]")
|
||||
_fail(f"gateway-down:{name}", name)
|
||||
else:
|
||||
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
|
||||
f"-> pid={pid} (running, reported only, NOT repaired) [CT {ct}]")
|
||||
continue # Skip the rest of the check for Koby
|
||||
|
||||
# Gateway state file
|
||||
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
||||
state, _, _ = _ssh_retry(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
||||
if state:
|
||||
try:
|
||||
st = json.loads(state)
|
||||
@@ -273,21 +444,22 @@ def check_agents():
|
||||
]
|
||||
streaming = "no"
|
||||
for p in adapter_paths:
|
||||
has_edit = ssh(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
|
||||
has_edit, _, _ = _ssh_retry(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
|
||||
if has_edit and has_edit != "0":
|
||||
streaming = "yes"
|
||||
break
|
||||
|
||||
# Recent errors
|
||||
recent_errors = ssh(host,
|
||||
recent_errors, _, _ = _ssh_retry(
|
||||
host,
|
||||
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
|
||||
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
|
||||
user=user)
|
||||
recent_errors = (recent_errors or "0").strip().split("\n")[-1]
|
||||
|
||||
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
|
||||
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
|
||||
f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
|
||||
f"{name}: probe: ssh {user}@{host} — gw={gw_state} zulip={zulip} "
|
||||
f"streaming={streaming} errors_10m={recent_errors.strip() or '0'} pid={pid} [CT {ct}]")
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
@@ -311,12 +483,12 @@ def check_ct_liveness():
|
||||
status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root")
|
||||
if not status:
|
||||
print(f" ❌ {name} (CT {ct} on {pve_node}): PVE UNREACHABLE")
|
||||
FAIL.append(f"ct-unreachable:{name}:{pve_ip}")
|
||||
_fail(f"ct-unreachable:{name}:{pve_ip}", name)
|
||||
elif "running" in status:
|
||||
print(f" ✅ {name} (CT {ct} on {pve_node}): running")
|
||||
elif "stopped" in status:
|
||||
print(f" ❌ {name} (CT {ct} on {pve_node}): STOPPED")
|
||||
FAIL.append(f"ct-stopped:{name}")
|
||||
_fail(f"ct-stopped:{name}", name)
|
||||
else:
|
||||
print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}")
|
||||
|
||||
@@ -328,6 +500,13 @@ def check_ct_liveness():
|
||||
def check_config_integrity():
|
||||
"""Verify agent config.yaml parses as valid YAML."""
|
||||
for name, agent in AGENTS.items():
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
|
||||
if agent.get("runtime") == "dsh":
|
||||
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
|
||||
continue
|
||||
if agent.get("runtime") == "pi":
|
||||
print(f" ⏭️ {name}: pi-only runtime — no Hermes config.yaml since the harness purge")
|
||||
continue
|
||||
host = agent.get("host")
|
||||
user = agent.get("user")
|
||||
if not host or not user:
|
||||
@@ -342,21 +521,48 @@ def check_config_integrity():
|
||||
user=user)
|
||||
if not yaml_ok:
|
||||
print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)")
|
||||
FAIL.append(f"config-unreachable:{name}")
|
||||
_fail(f"config-unreachable:{name}", name)
|
||||
elif "OK" in yaml_ok:
|
||||
print(f" ✅ {name}: config.yaml valid YAML")
|
||||
else:
|
||||
print(f" ❌ {name}: config.yaml YAML ERROR — {yaml_ok[:120]}")
|
||||
FAIL.append(f"config-yaml-error:{name}")
|
||||
_fail(f"config-yaml-error:{name}", name)
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
# CHECK 6: Wrapper/CLI Integrity (NEW)
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
|
||||
def _infisical_invocation_paths(wrapper_body):
|
||||
"""Absolute infisical paths the wrapper actually invokes.
|
||||
|
||||
Only executed (non-comment) lines count, and only a path followed by a real
|
||||
infisical subcommand (e.g. `/usr/bin/infisical run`) is treated as an
|
||||
invocation. A note such as `# migrated from /usr/local/bin/infisical` is
|
||||
prose, not a call, so it must not manufacture a dangling-path false alarm.
|
||||
"""
|
||||
paths = []
|
||||
for line in wrapper_body.splitlines():
|
||||
code = line.split("#", 1)[0]
|
||||
for _m in re.finditer(
|
||||
r"(/[A-Za-z0-9._/-]*infisical)\s+(?:run|export|secrets|login|logout)\b",
|
||||
code,
|
||||
):
|
||||
if _m.group(1) not in paths:
|
||||
paths.append(_m.group(1))
|
||||
return paths
|
||||
|
||||
|
||||
def check_wrapper_integrity():
|
||||
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
|
||||
for name, agent in AGENTS.items():
|
||||
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
|
||||
if agent.get("runtime") == "dsh":
|
||||
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
|
||||
continue
|
||||
if agent.get("runtime") == "pi":
|
||||
print(f" ⏭️ {name}: pi-only runtime — no hermes CLI wrapper since the harness purge")
|
||||
continue
|
||||
host = agent.get("host")
|
||||
user = agent.get("user")
|
||||
if not host or not user:
|
||||
@@ -370,24 +576,56 @@ def check_wrapper_integrity():
|
||||
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
|
||||
if not wrapper:
|
||||
print(f" ❌ {name}: NO HERMES CLI WRAPPER FOUND")
|
||||
FAIL.append(f"wrapper-missing:{name}")
|
||||
_fail(f"wrapper-missing:{name}", name)
|
||||
continue
|
||||
else:
|
||||
print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)")
|
||||
|
||||
# Check wrapper has correct infisical path
|
||||
infisical_path_valid = ssh(host,
|
||||
"head -20 /root/.local/bin/hermes 2>/dev/null | grep -q '/usr/bin/infisical' && echo OK || echo MISS",
|
||||
user=user)
|
||||
if infisical_path_valid == "MISS":
|
||||
# Check if infisical exists on path
|
||||
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
|
||||
if not inf_actual:
|
||||
print(f" ❌ {name}: INFISICAL NOT INSTALLED (wrapper broken)")
|
||||
FAIL.append(f"wrapper-no-infisical:{name}")
|
||||
# Credential-injection mechanism. The Hermes-era wrapper injected creds
|
||||
# with `/usr/bin/infisical run`, but the mechanism is not required to be
|
||||
# infisical at all: koby's wrapper sources the key from ~/.hermes/.env
|
||||
# and never mentions infisical, which is valid. The old check read only
|
||||
# the first 20 lines, so koonimo's wrapper — which DOES reference
|
||||
# /usr/bin/infisical, just past line 20 — false-failed as "path may be
|
||||
# wrong". Read the full body, accept a no-infisical wrapper, and verify
|
||||
# the absolute infisical path(s) the wrapper actually invokes. Only
|
||||
# executed (non-comment) lines count: a comment or dead prose mentioning
|
||||
# a removed path (litellm-api-keys.prose.md documents
|
||||
# `rm -f /usr/local/bin/infisical`) must neither produce a dangling path
|
||||
# nor trigger the PATH check — it is not an invocation.
|
||||
wrapper_body = ssh(host, "cat /root/.local/bin/hermes 2>/dev/null", user=user) or ""
|
||||
wrapper_code = "\n".join(line.split("#", 1)[0] for line in wrapper_body.splitlines())
|
||||
invoked_paths = _infisical_invocation_paths(wrapper_body)
|
||||
if "infisical" in wrapper_code:
|
||||
if invoked_paths:
|
||||
missing = []
|
||||
for _p in invoked_paths:
|
||||
_exists = ssh(host, f"test -x {_p} && echo OK || echo MISS", user=user)
|
||||
if not _exists or _exists.strip().splitlines()[-1] != "OK":
|
||||
missing.append(_p)
|
||||
if len(missing) == len(invoked_paths):
|
||||
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
|
||||
suffix = f" (infisical at {inf_actual})" if inf_actual else ""
|
||||
print(f" ❌ {name}: wrapper invokes infisical via missing path(s) "
|
||||
f"{', '.join(missing)}{suffix}")
|
||||
_fail(f"wrapper-infisical-path:{name}", name)
|
||||
elif missing:
|
||||
print(f" ⚠️ {name}: wrapper has an unused/missing infisical path "
|
||||
f"({', '.join(missing)}) but a working invocation — informational")
|
||||
elif "/usr/bin/infisical" not in invoked_paths:
|
||||
print(f" ⚠️ {name}: wrapper infisical path differs "
|
||||
f"({', '.join(invoked_paths)}) — informational")
|
||||
else:
|
||||
print(f" ✅ {name}: wrapper infisical path OK")
|
||||
else:
|
||||
print(f" ⚠️ {name}: wrapper infisical path may be wrong (infisical at {inf_actual})")
|
||||
FAIL.append(f"wrapper-infisical-path:{name}")
|
||||
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
|
||||
if not inf_actual:
|
||||
print(f" ❌ {name}: wrapper invokes infisical but the binary is MISSING")
|
||||
_fail(f"wrapper-no-infisical:{name}", name)
|
||||
else:
|
||||
print(f" ✅ {name}: wrapper infisical resolves via PATH ({inf_actual})")
|
||||
else:
|
||||
print(f" ℹ️ {name}: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK")
|
||||
|
||||
# Check hermes-real exists
|
||||
hermes_real = ssh(host,
|
||||
@@ -400,7 +638,7 @@ def check_wrapper_integrity():
|
||||
user=user)
|
||||
if not hermes_real or hermes_real.strip() == "MISS":
|
||||
print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)")
|
||||
FAIL.append(f"wrapper-no-hermes-real:{name}")
|
||||
_fail(f"wrapper-no-hermes-real:{name}", name)
|
||||
else:
|
||||
print(f" ✅ {name}: hermes-real at alt path")
|
||||
|
||||
@@ -428,10 +666,10 @@ def check_vault_secrets():
|
||||
key = agent.get("key")
|
||||
if not key:
|
||||
print(f" ❌ {name}: vault secret {vault_key_name} MISSING or EMPTY")
|
||||
FAIL.append(f"vault-empty:{name}:{vault_key_name}")
|
||||
_fail(f"vault-empty:{name}:{vault_key_name}", name)
|
||||
elif not key.startswith("sk-"):
|
||||
print(f" ❌ {name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')")
|
||||
FAIL.append(f"vault-bad-format:{name}:{vault_key_name}")
|
||||
_fail(f"vault-bad-format:{name}:{vault_key_name}", name)
|
||||
else:
|
||||
print(f" ✅ {name}: vault {vault_key_name}=sk-...{key[-4:]}")
|
||||
|
||||
@@ -464,18 +702,7 @@ def deploy_self():
|
||||
# MAIN
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
|
||||
def main():
|
||||
quiet = "--quiet" in sys.argv
|
||||
as_json = "--json" in sys.argv
|
||||
|
||||
# Self-deploy to canonical location
|
||||
if not quiet and "--no-deploy" not in sys.argv:
|
||||
deploy_self()
|
||||
|
||||
if not quiet:
|
||||
print(f"🏥 Agent Health Check v2 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}")
|
||||
print()
|
||||
|
||||
def _run_checks():
|
||||
print("🔑 LiteLLM Keys:")
|
||||
check_keys()
|
||||
print()
|
||||
@@ -503,16 +730,49 @@ def main():
|
||||
print("🔐 Vault Secrets:")
|
||||
check_vault_secrets()
|
||||
|
||||
|
||||
def main():
|
||||
quiet = "--quiet" in sys.argv
|
||||
as_json = "--json" in sys.argv
|
||||
|
||||
# Self-deploy to canonical location
|
||||
if not quiet and "--no-deploy" not in sys.argv:
|
||||
deploy_self()
|
||||
|
||||
# Provenance: a report is only actionable if the reader can tell WHICH copy
|
||||
# of this script produced it. A normal run carries it in the header, --json
|
||||
# carries it for machine consumers, and the cron ALERT line carries it on
|
||||
# failure. --quiet is documented as "only output on failure", so the header
|
||||
# is emitted only when not quiet and a healthy quiet run stays silent.
|
||||
script_path = os.path.abspath(__file__)
|
||||
cwd = os.getcwd()
|
||||
|
||||
if quiet:
|
||||
captured = io.StringIO()
|
||||
with contextlib.redirect_stdout(captured):
|
||||
load_agent_keys()
|
||||
_run_checks()
|
||||
if FAIL:
|
||||
sys.stdout.write(captured.getvalue())
|
||||
else:
|
||||
print(f"🏥 Agent Health Check v4 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}")
|
||||
print(f"📍 executed from: script={script_path} cwd={cwd}")
|
||||
print()
|
||||
load_agent_keys()
|
||||
_run_checks()
|
||||
|
||||
if FAIL:
|
||||
print(f"\n❌ {len(FAIL)} FAILURE(S): {' | '.join(FAIL)}")
|
||||
if quiet:
|
||||
print(f"ALERT agent-health:{','.join(FAIL)}")
|
||||
print(f"ALERT agent-health:{','.join(FAIL)} script={script_path} cwd={cwd}")
|
||||
elif not quiet:
|
||||
print("\n✅ All checks passed")
|
||||
|
||||
if as_json:
|
||||
print(json.dumps({"timestamp": datetime.now().isoformat(),
|
||||
"failures": FAIL, "healthy": len(FAIL) == 0}))
|
||||
"execution_path": script_path, "cwd": cwd,
|
||||
"failures": FAIL, "report_only": REPORT_ONLY,
|
||||
"healthy": len(FAIL) == 0}))
|
||||
|
||||
sys.exit(1 if FAIL else 0)
|
||||
|
||||
|
||||
Executable
+227
@@ -0,0 +1,227 @@
|
||||
#!/usr/bin/env bash
|
||||
# capture-dsh-token.sh — refresh the dsh-web login token WITHOUT restarting dsh-web.
|
||||
#
|
||||
# Context (CT 112 / tankodhs.sysloggh.net)
|
||||
# ----------------------------------------
|
||||
# The dsh-web UI (systemd unit `dsh-web.service`, 127.0.0.1:3080) prints a random
|
||||
# launch token to the journal on every start:
|
||||
#
|
||||
# dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||||
#
|
||||
# That token is the only way to bootstrap the authority-bound 30-day browser
|
||||
# cookie. It rotates on every dsh-web start, so the Authentik-gated
|
||||
# `location = /dsh-web-login` in /etc/nginx/sites-available/dsh must always
|
||||
# reference the token of the RUNNING process.
|
||||
#
|
||||
# This script:
|
||||
# 1. selects the launch token the RUNNING service actually accepts from the
|
||||
# current systemd invocation — it NEVER stops or starts dsh-web,
|
||||
# 2. records it in /etc/dsh-web/launch-token,
|
||||
# 3. regenerates the nginx include /etc/dsh-web/nginx-login.conf (the
|
||||
# `proxy_pass ...?token=` line consumed by /dsh-web-login),
|
||||
# 4. reloads nginx ONLY when the on-disk include differs from the generated
|
||||
# one or the applied-state stamp does not match the token (the stamp is
|
||||
# written only after a successful reload), rolling the include back on
|
||||
# failure so the next run retries,
|
||||
# 5. removes the legacy unauthenticated :8081 endpoint if it ever reappears.
|
||||
#
|
||||
# Idempotent and safe to run at any time (systemd ExecStartPost or timer).
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
|
||||
|
||||
JOURNAL_UNIT="dsh-web.service"
|
||||
TOKEN_FILE="/etc/dsh-web/launch-token"
|
||||
INCLUDE_FILE="/etc/dsh-web/nginx-login.conf"
|
||||
STAMP_FILE="/etc/dsh-web/nginx-login.conf.applied"
|
||||
PENDING_FILE="/etc/dsh-web/nginx-reload.pending"
|
||||
SITE_ENABLED="/etc/nginx/sites-enabled/dsh"
|
||||
LEGACY_8081="/etc/nginx/sites-enabled/dsh.token"
|
||||
STASH_DIR="/etc/nginx/sites-available"
|
||||
LOCK_FILE="/run/capture-dsh-token.lock"
|
||||
LOGIN_HOST="tankodhs.sysloggh.net"
|
||||
LOGIN_UPSTREAM="http://127.0.0.1:3080"
|
||||
TOKEN_WAIT=120
|
||||
|
||||
log() { printf 'capture-dsh-token: %s\n' "$*" >&2; }
|
||||
die() { printf 'capture-dsh-token: ERROR: %s\n' "$*" >&2; exit 1; }
|
||||
|
||||
[ "$(id -u)" -eq 0 ] || die "must run as root"
|
||||
|
||||
# ── 0. Serialize runs so timer/ExecStartPost/manual runs cannot interleave ──
|
||||
exec 9>"$LOCK_FILE"
|
||||
flock -n 9 || { log "another capture-dsh-token run holds $LOCK_FILE; exiting"; exit 0; }
|
||||
mkdir -p "$(dirname "$PENDING_FILE")"
|
||||
|
||||
# ── 0b. Guarantee the generated include exists before any `nginx -t` ──────
|
||||
# The :80 site includes /etc/dsh-web/nginx-login.conf by literal path, so a
|
||||
# missing include makes every `nginx -t` fail and can wedge recovery. Seed it
|
||||
# from the last known token (or a placeholder); step 4 replaces it.
|
||||
if [ ! -f "$INCLUDE_FILE" ]; then
|
||||
SEED="placeholder"
|
||||
if [ -f "$TOKEN_FILE" ]; then
|
||||
SEED="$(cat "$TOKEN_FILE" 2>/dev/null || true)"
|
||||
[ -n "$SEED" ] || SEED="placeholder"
|
||||
fi
|
||||
printf '%s' "$SEED" | grep -qE '^[A-Za-z0-9._~+/=:@-]+$' || SEED="placeholder"
|
||||
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$SEED" > "$INCLUDE_FILE"
|
||||
chmod 600 "$INCLUDE_FILE"
|
||||
log "created missing $INCLUDE_FILE"
|
||||
fi
|
||||
|
||||
# ── 1. Remove the legacy unauthenticated :8081 endpoint, if present ─────────
|
||||
# It bypassed Authentik entirely (listened on 0.0.0.0:8081 with no auth_request)
|
||||
# and must never come back. Stash it rather than delete so it is auditable.
|
||||
if [ -e "$LEGACY_8081" ] || [ -L "$LEGACY_8081" ]; then
|
||||
TS="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
STASHED="$STASH_DIR/dsh.token.disabled-$TS"
|
||||
mv "$LEGACY_8081" "$STASHED"
|
||||
chmod 600 "$STASHED" 2>/dev/null || true
|
||||
touch "$PENDING_FILE"
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
die "nginx config test failed after disabling $LEGACY_8081 (kept disabled at $STASHED): $NGINX_TEST_OUT; a pending reload is recorded so running nginx is reloaded once the config is fixed. The legacy :8081 endpoint will NOT be restored."
|
||||
fi
|
||||
if ! nginx -s reload; then
|
||||
die "nginx reload failed after disabling $LEGACY_8081 (kept disabled at $STASHED); a pending reload is recorded so running nginx is reloaded on the next run. The legacy :8081 endpoint will NOT be restored."
|
||||
fi
|
||||
rm -f "$PENDING_FILE"
|
||||
log "removed legacy :8081 endpoint -> $STASHED"
|
||||
fi
|
||||
|
||||
# ── 1b. Honor a recorded pending reload regardless of token selection ───────
|
||||
# A failed reload leaves PENDING_FILE set so a stashed legacy :8081 file can
|
||||
# never remain loaded in the running nginx while dsh-web is down or not yet
|
||||
# answering. Reconcile it before the token wait.
|
||||
if [ -e "$PENDING_FILE" ]; then
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
log "WARNING: pending nginx reload recorded but 'nginx -t' fails: $NGINX_TEST_OUT; continuing so the include can be regenerated; will retry next run"
|
||||
elif ! nginx -s reload; then
|
||||
log "WARNING: pending nginx reload recorded but 'nginx -s reload' failed; will retry next run"
|
||||
else
|
||||
rm -f "$PENDING_FILE"
|
||||
log "completed pending nginx reload"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── 2. Select the token the RUNNING service actually accepts ────────────────
|
||||
# Re-sample the service's CURRENT systemd invocation on every pass and read
|
||||
# candidates only from it, so a restart that lands during the wait immediately
|
||||
# switches to the new invocation; there is no whole-journal or cross-invocation
|
||||
# fallback, and an empty/unknown invocation just waits. Each candidate is then
|
||||
# functionally verified against the local dsh-web using the public authority,
|
||||
# exactly as the /dsh-web-login proxy does, and the first that answers 303 is
|
||||
# the live token. Candidates are re-probed newest-first on each pass (connection
|
||||
# failures stay eligible) until one is accepted or the wait elapses.
|
||||
journal_tokens() {
|
||||
journalctl -u "$JOURNAL_UNIT" "_SYSTEMD_INVOCATION_ID=$1" --no-pager -o cat 2>/dev/null \
|
||||
| grep -oE 'dsh web: https?://[^[:space:]]+[?&]token=[^[:space:]]+' \
|
||||
| sed -E 's/.*[?&]token=//' \
|
||||
| grep -E '^[A-Za-z0-9._~+/=:@-]+$' \
|
||||
| tac | awk '!seen[$0]++' || true
|
||||
}
|
||||
|
||||
TOKEN=""
|
||||
DEADLINE=$((SECONDS + TOKEN_WAIT))
|
||||
NO_INVOCATION_WARNED=0
|
||||
while [ -z "$TOKEN" ] && [ "$SECONDS" -lt "$DEADLINE" ]; do
|
||||
INVOCATION="$(systemctl show -p InvocationID --value "$JOURNAL_UNIT" 2>/dev/null || true)"
|
||||
if [ -z "$INVOCATION" ] || [ "$INVOCATION" = "n/a" ]; then
|
||||
if [ "$NO_INVOCATION_WARNED" -eq 0 ]; then
|
||||
log "WARNING: no invocation id for $JOURNAL_UNIT; waiting for a live invocation"
|
||||
NO_INVOCATION_WARNED=1
|
||||
fi
|
||||
sleep 2
|
||||
continue
|
||||
fi
|
||||
for cand in $(journal_tokens "$INVOCATION"); do
|
||||
code="$(curl -s -o /dev/null --max-time 5 -w '%{http_code}' \
|
||||
-H "Host: $LOGIN_HOST" "$LOGIN_UPSTREAM/?token=$cand" || true)"
|
||||
if [ "$code" = "303" ]; then
|
||||
TOKEN="$cand"
|
||||
break
|
||||
fi
|
||||
done
|
||||
[ -n "$TOKEN" ] && break
|
||||
sleep 2
|
||||
done
|
||||
|
||||
if [ -z "$TOKEN" ]; then
|
||||
log "no accepted launch token in the current invocation within ${TOKEN_WAIT}s; leaving the include untouched for the next run"
|
||||
[ -e "$PENDING_FILE" ] && die "pending nginx reload could not be completed; will retry next run"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ── 3. Record the token (atomic, private) ──────────────────────────────────
|
||||
mkdir -p "$(dirname "$TOKEN_FILE")"
|
||||
if ! printf '%s\n' "$TOKEN" | cmp -s - "$TOKEN_FILE" 2>/dev/null; then
|
||||
printf '%s\n' "$TOKEN" > "$TOKEN_FILE.tmp"
|
||||
chmod 600 "$TOKEN_FILE.tmp"
|
||||
mv "$TOKEN_FILE.tmp" "$TOKEN_FILE"
|
||||
log "recorded live launch token in $TOKEN_FILE"
|
||||
fi
|
||||
chmod 600 "$TOKEN_FILE"
|
||||
|
||||
# ── 4. Regenerate the nginx login include (reload only when it changes) ────
|
||||
NEW_INCLUDE="$(mktemp "$INCLUDE_FILE.XXXXXX")"
|
||||
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$TOKEN" > "$NEW_INCLUDE"
|
||||
chmod 600 "$NEW_INCLUDE"
|
||||
|
||||
# The stamp records the token nginx actually loaded. It is written only after a
|
||||
# successful reload, so the early exit is safe only when both the stamp and the
|
||||
# on-disk include agree with the live token; anything else falls through to the
|
||||
# reload path so the include can never silently diverge from what nginx serves.
|
||||
APPLIED=""
|
||||
[ -f "$STAMP_FILE" ] && APPLIED="$(cat "$STAMP_FILE" 2>/dev/null || true)"
|
||||
[ -f "$INCLUDE_FILE" ] && chmod 600 "$INCLUDE_FILE"
|
||||
[ -f "$STAMP_FILE" ] && chmod 600 "$STAMP_FILE"
|
||||
|
||||
if [ "$APPLIED" = "$TOKEN" ] && [ -f "$INCLUDE_FILE" ] && cmp -s "$NEW_INCLUDE" "$INCLUDE_FILE" \
|
||||
&& [ ! -e "$PENDING_FILE" ]; then
|
||||
rm -f "$NEW_INCLUDE"
|
||||
log "token unchanged; nginx not reloaded"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[ -e "$SITE_ENABLED" ] || { rm -f "$NEW_INCLUDE"; die "$SITE_ENABLED missing; refusing to reload"; }
|
||||
|
||||
RESTORE=""
|
||||
if [ -f "$INCLUDE_FILE" ]; then
|
||||
RESTORE="$(mktemp "$INCLUDE_FILE.bak.XXXXXX")"
|
||||
cp -p "$INCLUDE_FILE" "$RESTORE"
|
||||
chmod 600 "$RESTORE"
|
||||
fi
|
||||
|
||||
mv "$NEW_INCLUDE" "$INCLUDE_FILE"
|
||||
chmod 600 "$INCLUDE_FILE"
|
||||
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
if [ -n "$RESTORE" ]; then
|
||||
mv "$RESTORE" "$INCLUDE_FILE"
|
||||
else
|
||||
rm -f "$INCLUDE_FILE"
|
||||
fi
|
||||
die "nginx config test failed: $NGINX_TEST_OUT; previous include restored"
|
||||
fi
|
||||
|
||||
if ! nginx -s reload; then
|
||||
if [ -n "$RESTORE" ]; then
|
||||
mv "$RESTORE" "$INCLUDE_FILE"
|
||||
else
|
||||
rm -f "$INCLUDE_FILE"
|
||||
fi
|
||||
touch "$PENDING_FILE"
|
||||
die "nginx reload failed; previous include restored; a pending reload is recorded so the next run retries"
|
||||
fi
|
||||
|
||||
if [ -n "$RESTORE" ]; then
|
||||
rm -f "$RESTORE"
|
||||
fi
|
||||
|
||||
rm -f "$PENDING_FILE"
|
||||
|
||||
printf '%s\n' "$TOKEN" > "$STAMP_FILE.tmp"
|
||||
chmod 600 "$STAMP_FILE.tmp"
|
||||
mv "$STAMP_FILE.tmp" "$STAMP_FILE"
|
||||
|
||||
log "token changed; nginx reloaded"
|
||||
log "login endpoint: https://$LOGIN_HOST/dsh-web-login (Authentik-gated)"
|
||||
+116
-85
@@ -15,21 +15,31 @@ import smtplib, json, subprocess, os, sys, datetime, re
|
||||
from email.mime.text import MIMEText
|
||||
from email.mime.multipart import MIMEMultipart
|
||||
|
||||
PVE = "https://minipve.sysloggh.net"
|
||||
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
PVE = "https://192.168.68.12:8006"
|
||||
AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
|
||||
|
||||
# ── Shared credentials —─
|
||||
|
||||
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_KEY = "cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_KEY}"
|
||||
# Note: /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential).
|
||||
# No Zulip API key is required for this call. If a future leg genuinely needs abiba-bot's key,
|
||||
# it must prove it with a 200 from /api/v1/users/me as abiba-bot and label itself degraded when it cannot.
|
||||
# Never fall back to the vault's shared ZULIP_API_KEY.
|
||||
ZULIP_AUTH = None
|
||||
DEGRADED_LEGS = []
|
||||
|
||||
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||
LITELLM_BACKEND = "192.168.68.116"
|
||||
AUTH_HOST = "192.168.68.11"
|
||||
|
||||
SYNTHETIC_API_KEY = "sk-U_ydi3B-wfGU-_xESkoU1Q"
|
||||
# Load LiteLLM API key from file (durable, works in cron)
|
||||
LITELLM_KEY_FILE = "/root/.abiba-workspace/secrets/litellm-key.txt"
|
||||
try:
|
||||
with open(LITELLM_KEY_FILE) as f:
|
||||
SYNTHETIC_API_KEY = f.read().strip()
|
||||
except:
|
||||
SYNTHETIC_API_KEY = None # Fail loudly: report "no-key-file" in check
|
||||
|
||||
NOW = datetime.datetime.now()
|
||||
DATE_STR = NOW.strftime("%Y-%m-%d")
|
||||
@@ -38,10 +48,16 @@ TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
|
||||
# ── Helpers ──
|
||||
|
||||
def pve_get(path):
|
||||
cmd = f'curl -sfk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list)."""
|
||||
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||
try:
|
||||
return json.loads(subprocess.check_output(cmd, shell=True))["data"]
|
||||
except: return []
|
||||
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
|
||||
if r.returncode != 0:
|
||||
return None
|
||||
data = json.loads(r.stdout)
|
||||
return data.get("data", [])
|
||||
except:
|
||||
return None
|
||||
|
||||
def ssh(host, cmd):
|
||||
try:
|
||||
@@ -62,7 +78,7 @@ def http_get(url, auth=None, timeout=10):
|
||||
cmd = f'curl -sfk --connect-timeout {timeout} -o /dev/null -w "%{{http_code}}" "{url}"'
|
||||
if auth:
|
||||
cmd = cmd.replace('"', '\\"')
|
||||
cmd = f'curl -sfk --connect-timeout {timeout} -u "{auth}" -o /dev/null -w "%{{http_code}}" "{url}"'
|
||||
cmd = f'curl -sfk --connect-timeout {timeout} -H "Authorization: Bearer {auth}" -o /dev/null -w "%{{http_code}}" "{url}"'
|
||||
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout+2)
|
||||
return r.stdout.strip() or "000"
|
||||
except:
|
||||
@@ -97,21 +113,33 @@ def collect():
|
||||
|
||||
# ── Proxmox Nodes ──
|
||||
nodes = pve_get("/api2/json/nodes")
|
||||
report["nodes"] = {n["node"]: {
|
||||
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
|
||||
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
|
||||
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
|
||||
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
|
||||
"uptime_h": n.get('uptime',0)//3600,
|
||||
"status": n["status"]
|
||||
} for n in nodes}
|
||||
report["node_count"] = len(nodes)
|
||||
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
|
||||
if nodes is None:
|
||||
report["nodes"] = {}
|
||||
report["node_count"] = 0
|
||||
report["nodes_online"] = 0
|
||||
report["pve_probe_status"] = "unreachable"
|
||||
else:
|
||||
report["nodes"] = {n["node"]: {
|
||||
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
|
||||
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
|
||||
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
|
||||
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
|
||||
"uptime_h": n.get('uptime',0)//3600,
|
||||
"status": n["status"]
|
||||
} for n in nodes}
|
||||
report["node_count"] = len(nodes)
|
||||
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
|
||||
report["pve_probe_status"] = "ok"
|
||||
|
||||
# ── VMs/CTs ──
|
||||
resources = pve_get("/api2/json/cluster/resources")
|
||||
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||
if resources is None:
|
||||
vms = []
|
||||
report["resources_probe_status"] = "unreachable"
|
||||
else:
|
||||
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||
report["resources_probe_status"] = "ok"
|
||||
report["total_vms"] = len(vms)
|
||||
report["running_vms"] = sum(1 for v in vms if v.get("status") == "running")
|
||||
stopped = [v for v in vms if v.get("status") != "running"]
|
||||
@@ -132,7 +160,7 @@ def collect():
|
||||
# ── Storage ──
|
||||
storages = pve_get("/api2/json/nodes/storepve/storage")
|
||||
report["storage"] = []
|
||||
for s in storages:
|
||||
for s in (storages or []):
|
||||
total = s.get("total",0) or 1
|
||||
used = s.get("used",0)
|
||||
pct = used/total*100
|
||||
@@ -183,7 +211,7 @@ def collect():
|
||||
("Authentik", "https://auth.sysloggh.net"),
|
||||
("Zulip", "https://chat.sysloggh.net"),
|
||||
("Pulse", "https://pulse.sysloggh.net"),
|
||||
("Proxmox", "https://minipve.sysloggh.net"),
|
||||
("Proxmox", "https://192.168.68.12:8006"),
|
||||
("SearXNG", "http://192.168.68.7:8888"),
|
||||
("Firecrawl", "http://192.168.68.7:3002/health"),
|
||||
]
|
||||
@@ -195,10 +223,11 @@ def collect():
|
||||
# ── LiteLLM Specific Checks (from litellm-health prose contract) ──
|
||||
report["litellm"] = {"checks": []}
|
||||
|
||||
# Check 1: LiteLLM aggregate health endpoint (router binds to 127.0.0.1, check via SSH)
|
||||
health_unified = ssh(LITELLM_BACKEND, "curl -sf http://127.0.0.1:9000/health/unified -o /dev/null -w '%{http_code}' 2>/dev/null")
|
||||
# Check 1: Fleet health via nginx /health/unified (301 -> /gpu/gpu-data served by gpu-monitor on .24:9100).
|
||||
# Router (:9000) was decommissioned 2026-09-11; probing it was a guaranteed daily failure.
|
||||
health_unified = ssh(LITELLM_BACKEND, "curl -s -o /dev/null -w '%{http_code}' --max-time 8 http://127.0.0.1/health/unified 2>/dev/null")
|
||||
report["litellm"]["health_unified"] = health_unified or "000"
|
||||
report["litellm"]["checks"].append({"name": "unified-health", "status": "pass" if health_unified == "200" else "fail", "code": health_unified or "000"})
|
||||
report["litellm"]["checks"].append({"name": "fleet-health-via-nginx", "status": "pass" if health_unified in ("200", "301") else "fail", "code": health_unified or "000"})
|
||||
|
||||
# Check 2: Nginx-proxied internal endpoints
|
||||
for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]:
|
||||
@@ -206,8 +235,11 @@ def collect():
|
||||
report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code})
|
||||
|
||||
# Check 3: Docker container health for LiteLLM stack
|
||||
expected_containers = ["harness-litellm", "harness-nginx", "harness-router",
|
||||
"harness-postgres", "harness-redis", "harness-dashboard"]
|
||||
expected_containers = ["harness-litellm", "harness-nginx", "harness-postgres",
|
||||
"harness-redis", "harness-dashboard", "harness-grafana",
|
||||
"harness-prometheus", "harness-alertmanager",
|
||||
"harness-zulip-bridge", "harness-docker-stats",
|
||||
"harness-pve-exporter"]
|
||||
actual_names = [c["name"] for c in containers2]
|
||||
report["litellm"]["expected_containers"] = expected_containers
|
||||
report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names]
|
||||
@@ -225,9 +257,13 @@ def collect():
|
||||
report["litellm"]["checks"].append({"name": "oidc-auth", "status": "pass" if auth_code in ("200","302") else "fail", "code": auth_code})
|
||||
|
||||
# Check 5: Synthetic API call through LiteLLM
|
||||
api_check = http_get(f"{LITELLM_PUBLIC}/v1/models", auth=SYNTHETIC_API_KEY)
|
||||
report["litellm"]["api_models"] = api_check
|
||||
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "pass" if api_check == "200" else "fail", "code": api_check})
|
||||
if SYNTHETIC_API_KEY is None:
|
||||
report["litellm"]["api_models"] = None
|
||||
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "fail", "code": "no-key-file"})
|
||||
else:
|
||||
api_check = http_get(f"{LITELLM_PUBLIC}/v1/models", auth=SYNTHETIC_API_KEY)
|
||||
report["litellm"]["api_models"] = api_check
|
||||
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "pass" if api_check == "200" else "fail", "code": api_check})
|
||||
|
||||
# ── NFS Mounts ──
|
||||
nfs = ssh("192.168.68.7", "df -h /media/storage /media/mediastore 2>/dev/null | tail -n +2")
|
||||
@@ -247,11 +283,13 @@ def collect():
|
||||
zulip_health = json.loads(health_body) if health_body else {}
|
||||
except:
|
||||
zulip_health = {}
|
||||
report["zulip_ext"]["connected"] = zulip_health.get("connected", False)
|
||||
report["zulip_ext"]["queue_id"] = zulip_health.get("queue_id")
|
||||
report["zulip_ext"]["last_error"] = zulip_health.get("last_error")
|
||||
report["zulip_ext"]["messages_processed"] = zulip_health.get("messages_processed", 0)
|
||||
report["zulip_ext"]["retry_count"] = zulip_health.get("retry_count", 0)
|
||||
# Live state is nested under 'zulip' key
|
||||
zulip_state = zulip_health.get("zulip", {})
|
||||
report["zulip_ext"]["connected"] = zulip_state.get("connected", False)
|
||||
report["zulip_ext"]["queue_id"] = zulip_state.get("queue_id")
|
||||
report["zulip_ext"]["last_error"] = zulip_state.get("last_error")
|
||||
report["zulip_ext"]["messages_processed"] = zulip_state.get("messages_processed", 0)
|
||||
report["zulip_ext"]["skipped"] = zulip_state.get("skipped", 0)
|
||||
|
||||
# Phase 2: PM2 process check
|
||||
pm2_raw = subprocess.check_output(
|
||||
@@ -289,50 +327,32 @@ def collect():
|
||||
# Abiba (pi)
|
||||
report["agents"]["abiba"] = {
|
||||
"platform": "pi", "ct": 100, "ip": "192.168.68.24",
|
||||
"zulip_connected": zulip_health.get("connected", False),
|
||||
"zulip_processed": zulip_health.get("messages_processed", 0),
|
||||
"zulip_connected": zulip_state.get("connected", False),
|
||||
"zulip_processed": zulip_state.get("messages_processed", 0),
|
||||
"pm2_status": pm2.get("status", "unknown"),
|
||||
"pm2_restarts": pm2.get("restarts", "?"),
|
||||
"pm2_uptime": pm2.get("uptime", "?"),
|
||||
}
|
||||
|
||||
# Tanko (CT 122)
|
||||
tanko_state = ssh_jerome("192.168.68.122", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
||||
tanko_data = {}
|
||||
try:
|
||||
tanko_data = json.loads(tanko_state) if tanko_state else {}
|
||||
except:
|
||||
tanko_data = {}
|
||||
platforms = tanko_data.get("platforms", {})
|
||||
# Tanko (CT 112, IP 192.168.68.122) — DSH (DeepSeek Harness), no Hermes gateway
|
||||
# since 2026-08-27. There is no ~/.hermes/gateway_state.json on CT 112 anymore;
|
||||
# Zulip/Telegram connectivity is managed by the DSH harness, not the Hermes gateway.
|
||||
report["agents"]["tanko"] = {
|
||||
"platform": "hermes", "ct": 112, "ip": "192.168.68.122",
|
||||
"gateway_state": tanko_data.get("gateway_state", "unknown"),
|
||||
"zulip_state": platforms.get("zulip", {}).get("state", "unknown"),
|
||||
"telegram_state": platforms.get("telegram", {}).get("state", "unknown"),
|
||||
"gateway_pid": tanko_data.get("pid"),
|
||||
"updated_at": tanko_data.get("updated_at"),
|
||||
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
|
||||
"gateway_state": "n/a (DSH)",
|
||||
"zulip_state": "unknown",
|
||||
"telegram_state": "unknown",
|
||||
"gateway_pid": None,
|
||||
"updated_at": "",
|
||||
}
|
||||
|
||||
# Mumuni (CT 100, IP 192.168.68.24)
|
||||
mumuni_state = ssh("192.168.68.24", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
||||
mumuni_data = {}
|
||||
try:
|
||||
mumuni_data = json.loads(mumuni_state) if mumuni_state else {}
|
||||
except:
|
||||
mumuni_data = {}
|
||||
mumuni_platforms = mumuni_data.get("platforms", {})
|
||||
report["agents"]["mumuni"] = {
|
||||
"platform": "hermes", "ct": 114, "ip": "192.168.68.24",
|
||||
"gateway_state": mumuni_data.get("gateway_state", "unknown"),
|
||||
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
|
||||
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
|
||||
"email_state": mumuni_platforms.get("email", {}).get("state", "unknown"),
|
||||
"hermes_version": "",
|
||||
}
|
||||
# Get Hermes version
|
||||
ver = ssh("192.168.68.24", "hermes --version 2>/dev/null | head -1")
|
||||
if ver:
|
||||
report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip()
|
||||
# Mumuni is deliberately absent from this digest: captain ruling 2026-09-10.
|
||||
# She moved off this host onto her own container (kagentz CT 105 on minipve,
|
||||
# 192.168.68.14, dedicated `hermes` user) and is monitored from her side. The
|
||||
# former probe ssh'd to 192.168.68.24 for the decommissioned deployment's
|
||||
# ~/.hermes/gateway_state.json, always read "unknown", and published a false
|
||||
# "mumuni:unknown" line in the agent table and the gateway-unknown issue
|
||||
# count of every digest. Do NOT re-add an .24 / gateway_state probe.
|
||||
|
||||
return report
|
||||
|
||||
@@ -424,7 +444,7 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
<div class="alert {'good' if not issues else 'bad' if any('🔴' in i for i in issues) else 'warn'}">
|
||||
<p style="margin:0;font-size:16px"><b>{status}</b></p>
|
||||
<p style="margin:4px 0 0 0;font-size:13px">
|
||||
{r['node_count']} PVE nodes · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
||||
Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']}) · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
||||
{r['docker_vm']['total'] + r['docker_syslog']['total'] + r['docker_netbird']['total']} containers ·
|
||||
{len(r['endpoints'])} endpoints · {len(r.get('agents',{}))} agents
|
||||
</p>
|
||||
@@ -440,8 +460,10 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
|
||||
# ── Quick Stats ──
|
||||
html += '<div class="card"><h2>📊 Quick Stats</h2><div class="grid">'
|
||||
pve_status_label = "unreachable" if r.get('pve_probe_status') == 'unreachable' else f"{r['nodes_online']}/{r['node_count']}"
|
||||
pve_status_color = "red" if r.get('pve_probe_status') == 'unreachable' or r['nodes_online'] != r['node_count'] else "green"
|
||||
stats = [
|
||||
("PVE Nodes", f"{r['nodes_online']}/{r['node_count']}", "green" if r['nodes_online'] == r['node_count'] else "red"),
|
||||
("PVE Nodes", pve_status_label, pve_status_color),
|
||||
("VMs/CTs", f"{r['running_vms']}/{r['total_vms']}", "green" if r['running_vms'] == r['total_vms'] else "red"),
|
||||
("Containers", f"{r['docker_vm']['running']}/{r['docker_vm']['total']}", "green" if r['docker_vm']['running'] == r['docker_vm']['total'] else "yellow"),
|
||||
("LiteLLM Ctrs", f"{r['docker_syslog']['running']}/{r['docker_syslog']['total']}", "green" if r['docker_syslog']['running'] == r['docker_syslog']['total'] else "red"),
|
||||
@@ -543,13 +565,7 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
elif name == "tanko":
|
||||
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
processed = agent.get("updated_at", "")[:10]
|
||||
elif name == "mumuni":
|
||||
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
tg = "✅" if agent.get("telegram_state") == "connected" else "❌"
|
||||
ver = agent.get("hermes_version", "")
|
||||
processed = f"TG:{tg} v{ver}"
|
||||
processed = "DSH"
|
||||
else:
|
||||
zulip_state = "⬜"
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
@@ -657,7 +673,11 @@ def send_email(html_content, subject_prefix=""):
|
||||
msg.attach(MIMEText(html_content, "html"))
|
||||
|
||||
try:
|
||||
EMAIL_PASSWORD = "rgbuomwcydxwbszd"
|
||||
EMAIL_PASSWORD = os.environ.get("EMAIL_PASSWORD") or os.environ.get("SMTP_PASSWORD") or os.environ.get("MAIL_PASSWORD")
|
||||
if not EMAIL_PASSWORD:
|
||||
print(" ⚠️ Degraded leg: credential-missing: EMAIL_PASSWORD (or SMTP_PASSWORD/MAIL_PASSWORD)", file=sys.stderr)
|
||||
DEGRADED_LEGS.append("credential-missing: EMAIL_PASSWORD")
|
||||
return True, "✅ Email leg degraded (no credential) — report still produced"
|
||||
GMAIL_EMAIL = "jtabiri@gmail.com"
|
||||
|
||||
server = smtplib.SMTP("smtp.gmail.com", 587)
|
||||
@@ -696,6 +716,17 @@ if __name__ == "__main__":
|
||||
print(f" {msg}")
|
||||
|
||||
# Show summary
|
||||
if DEGRADED_LEGS:
|
||||
print(f"\n⚠️ Degraded legs ({len(DEGRADED_LEGS)}):")
|
||||
for leg in DEGRADED_LEGS:
|
||||
print(f" - {leg}")
|
||||
else:
|
||||
print("\n✅ All legs fully credentialed")
|
||||
|
||||
# A failed send must exit non-zero; a degraded leg (no credential) must stay exit 0
|
||||
if not ok:
|
||||
sys.exit(1)
|
||||
|
||||
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
|
||||
print(f"\n📋 Summary:")
|
||||
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
|
||||
@@ -703,6 +734,6 @@ if __name__ == "__main__":
|
||||
print(f" Zulip Ext: {'✅' if report.get('zulip_ext',{}).get('connected') else '❌'}")
|
||||
print(f" LiteLLM: {sum(1 for c in report.get('litellm',{}).get('checks',[]) if c['status']=='pass')}/{len(report.get('litellm',{}).get('checks',[]))} checks pass")
|
||||
agent_parts = []
|
||||
for k,v in report.get('agents',{}).items():
|
||||
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
||||
print(f" Agents: {', '.join(agent_parts)}")
|
||||
for k,v in report.get('agents',{}).items():
|
||||
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
||||
print(f" Agents: {', '.join(agent_parts)}")
|
||||
|
||||
Executable
+206
@@ -0,0 +1,206 @@
|
||||
#!/usr/bin/env python3
|
||||
"""disk-gc-plan — turn a fleet disk scan into the GC action plan.
|
||||
|
||||
This is the executable side of `disk-gc-threat-response.prose.md`. It exists so the
|
||||
report-only gate is enforced by code that can be tested, rather than by prose the
|
||||
executor might misread.
|
||||
|
||||
THE HARD GATE: guests listed in the contract's `report_only_guests` block are
|
||||
DETECT-AND-REPORT-ONLY at EVERY level (AMBER, RED, CRITICAL). This tool will never
|
||||
emit a `gc-executor` action for one, so no GC command can be constructed for it.
|
||||
|
||||
The gate is keyed on GUEST identity — guest id, hostname, or IP — never on an agent
|
||||
name. An agent-name marker can silently miss the guest it lives on; a guest marker
|
||||
cannot.
|
||||
|
||||
The authoritative exclusion list lives in the contract itself (the fenced ```yaml
|
||||
block containing `report_only_guests:`). This tool reads it from there so there is
|
||||
only ever one copy.
|
||||
|
||||
Usage:
|
||||
disk-gc-plan.py --scan scan.json # [{"id":111,"usage_pct":84}, ...]
|
||||
cat scan.json | disk-gc-plan.py # same, via stdin
|
||||
disk-gc-plan.py --scan scan.json --json # machine-readable plan
|
||||
|
||||
Scan entries may carry any of: id / guest / vmid / ct / ctid, hostname / name, ip.
|
||||
A threshold-crossing entry with no recognizable identity is reported, never GC'd.
|
||||
Exit codes: 0 ok, 1 usage/parse error.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import pathlib
|
||||
import re
|
||||
import sys
|
||||
|
||||
import yaml
|
||||
|
||||
REPO = pathlib.Path(__file__).resolve().parent.parent
|
||||
DEFAULT_CONTRACT = REPO / "disk-gc-threat-response.prose.md"
|
||||
|
||||
AMBER, RED, CRITICAL = 75, 85, 95
|
||||
|
||||
IDENTITY_FIELDS = ("id", "guest", "vmid", "ct", "ctid", "hostname", "name", "ip")
|
||||
IDENTITY_TYPE_PREFIX = re.compile(r"^(?:lxc|qemu)/")
|
||||
UNIDENTIFIED_REASON = "unidentified target - refusing to schedule GC"
|
||||
|
||||
|
||||
def load_report_only_guests(contract_path: pathlib.Path) -> list[dict]:
|
||||
"""Read the authoritative report_only_guests block out of the contract.
|
||||
|
||||
The contract carries it as a fenced ```yaml block. Parsing the declared,
|
||||
machine-readable block is the intended interface — the contract owns the list.
|
||||
"""
|
||||
text = contract_path.read_text(encoding="utf-8")
|
||||
for block in re.findall(r"```yaml\n(.*?)```", text, re.S):
|
||||
if "report_only_guests:" in block:
|
||||
data = yaml.safe_load(block)
|
||||
guests = data.get("report_only_guests") or []
|
||||
if not isinstance(guests, list):
|
||||
raise SystemExit("report_only_guests must be a list")
|
||||
if not guests:
|
||||
raise SystemExit(
|
||||
"report_only_guests is empty or missing - refusing to plan GC "
|
||||
"without the report-only gate"
|
||||
)
|
||||
for guest in guests:
|
||||
if not isinstance(guest, dict) or not _keys(guest):
|
||||
raise SystemExit(
|
||||
"report_only_guests entry has no recognizable identity key "
|
||||
f"(expected one of: {', '.join(IDENTITY_FIELDS)}): {guest!r}"
|
||||
)
|
||||
return guests
|
||||
raise SystemExit(
|
||||
f"no authoritative report_only_guests block found in {contract_path}"
|
||||
)
|
||||
|
||||
|
||||
def _canonical_number(number: float) -> str:
|
||||
if float(number).is_integer():
|
||||
return str(int(number))
|
||||
return str(number).strip().lower()
|
||||
|
||||
|
||||
def _normalize_identity(value: object) -> str:
|
||||
"""Canonicalise a guest identity so differently-encoded ids compare equal:
|
||||
numeric and numeric-string ids collapse to an integer string, Proxmox
|
||||
type prefixes and leading zeros are stripped, and hostnames/IPs are only
|
||||
trimmed and lowercased."""
|
||||
if isinstance(value, bool):
|
||||
return str(value).strip().lower()
|
||||
if isinstance(value, (int, float)):
|
||||
return _canonical_number(float(value))
|
||||
text = str(value).strip().lower()
|
||||
text = IDENTITY_TYPE_PREFIX.sub("", text)
|
||||
try:
|
||||
return _canonical_number(float(text))
|
||||
except ValueError:
|
||||
return text
|
||||
|
||||
|
||||
def _keys(entry: dict) -> set[str]:
|
||||
"""Guest/host identity keys, shared by exclusions and scan entries so the two
|
||||
sides of the gate can never key on different fields."""
|
||||
out: set[str] = set()
|
||||
for field in IDENTITY_FIELDS:
|
||||
value = entry.get(field)
|
||||
if value is None:
|
||||
continue
|
||||
key = _normalize_identity(value)
|
||||
if key:
|
||||
out.add(key)
|
||||
return out
|
||||
|
||||
|
||||
def level_for(pct: float) -> str | None:
|
||||
if pct >= CRITICAL:
|
||||
return "CRITICAL"
|
||||
if pct >= RED:
|
||||
return "RED"
|
||||
if pct >= AMBER:
|
||||
return "AMBER"
|
||||
return None
|
||||
|
||||
|
||||
def build_plan(scan: list[dict], report_only: list[dict]) -> list[dict]:
|
||||
excluded = [(e, _keys(e)) for e in report_only]
|
||||
plan: list[dict] = []
|
||||
for entry in scan:
|
||||
pct = entry.get("usage_pct")
|
||||
if pct is None:
|
||||
continue
|
||||
level = level_for(float(pct))
|
||||
if level is None:
|
||||
continue # GREEN: log only, no action
|
||||
scan_keys = _keys(entry)
|
||||
target = next(
|
||||
(entry.get(k) for k in IDENTITY_FIELDS if entry.get(k) not in (None, "")),
|
||||
"?",
|
||||
)
|
||||
if not scan_keys:
|
||||
plan.append({
|
||||
"target": target,
|
||||
"level": level,
|
||||
"pct": float(pct),
|
||||
"action": "report-only",
|
||||
"reason": UNIDENTIFIED_REASON,
|
||||
})
|
||||
continue
|
||||
match = next((e for e, keys in excluded if keys & scan_keys), None)
|
||||
if match is not None:
|
||||
plan.append({
|
||||
"target": target,
|
||||
"level": level,
|
||||
"pct": float(pct),
|
||||
"action": "report-only",
|
||||
"reason": match.get("reason", "").strip(),
|
||||
})
|
||||
else:
|
||||
plan.append({
|
||||
"target": target,
|
||||
"level": level,
|
||||
"pct": float(pct),
|
||||
"action": "gc-executor",
|
||||
})
|
||||
plan.sort(key=lambda row: row["pct"], reverse=True)
|
||||
return plan
|
||||
|
||||
|
||||
def main() -> int:
|
||||
ap = argparse.ArgumentParser(description="Plan disk GC actions with the report-only gate.")
|
||||
ap.add_argument("--scan", help="JSON file: list of {id|ct|hostname|ip, usage_pct}")
|
||||
ap.add_argument("--contract", default=str(DEFAULT_CONTRACT))
|
||||
ap.add_argument("--json", action="store_true", help="emit the plan as JSON")
|
||||
args = ap.parse_args()
|
||||
|
||||
raw = pathlib.Path(args.scan).read_text() if args.scan else sys.stdin.read()
|
||||
try:
|
||||
scan = json.loads(raw)
|
||||
except json.JSONDecodeError as exc:
|
||||
print(f"invalid scan JSON: {exc}", file=sys.stderr)
|
||||
return 1
|
||||
if not isinstance(scan, list):
|
||||
print("scan must be a JSON list", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
report_only = load_report_only_guests(pathlib.Path(args.contract))
|
||||
plan = build_plan(scan, report_only)
|
||||
|
||||
if args.json:
|
||||
print(json.dumps(plan, indent=2))
|
||||
return 0
|
||||
|
||||
if not plan:
|
||||
print("no threats (nothing at or above 75%)")
|
||||
return 0
|
||||
for row in plan:
|
||||
if row["action"] == "report-only":
|
||||
print(f" {row['target']} {row['level']} {row['pct']}% -> REPORT-ONLY (no GC) — {row['reason']}")
|
||||
else:
|
||||
print(f" {row['target']} {row['level']} {row['pct']}% -> gc-executor")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,569 @@
|
||||
#!/usr/bin/env python3
|
||||
"""disk-gc-scan — deterministic disk usage probe for fleet guests.
|
||||
|
||||
This is the executable scanner side of `disk-gc-threat-response.prose.md`. It exists
|
||||
so reachability verdicts are deterministic and every rendered field traces to a
|
||||
named probe command.
|
||||
|
||||
DESIGN PRINCIPLES (per task disk-gc-probe-false-unreachable-20260913):
|
||||
1. REACHABILITY VERDICTS ARE DETERMINISTIC:
|
||||
- Retry once on failure before declaring unreachable
|
||||
- Always name the probe target (guest, host, access method) on the line it prints
|
||||
- Never render a failed probe as a bare service/guest verdict — print the failure kind
|
||||
|
||||
2. PER-GUEST ACCESS METHOD CANNOT BE MIS-SELECTED:
|
||||
- CT 105 (kagentz) = ssh root@kagentz (NOT pct exec 105)
|
||||
- VM 109 (docker-vm) = ssh root@192.168.68.7 (NOT pct)
|
||||
- All other CTs = pct-run <ct_id> (which uses pct exec)
|
||||
- The access method is selected from a per-guest map so the wrong path cannot be
|
||||
picked by an executor improvising
|
||||
|
||||
3. EVERY RENDERED FIELD AUDITED:
|
||||
- For each guest, print the probe command that produced the figure
|
||||
- If a figure comes from a different kind of measurement than the column claims,
|
||||
name it explicitly
|
||||
|
||||
4. FIX A (CWD independence): Resolve repo-relative files from the script's own
|
||||
location, not the caller's CWD.
|
||||
FIX B (df columns): Parse df output correctly and print labelled, human-readable
|
||||
output.
|
||||
|
||||
Usage:
|
||||
disk-gc-scan.py # scan all guests
|
||||
disk-gc-scan.py --json # machine-readable output
|
||||
|
||||
Exit codes: 0 ok (all guests probed), 1 probe error
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import os
|
||||
from dataclasses import dataclass
|
||||
from typing import Optional
|
||||
|
||||
# Resolve repo-relative files from the script's own location, not the caller's CWD
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
HELPER_PCT_RUN = SCRIPT_DIR / "pct-run.sh"
|
||||
|
||||
# Per-guest access method map. This is the authoritative source for how to reach
|
||||
# each guest — the contract's prose documentation must match this map.
|
||||
#
|
||||
# Access methods:
|
||||
# - "pct-run": use pct-run.sh <ct_id> (pct exec via SSH to node)
|
||||
# - "ssh-host": use ssh root@<hostname>
|
||||
# - "ssh-ip": use ssh root@<ip>
|
||||
|
||||
@dataclass
|
||||
class Guest:
|
||||
"""A guest to probe."""
|
||||
ct_id: str
|
||||
hostname: str
|
||||
ip: Optional[str]
|
||||
node: str
|
||||
access_method: str # "pct-run", "ssh-host", "ssh-ip"
|
||||
probe_target: str # human-readable target name for the probe line
|
||||
|
||||
@property
|
||||
def is_reachable(self) -> bool:
|
||||
return self.probe_result is not None and self.probe_result.exit_code == 0
|
||||
|
||||
@property
|
||||
def usage_pct(self) -> Optional[float]:
|
||||
return self.probe_result.usage_pct if self.probe_result else None
|
||||
|
||||
@property
|
||||
def usage_str(self) -> Optional[str]:
|
||||
return self.probe_result.usage_str if self.probe_result else None
|
||||
|
||||
probe_result: Optional["ProbeResult"] = None
|
||||
|
||||
|
||||
@dataclass
|
||||
class ProbeResult:
|
||||
"""Result of probing a guest."""
|
||||
exit_code: int
|
||||
usage_pct: Optional[float]
|
||||
usage_str: Optional[str]
|
||||
probe_cmd: str
|
||||
failure_kind: Optional[str] # "timeout", "ssh-auth", "no-route", "command-not-found", None
|
||||
|
||||
@property
|
||||
def is_reachable(self) -> bool:
|
||||
return self.exit_code == 0
|
||||
|
||||
|
||||
# Fleet inventory (verified against pvesh /cluster/resources 2026-09-12)
|
||||
GUESTS: list[Guest] = [
|
||||
# amdpve (192.168.68.15)
|
||||
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
|
||||
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
|
||||
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
|
||||
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
|
||||
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
|
||||
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
|
||||
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
|
||||
access_method="pct-run", probe_target="scottdenya (CT 115, amdpve)"),
|
||||
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
|
||||
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
|
||||
# minipve (192.168.68.12)
|
||||
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
|
||||
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
|
||||
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
|
||||
access_method="pct-run", probe_target="adguard (CT 102, minipve)"),
|
||||
Guest(ct_id="104", hostname="authentik", ip="192.168.68.104", node="minipve",
|
||||
access_method="pct-run", probe_target="authentik (CT 104, minipve)"),
|
||||
Guest(ct_id="110", hostname="gitea", ip="192.168.68.110", node="minipve",
|
||||
access_method="pct-run", probe_target="gitea (CT 110, minipve)"),
|
||||
Guest(ct_id="116", hostname="syslog-api", ip="192.168.68.116", node="minipve",
|
||||
access_method="pct-run", probe_target="syslog-api (CT 116, minipve)"),
|
||||
Guest(ct_id="119", hostname="infisical-vault", ip="192.168.68.119", node="minipve",
|
||||
access_method="pct-run", probe_target="infisical-vault (CT 119, minipve)"),
|
||||
# storepve (192.168.68.6)
|
||||
Guest(ct_id="106", hostname="ra-h-os", ip="192.168.68.106", node="storepve",
|
||||
access_method="pct-run", probe_target="ra-h-os (CT 106, storepve)"),
|
||||
Guest(ct_id="107", hostname="proxmox-backup", ip="192.168.68.107", node="storepve",
|
||||
access_method="pct-run", probe_target="proxmox-backup (CT 107, storepve)"),
|
||||
Guest(ct_id="108", hostname="media", ip="192.168.68.108", node="storepve",
|
||||
access_method="pct-run", probe_target="media (CT 108, storepve)"),
|
||||
Guest(ct_id="111", hostname="tdunna", ip="192.168.68.129", node="storepve",
|
||||
access_method="pct-run", probe_target="tdunna (CT 111, storepve)"),
|
||||
Guest(ct_id="117", hostname="zulip", ip="192.168.68.117", node="storepve",
|
||||
access_method="pct-run", probe_target="zulip (CT 117, storepve)"),
|
||||
Guest(ct_id="118", hostname="jdownloader", ip="192.168.68.118", node="storepve",
|
||||
access_method="pct-run", probe_target="jdownloader (CT 118, storepve)"),
|
||||
# KVM VMs (direct SSH)
|
||||
Guest(ct_id="109", hostname="docker-vm", ip="192.168.68.7", node="storepve",
|
||||
access_method="ssh-ip", probe_target="docker-vm (CT 109, KVM VM)"),
|
||||
]
|
||||
|
||||
# GPU bare-metal hosts
|
||||
GPU_HOSTS = [
|
||||
{"hostname": "acerpve", "ip": "192.168.68.9", "gpu": "RTX 3090",
|
||||
"probe_target": "RTX 3090 (bare metal .9)"},
|
||||
{"hostname": "ocupve", "ip": "192.168.68.110", "gpu": "RTX 5070",
|
||||
"probe_target": "RTX 5070 (bare metal .110)"},
|
||||
{"hostname": "amdpve", "ip": "192.168.68.15", "gpu": "Strix Halo",
|
||||
"probe_target": "Strix Halo (bare metal .15)"},
|
||||
]
|
||||
|
||||
CONNECT_TIMEOUT = 5
|
||||
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
|
||||
|
||||
# Host filesystem thresholds (from contract)
|
||||
HOST_THRESHOLDS = {
|
||||
"WARN": 85,
|
||||
"AMBER": 90,
|
||||
"RED": 95,
|
||||
}
|
||||
|
||||
# State file path (absolute, so execution context doesn't matter)
|
||||
STATE_FILE = pathlib.Path(__file__).resolve().parent.parent / "state" / "host-disk-bands.json"
|
||||
|
||||
# PVE nodes to probe for host filesystems
|
||||
HOST_NODES = [
|
||||
{"hostname": "acerpve", "ip": "192.168.68.9"},
|
||||
{"hostname": "amdpve", "ip": "192.168.68.15"},
|
||||
{"hostname": "storepve", "ip": "192.168.68.6"},
|
||||
{"hostname": "minipve", "ip": "192.168.68.12"},
|
||||
{"hostname": "ocupve", "ip": "192.168.68.5"},
|
||||
]
|
||||
|
||||
|
||||
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
|
||||
"""Run a command and return (exit_code, stdout, stderr)."""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
cmd, shell=True, capture_output=True, text=True, timeout=timeout
|
||||
)
|
||||
return result.returncode, result.stdout.strip(), result.stderr.strip()
|
||||
except subprocess.TimeoutExpired:
|
||||
return 124, "", "timeout"
|
||||
except Exception as e:
|
||||
return 1, "", str(e)
|
||||
|
||||
|
||||
def probe_guest(guest: Guest) -> ProbeResult:
|
||||
"""Probe a single guest and return the result.
|
||||
|
||||
Access method is selected from guest.access_method:
|
||||
- "pct-run": pct-run.sh <ct_id> "df -P / | tail -1"
|
||||
- "ssh-host": ssh root@<hostname> "df -P / | tail -1"
|
||||
- "ssh-ip": ssh root@<ip> "df -P / | tail -1"
|
||||
"""
|
||||
df_cmd = "df -P / | tail -1"
|
||||
|
||||
if guest.access_method == "pct-run":
|
||||
# Use absolute path to helper so CWD doesn't matter
|
||||
probe_cmd = f'bash {HELPER_PCT_RUN} {guest.ct_id} "{df_cmd}"'
|
||||
elif guest.access_method == "ssh-host":
|
||||
probe_cmd = f'ssh {SSH_OPTS} root@{guest.hostname} "{df_cmd}"'
|
||||
elif guest.access_method == "ssh-ip":
|
||||
probe_cmd = f'ssh {SSH_OPTS} root@{guest.ip} "{df_cmd}"'
|
||||
else:
|
||||
raise ValueError(f"unknown access_method: {guest.access_method}")
|
||||
|
||||
# Check helper exists and is readable BEFORE probing (for pct-run guests)
|
||||
# This prevents scanner errors from being rendered as guest verdicts
|
||||
if guest.access_method == "pct-run":
|
||||
if not HELPER_PCT_RUN.exists():
|
||||
print(f"SCANNER ERROR: helper not found: {HELPER_PCT_RUN}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
if not os.access(str(HELPER_PCT_RUN), os.R_OK):
|
||||
print(f"SCANNER ERROR: helper not readable: {HELPER_PCT_RUN}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
# Retry once on failure before declaring unreachable
|
||||
for attempt in range(2):
|
||||
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
|
||||
|
||||
if exit_code == 0:
|
||||
# Parse df output: Filesystem 1024-blocks Used Available Capacity Mounted on
|
||||
# parts[0]=Filesystem, parts[1]=Total (1K blocks), parts[2]=Used, parts[3]=Available, parts[4]=Capacity
|
||||
parts = stdout.split()
|
||||
if len(parts) >= 5:
|
||||
capacity_str = parts[4] # e.g., "34%"
|
||||
usage_pct = float(capacity_str.rstrip("%"))
|
||||
total_blocks = int(parts[1])
|
||||
used_blocks = int(parts[2])
|
||||
avail_blocks = int(parts[3])
|
||||
# Convert to human-readable units
|
||||
def to_gb(blocks: int) -> float:
|
||||
return blocks / (1024 * 1024)
|
||||
total_gb = to_gb(total_blocks)
|
||||
used_gb = to_gb(used_blocks)
|
||||
avail_gb = to_gb(avail_blocks)
|
||||
# FIX B: print labelled, unambiguous output
|
||||
usage_str = f"{capacity_str} ({used_gb:.1f}G used of {total_gb:.1f}G total, {avail_gb:.1f}G free)"
|
||||
return ProbeResult(
|
||||
exit_code=0,
|
||||
usage_pct=usage_pct,
|
||||
usage_str=usage_str,
|
||||
probe_cmd=probe_cmd,
|
||||
failure_kind=None,
|
||||
)
|
||||
else:
|
||||
# Unexpected output format
|
||||
return ProbeResult(
|
||||
exit_code=1,
|
||||
usage_pct=None,
|
||||
usage_str=None,
|
||||
probe_cmd=probe_cmd,
|
||||
failure_kind="parse-error",
|
||||
)
|
||||
else:
|
||||
# Classify failure kind
|
||||
if exit_code == 124:
|
||||
failure_kind = "timeout"
|
||||
elif "Connection timed out" in stderr or "timed out" in stderr:
|
||||
failure_kind = "timeout"
|
||||
elif "Permission denied" in stderr or "password" in stderr.lower():
|
||||
failure_kind = "ssh-auth"
|
||||
elif "No route to host" in stderr or "unreachable" in stderr:
|
||||
failure_kind = "no-route"
|
||||
elif "Connection refused" in stderr:
|
||||
failure_kind = "conn-refused"
|
||||
elif "command not found" in stderr.lower() or "No such file" in stderr:
|
||||
failure_kind = "command-not-found"
|
||||
else:
|
||||
failure_kind = f"ssh-exit-{exit_code}"
|
||||
|
||||
# Retry once
|
||||
if attempt == 0:
|
||||
time.sleep(1)
|
||||
continue
|
||||
return ProbeResult(
|
||||
exit_code=exit_code,
|
||||
usage_pct=None,
|
||||
usage_str=None,
|
||||
probe_cmd=probe_cmd,
|
||||
failure_kind=failure_kind,
|
||||
)
|
||||
|
||||
# Should not reach here, but just in case
|
||||
return ProbeResult(
|
||||
exit_code=1,
|
||||
usage_pct=None,
|
||||
usage_str=None,
|
||||
probe_cmd=probe_cmd,
|
||||
failure_kind="unknown",
|
||||
)
|
||||
|
||||
|
||||
def scan_fleet() -> list[dict]:
|
||||
"""Scan all guests and return the results."""
|
||||
results = []
|
||||
for guest in GUESTS:
|
||||
probe_result = probe_guest(guest)
|
||||
guest.probe_result = probe_result
|
||||
|
||||
row = {
|
||||
"target": guest.probe_target,
|
||||
"ct_id": guest.ct_id,
|
||||
"hostname": guest.hostname,
|
||||
"node": guest.node,
|
||||
"access_method": guest.access_method,
|
||||
"reachable": probe_result.is_reachable,
|
||||
"usage_pct": probe_result.usage_pct,
|
||||
"usage_str": probe_result.usage_str,
|
||||
"probe_cmd": probe_result.probe_cmd,
|
||||
"failure_kind": probe_result.failure_kind,
|
||||
}
|
||||
results.append(row)
|
||||
|
||||
return results
|
||||
|
||||
|
||||
def classify_band(usage_pct: float) -> str:
|
||||
"""Classify a percentage into a band."""
|
||||
if usage_pct >= HOST_THRESHOLDS["RED"]:
|
||||
return "HOST-RED"
|
||||
elif usage_pct >= HOST_THRESHOLDS["AMBER"]:
|
||||
return "HOST-AMBER"
|
||||
elif usage_pct >= HOST_THRESHOLDS["WARN"]:
|
||||
return "HOST-WARN"
|
||||
else:
|
||||
return "GREEN"
|
||||
|
||||
|
||||
def probe_host_filesystems() -> tuple[list[dict], dict[str, str]]:
|
||||
"""Probe host filesystems on all PVE nodes.
|
||||
|
||||
Returns:
|
||||
- List of host filesystem results
|
||||
- Dict of volume_key -> current_band (for state file)
|
||||
"""
|
||||
results = []
|
||||
current_bands = {}
|
||||
|
||||
for node in HOST_NODES:
|
||||
ip = node["ip"]
|
||||
hostname = node["hostname"]
|
||||
|
||||
# Probe df for host filesystems
|
||||
probe_cmd = f'ssh {SSH_OPTS} root@{ip} "df -hP / /media/* tank 2>/dev/null | tail -n +2"'
|
||||
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
|
||||
|
||||
if exit_code != 0:
|
||||
results.append({
|
||||
"target": f"{hostname} ({ip})",
|
||||
"hostname": hostname,
|
||||
"ip": ip,
|
||||
"reachable": False,
|
||||
"volumes": [],
|
||||
"probe_cmd": probe_cmd,
|
||||
"failure_kind": "ssh-error",
|
||||
})
|
||||
continue
|
||||
|
||||
# Parse df output and classify each volume
|
||||
volumes = []
|
||||
for line in stdout.splitlines():
|
||||
parts = line.split()
|
||||
if len(parts) < 6:
|
||||
continue
|
||||
|
||||
dev, size, used, avail, pct_str, mount = parts[:6]
|
||||
pct = float(pct_str.rstrip("%"))
|
||||
band = classify_band(pct)
|
||||
|
||||
# Volume type classification
|
||||
if mount.startswith("/media/"):
|
||||
vol_type = "media"
|
||||
elif mount == "/" or "pve" in dev:
|
||||
vol_type = "host-root"
|
||||
elif mount == "tank" or "tank" in mount:
|
||||
vol_type = "pbs-datastore"
|
||||
else:
|
||||
vol_type = "other"
|
||||
|
||||
# Volume key for state file (host/volume)
|
||||
volume_key = f"{hostname}/{mount}"
|
||||
current_bands[volume_key] = band
|
||||
|
||||
volumes.append({
|
||||
"mount": mount,
|
||||
"device": dev,
|
||||
"size": size,
|
||||
"used": used,
|
||||
"avail": avail,
|
||||
"pct": pct,
|
||||
"band": band,
|
||||
"type": vol_type,
|
||||
})
|
||||
|
||||
results.append({
|
||||
"target": f"{hostname} ({ip})",
|
||||
"hostname": hostname,
|
||||
"ip": ip,
|
||||
"reachable": True,
|
||||
"volumes": volumes,
|
||||
"probe_cmd": probe_cmd,
|
||||
"failure_kind": None,
|
||||
})
|
||||
|
||||
return results, current_bands
|
||||
|
||||
|
||||
def read_state_file() -> Optional[dict[str, str]]:
|
||||
"""Read the state file if it exists."""
|
||||
if not STATE_FILE.exists():
|
||||
return None
|
||||
try:
|
||||
with open(STATE_FILE) as f:
|
||||
return json.load(f)
|
||||
except (json.JSONDecodeError, IOError) as e:
|
||||
print(f"⚠️ State file exists but unreadable: {e}", file=sys.stderr)
|
||||
return {}
|
||||
|
||||
|
||||
def write_state_file(bands: dict[str, str]) -> None:
|
||||
"""Write the state file."""
|
||||
STATE_FILE.parent.mkdir(parents=True, exist_ok=True)
|
||||
try:
|
||||
with open(STATE_FILE, "w") as f:
|
||||
json.dump(bands, f, indent=2)
|
||||
except IOError as e:
|
||||
print(f"⚠️ State file write failed: {e}", file=sys.stderr)
|
||||
|
||||
|
||||
def detect_transitions(current_bands: dict[str, str], prior_bands: Optional[dict[str, str]]) -> list[dict]:
|
||||
"""Detect band transitions (current vs. prior)."""
|
||||
if prior_bands is None:
|
||||
# First run — no transitions, just establish baseline
|
||||
return []
|
||||
|
||||
transitions = []
|
||||
# Check for volumes that moved to a higher band (escalation)
|
||||
for volume, current_band in current_bands.items():
|
||||
prior_band = prior_bands.get(volume, "GREEN")
|
||||
|
||||
# Band ordering: GREEN < HOST-WARN < HOST-AMBER < HOST-RED
|
||||
band_order = {"GREEN": 0, "HOST-WARN": 1, "HOST-AMBER": 2, "HOST-RED": 3}
|
||||
|
||||
if band_order[current_band] > band_order[prior_band]:
|
||||
transitions.append({
|
||||
"type": "escalation",
|
||||
"volume": volume,
|
||||
"from": prior_band,
|
||||
"to": current_band,
|
||||
})
|
||||
elif band_order[current_band] < band_order[prior_band]:
|
||||
transitions.append({
|
||||
"type": "recovery",
|
||||
"volume": volume,
|
||||
"from": prior_band,
|
||||
"to": current_band,
|
||||
})
|
||||
|
||||
return transitions
|
||||
|
||||
|
||||
def render_results(results: list[dict]) -> str:
|
||||
"""Render scan results in human-readable format."""
|
||||
lines = []
|
||||
lines.append("=== Disk GC Scan ===")
|
||||
lines.append("")
|
||||
|
||||
for row in results:
|
||||
if row["reachable"]:
|
||||
lines.append(f" ✅ {row['target']}: {row['usage_str']}")
|
||||
lines.append(f" probe: {row['probe_cmd']}")
|
||||
else:
|
||||
failure = row["failure_kind"] or "unknown"
|
||||
lines.append(f" ❌ {row['target']}: UNREACHABLE ({failure})")
|
||||
lines.append(f" probe: {row['probe_cmd']}")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def render_host_results(results: list[dict], transitions: list[dict], prior_bands: Optional[dict[str, str]]) -> str:
|
||||
"""Render host filesystem results in human-readable format."""
|
||||
lines = []
|
||||
lines.append("")
|
||||
lines.append("=== Host Filesystem Bands ===")
|
||||
lines.append("")
|
||||
|
||||
# Render transitions first (they're the actionable alerts)
|
||||
if prior_bands is None:
|
||||
lines.append(" (first run — recording baseline, no alerts)")
|
||||
elif not transitions:
|
||||
lines.append(" (no band changes since last scan)")
|
||||
else:
|
||||
for t in transitions:
|
||||
volume, from_band, to_band = t["volume"], t["from"], t["to"]
|
||||
if t["type"] == "escalation":
|
||||
lines.append(f" ⚠️ {volume}: {from_band} → {to_band} (ESCALATION)")
|
||||
else:
|
||||
lines.append(f" ✅ {volume}: {from_band} → {to_band} (RECOVERY)")
|
||||
|
||||
# Render all volumes with their bands
|
||||
lines.append("")
|
||||
for node_result in results:
|
||||
if not node_result["reachable"]:
|
||||
lines.append(f" ❌ {node_result['target']}: UNREACHABLE ({node_result['failure_kind']})")
|
||||
continue
|
||||
|
||||
lines.append(f" {node_result['target']}:")
|
||||
for vol in node_result["volumes"]:
|
||||
lines.append(f" {vol['mount']} ({vol['type']}): {vol['pct']}% ({vol['used']}/{vol['size']}, {vol['avail']} free) -> {vol['band']}")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
import argparse
|
||||
|
||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests and host filesystems.")
|
||||
ap.add_argument("--json", action="store_true", help="machine-readable output")
|
||||
ap.add_argument("--hosts-only", action="store_true", help="scan host filesystems only")
|
||||
ap.add_argument("--guests-only", action="store_true", help="scan guests only (skip host filesystems)")
|
||||
args = ap.parse_args()
|
||||
|
||||
# Scan guests (unless --hosts-only)
|
||||
guest_results = []
|
||||
if not args.hosts_only:
|
||||
guest_results = scan_fleet()
|
||||
|
||||
# Scan host filesystems (unless --guests-only)
|
||||
host_results = []
|
||||
current_bands = {}
|
||||
if not args.guests_only:
|
||||
host_results, current_bands = probe_host_filesystems()
|
||||
|
||||
# Read prior state and detect transitions
|
||||
prior_bands = read_state_file()
|
||||
transitions = detect_transitions(current_bands, prior_bands)
|
||||
|
||||
# Write new state
|
||||
write_state_file(current_bands)
|
||||
else:
|
||||
prior_bands = None
|
||||
transitions = []
|
||||
|
||||
if args.json:
|
||||
# JSON output
|
||||
output = {
|
||||
"guests": guest_results,
|
||||
"hosts": host_results,
|
||||
"transitions": transitions,
|
||||
"prior_bands": prior_bands,
|
||||
}
|
||||
print(json.dumps(output, indent=2))
|
||||
else:
|
||||
# Human-readable output
|
||||
if guest_results:
|
||||
print(render_results(guest_results))
|
||||
|
||||
if host_results:
|
||||
print(render_host_results(host_results, transitions, prior_bands))
|
||||
|
||||
# Exit 0 if all probed (reachable or not), 1 if any probe error
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Executable
+32
@@ -0,0 +1,32 @@
|
||||
#!/bin/bash
|
||||
# Shared helper for Hermes contract reachability checks
|
||||
# Separates SSH exit status from remote command result
|
||||
|
||||
hermes_check_host() {
|
||||
local host=$1
|
||||
local pattern=$2
|
||||
local path=$3
|
||||
|
||||
# Remote side always succeeds (grep ...; true), so ssh exit code = connection status only
|
||||
local out
|
||||
out=$(ssh -o BatchMode=yes -o ConnectTimeout=3 root@"$host" "grep -RIn '$pattern' '$path' 2>/dev/null; true" 2>/dev/null)
|
||||
local status=$?
|
||||
|
||||
if [ $status -ne 0 ]; then
|
||||
echo "$host: UNREACHABLE (ssh exit $status)"
|
||||
elif [ -n "$out" ]; then
|
||||
echo "$host: VIOLATION: $out"
|
||||
else
|
||||
echo "$host: COMPLIANT (no matches found)"
|
||||
fi
|
||||
}
|
||||
|
||||
# Standalone mode: scripts/hermes-reachability-check.sh <host> <pattern> <path>
|
||||
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
|
||||
if [ $# -ne 3 ]; then
|
||||
echo "Usage: $0 <host> <pattern> <path>" >&2
|
||||
exit 2
|
||||
fi
|
||||
hermes_check_host "$1" "$2" "$3"
|
||||
exit 0
|
||||
fi
|
||||
Executable
+234
@@ -0,0 +1,234 @@
|
||||
#!/bin/bash
|
||||
# infrastructure-monitoring.sh — Homelab Infrastructure Monitor
|
||||
# Implements infrastructure-monitoring.prose.md (check-health section)
|
||||
#
|
||||
# Legs: Grafana, Prometheus, LiteLLM, PVE API (5 nodes), GPU exporters,
|
||||
# Docker Stats, PVE Exporter
|
||||
#
|
||||
# Design:
|
||||
# - Every target, port, path, and expected status is defined in code
|
||||
# - Liveness rule: any HTTP status = ALIVE for auth-gated/redirect endpoints;
|
||||
# only connection failures (000/timeout) = probe-failed
|
||||
# - Bare-200 rule: expected status must match exactly (200); anything else = alert
|
||||
# - PVE API uses -k flag (self-signed certs), probes /api2/json/version
|
||||
# - Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116, probed via SSH
|
||||
# with one retry at longer timeout (25s connect, 30s max) to distinguish
|
||||
# transient timeout from host-down
|
||||
# - Non-zero exit naming every failed target; no "OK" summary when any leg failed
|
||||
#
|
||||
# Output shape per leg:
|
||||
# ✅ <name>: alive
|
||||
# 🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)
|
||||
#
|
||||
# Failure kinds: timeout | refused | tls | unexpected:<code> (printed in the failure line)
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
# ── Configuration (documented in infrastructure-monitoring.prose.md) ────────
|
||||
# Change these in ONE place; test_infra_monitoring.sh asserts against these.
|
||||
|
||||
GRAFANA_HOST="192.168.68.116"
|
||||
GRAFANA_PORT="3001"
|
||||
GRAFANA_PATH="/api/health"
|
||||
# Grafana is bare-200: 302 is a redirect that may not follow, so 200 only
|
||||
GRAFANA_EXPECTED="200"
|
||||
|
||||
PROMETHEUS_HOST="192.168.68.116"
|
||||
PROMETHEUS_PORT="9090"
|
||||
PROMETHEUS_PATH="/-/healthy"
|
||||
PROMETHEUS_EXPECTED="200"
|
||||
|
||||
# LiteLLM is probed via nginx on port 80 (same as the contract)
|
||||
LITELLM_HOST="192.168.68.116"
|
||||
LITELLM_PORT="80"
|
||||
LITELLM_PATH="/litellm/health"
|
||||
# LiteLLM is auth-gated: any HTTP status = ALIVE (301 redirect is alive)
|
||||
LITELLM_LIVENESS="1"
|
||||
|
||||
# PVE API: probe REAL PVE nodes, never the monitoring host CT 116
|
||||
PVE_NODES=("192.168.68.9" "192.168.68.12" "192.168.68.6" "192.168.68.15" "192.168.68.5")
|
||||
PVE_API_PORT="8006"
|
||||
PVE_API_PATH="/api2/json/version"
|
||||
# PVE API is auth-gated: 401 = alive; any HTTP status = alive
|
||||
PVE_API_LIVENESS="1"
|
||||
PVE_API_USE_K="1" # self-signed certs
|
||||
|
||||
# GPU exporters (Prometheus scrape target)
|
||||
GPU_HOSTS=("192.168.68.8" "192.168.68.110" "192.168.68.15")
|
||||
GPU_PORT="9400"
|
||||
GPU_PATH="/metrics"
|
||||
GPU_EXPECTED="200"
|
||||
|
||||
# Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116
|
||||
DOCKER_STATS_PORT="9324" # harness-docker-stats (docker_container_* metrics)
|
||||
PVE_EXPORTER_PORT="9221" # harness-pve-exporter (5 pve_* metrics)
|
||||
CT116_SSH_HOST="192.168.68.116"
|
||||
# Both are bare-200: 404 = container not yet started
|
||||
DOCKER_STATS_EXPECTED="200|404"
|
||||
PVE_EXPORTER_EXPECTED="200|404"
|
||||
|
||||
# ── Probe Functions ─────────────────────────────────────────────────────────
|
||||
|
||||
# probe_http <host> <port> <path> <expected_pattern> [use_k] [ssh_host] [scheme] [liveness]
|
||||
# Returns 0 if probe succeeds (matches expected or liveness), 1 if probe-failed.
|
||||
# Prints the result line.
|
||||
#
|
||||
# FIX C1: The kind value is computed and printed in the failure line.
|
||||
# FIX C2: SSH retry logic is in the first attempt branch (not unreachable).
|
||||
|
||||
LAST_KIND=""
|
||||
probe_http() {
|
||||
local host="$1" port="$2" path="$3" expected="$4"
|
||||
local use_k="${5:-}" ssh_host="${6:-}" scheme="${7:-http}" liveness="${8:-0}"
|
||||
local url="${scheme}://${host}:${port}${path}"
|
||||
local code="" kind=""
|
||||
LAST_KIND=""
|
||||
|
||||
# Single invocation that captures both output and status
|
||||
if [ -n "$ssh_host" ]; then
|
||||
out=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||
rc=$?
|
||||
else
|
||||
out=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} "$url" 2>/dev/null)
|
||||
rc=$?
|
||||
fi
|
||||
code=$(printf '%s' "$out" | tr -d '[:space:]')
|
||||
|
||||
# Classify failure kind and retry if needed
|
||||
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||
# Distinguish timeout from TLS error from refused
|
||||
case "$rc" in
|
||||
35|51|58|59|60|77|83) kind="tls" ;;
|
||||
*) kind="timeout" ;;
|
||||
esac
|
||||
# Retry once at longer timeout (25s connect, 30s max)
|
||||
if [ -n "$ssh_host" ]; then
|
||||
code=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||
rc=$?
|
||||
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||
else
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} "$url" 2>/dev/null)
|
||||
rc=$?
|
||||
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||
# On retry, classify: still 000 = keep existing kind (or timeout if empty), unexpected status = refused
|
||||
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||
[ -z "$kind" ] && kind="timeout"
|
||||
elif ! echo "$code" | grep -qE "^(${expected})$"; then
|
||||
kind="refused"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# Check result
|
||||
if [ -n "$code" ] && [ "$code" != "000" ]; then
|
||||
if [ "$liveness" = "1" ]; then
|
||||
# Any HTTP status = ALIVE for auth-gated/redirect endpoints
|
||||
return 0
|
||||
else
|
||||
# Bare-200 or specific expected pattern
|
||||
if echo "$code" | grep -qE "^(${expected})$"; then
|
||||
return 0
|
||||
else
|
||||
kind="unexpected:$code"
|
||||
LAST_KIND="$kind"
|
||||
return 1
|
||||
fi
|
||||
fi
|
||||
else
|
||||
[ -z "$kind" ] && kind="refused"
|
||||
LAST_KIND="$kind"
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
|
||||
# ── Main ────────────────────────────────────────────────────────────────────
|
||||
|
||||
FAILED=()
|
||||
FAILED_KIND=()
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
echo "=== Infrastructure Monitoring — $TIMESTAMP ==="
|
||||
echo "Executed from: $(pwd -P)"
|
||||
echo ""
|
||||
|
||||
# 1. Grafana (CT 116 :3001 /api/health) — bare-200
|
||||
if probe_http "$GRAFANA_HOST" "$GRAFANA_PORT" "$GRAFANA_PATH" "$GRAFANA_EXPECTED"; then
|
||||
echo " ✅ Grafana: alive"
|
||||
else
|
||||
echo " 🔴 Grafana: probe-failed: ${GRAFANA_HOST}:${GRAFANA_PORT} (expected ${GRAFANA_EXPECTED}) (<${LAST_KIND}>)"
|
||||
FAILED+=("grafana")
|
||||
fi
|
||||
|
||||
# 2. Prometheus (CT 116 :9090 /-/healthy) — bare-200
|
||||
if probe_http "$PROMETHEUS_HOST" "$PROMETHEUS_PORT" "$PROMETHEUS_PATH" "$PROMETHEUS_EXPECTED"; then
|
||||
echo " ✅ Prometheus: alive"
|
||||
else
|
||||
echo " 🔴 Prometheus: probe-failed: ${PROMETHEUS_HOST}:${PROMETHEUS_PORT} (expected ${PROMETHEUS_EXPECTED}) (<${LAST_KIND}>)"
|
||||
FAILED+=("prometheus")
|
||||
fi
|
||||
|
||||
# 3. LiteLLM (CT 116 :80/litellm/health via nginx) — liveness (any HTTP = alive)
|
||||
if probe_http "$LITELLM_HOST" "$LITELLM_PORT" "$LITELLM_PATH" "" "" "" "http" "$LITELLM_LIVENESS"; then
|
||||
echo " ✅ LiteLLM: alive"
|
||||
else
|
||||
echo " 🔴 LiteLLM: probe-failed: ${LITELLM_HOST}:${LITELLM_PORT}${LITELLM_PATH} (any-HTTP liveness) (<${LAST_KIND}>)"
|
||||
FAILED+=("litellm")
|
||||
fi
|
||||
|
||||
# 4. PVE API (5 real nodes :8006 /api2/json/version, -k, liveness)
|
||||
PVE_FAILED=()
|
||||
for node in "${PVE_NODES[@]}"; do
|
||||
if probe_http "$node" "$PVE_API_PORT" "$PVE_API_PATH" "" "$PVE_API_USE_K" "" "https" "$PVE_API_LIVENESS"; then
|
||||
echo " ✅ PVE API ${node}: alive"
|
||||
else
|
||||
echo " 🔴 PVE API ${node}: probe-failed: ${node}:${PVE_API_PORT} (any-HTTP liveness, -k for self-signed) (<${LAST_KIND}>)"
|
||||
PVE_FAILED+=("$node")
|
||||
fi
|
||||
done
|
||||
if [ ${#PVE_FAILED[@]} -gt 0 ]; then
|
||||
FAILED+=("pve-api: ${PVE_FAILED[*]}")
|
||||
fi
|
||||
|
||||
# 5. GPU exporters (:9400/metrics) — bare-200
|
||||
GPU_FAILED=()
|
||||
for host in "${GPU_HOSTS[@]}"; do
|
||||
if probe_http "$host" "$GPU_PORT" "$GPU_PATH" "$GPU_EXPECTED"; then
|
||||
echo " ✅ GPU exporter ${host}: alive"
|
||||
else
|
||||
echo " 🔴 GPU exporter ${host}: probe-failed: ${host}:${GPU_PORT} (expected 200) (<${LAST_KIND}>)"
|
||||
GPU_FAILED+=("$host")
|
||||
fi
|
||||
done
|
||||
if [ ${#GPU_FAILED[@]} -gt 0 ]; then
|
||||
FAILED+=("gpu-exporters: ${GPU_FAILED[*]}")
|
||||
fi
|
||||
|
||||
# 6. Docker Stats (CT 116 :9324, 127.0.0.1 via SSH) — 200|404
|
||||
if probe_http "127.0.0.1" "$DOCKER_STATS_PORT" "/" "$DOCKER_STATS_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||
echo " ✅ Docker Stats: alive"
|
||||
else
|
||||
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:${DOCKER_STATS_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||
FAILED+=("docker-stats")
|
||||
fi
|
||||
|
||||
# 7. PVE Exporter (CT 116 :9221, 127.0.0.1 via SSH) — 200|404
|
||||
if probe_http "127.0.0.1" "$PVE_EXPORTER_PORT" "/" "$PVE_EXPORTER_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||
echo " ✅ PVE Exporter: alive"
|
||||
else
|
||||
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:${PVE_EXPORTER_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||
FAILED+=("pve-exporter")
|
||||
fi
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
|
||||
echo ""
|
||||
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||
echo " ✅ All legs OK"
|
||||
exit 0
|
||||
else
|
||||
for f in "${FAILED[@]}"; do
|
||||
echo " 🔴 FAILED: $f"
|
||||
done
|
||||
exit 1
|
||||
fi
|
||||
Executable
+358
@@ -0,0 +1,358 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
LiteLLM Health Check - Contract executor
|
||||
Runs all checks defined in litellm-health.prose.md and reports results.
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import sys
|
||||
import json
|
||||
import time
|
||||
import random
|
||||
|
||||
# Configuration
|
||||
BACKEND_HOST = "192.168.68.116"
|
||||
GPU_HOSTS = {
|
||||
"gpu-dense": "192.168.68.8",
|
||||
"gpu-vision": "192.168.68.110",
|
||||
"strix-moe": "192.168.68.15"
|
||||
}
|
||||
|
||||
def run_command(cmd, timeout=15):
|
||||
"""Run a command and return (exit_code, stdout, stderr)"""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
cmd,
|
||||
shell=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=timeout
|
||||
)
|
||||
return result.returncode, result.stdout.strip(), result.stderr.strip()
|
||||
except subprocess.TimeoutExpired:
|
||||
return 1, "", "TIMEOUT"
|
||||
except Exception as e:
|
||||
return 1, "", str(e)
|
||||
|
||||
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
|
||||
"""Probe HTTP endpoint and return (status_code, failure_kind)
|
||||
|
||||
Returns:
|
||||
(code, None) if successful or HTTP response received
|
||||
(000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc.
|
||||
"""
|
||||
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
|
||||
if method == "POST":
|
||||
cmd += " -X POST"
|
||||
if bearer_token:
|
||||
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
|
||||
if data:
|
||||
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
|
||||
if follow_redirects:
|
||||
cmd += " -L"
|
||||
cmd += " '" + url + "'"
|
||||
|
||||
try:
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
if rc != 0:
|
||||
# Check if this is a timeout from run_command (rc=1, stderr="TIMEOUT")
|
||||
if rc == 1 and stderr == "TIMEOUT":
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
# Otherwise, determine failure kind from curl exit code
|
||||
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
||||
elif rc == 28:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
elif rc == 7:
|
||||
return (000, "connection refused")
|
||||
elif rc == 6:
|
||||
return (000, "dns failure")
|
||||
elif rc == 35:
|
||||
return (000, "ssl error")
|
||||
elif rc == 52:
|
||||
return (000, "empty response")
|
||||
else:
|
||||
return (000, "curl exit " + str(rc))
|
||||
return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response")
|
||||
except subprocess.TimeoutExpired:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
|
||||
def check_host_health(host_ip):
|
||||
"""Check if the GPU host's llama-chat-api health endpoint is reachable
|
||||
|
||||
Returns: (healthy: bool, detail: str)
|
||||
"""
|
||||
code, _ = probe_http("http://" + host_ip + ":8080/health", timeout=10)
|
||||
if code == 200:
|
||||
return True, "host healthy (200)"
|
||||
elif code == 000:
|
||||
return False, "host unreachable (timeout or refused)"
|
||||
else:
|
||||
return False, "host unhealthy (HTTP " + str(code) + ")"
|
||||
|
||||
|
||||
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
||||
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
||||
cmd = "curl -s -m " + str(timeout)
|
||||
if method == "POST":
|
||||
cmd += " -X POST"
|
||||
if bearer_token:
|
||||
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
|
||||
if data:
|
||||
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
|
||||
cmd += " '" + url + "'"
|
||||
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
# Return first 200 chars, single line
|
||||
body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else ""
|
||||
return body
|
||||
|
||||
|
||||
def check_liveliness():
|
||||
"""Step 1: Liveliness probe"""
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
|
||||
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
|
||||
|
||||
def check_containers():
|
||||
"""Step 2: Container health via SSH"""
|
||||
cmd = "ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@192.168.68.116 'docker ps --format \"{{.Names}} {{.Status}}\"'"
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Containers", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
lines = stdout.split('\n') if stdout else []
|
||||
container_count = len([l for l in lines if l.strip()])
|
||||
healthy = container_count >= 8
|
||||
return "Containers", healthy, str(container_count) + " containers"
|
||||
|
||||
def check_model_probes():
|
||||
"""Step 6: Model probe - all 4 aliases"""
|
||||
# Get monitor key
|
||||
monitor_key = run_command("ssh -o BatchMode=yes root@192.168.68.116 \"grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2\"")[1]
|
||||
|
||||
results = []
|
||||
|
||||
# Host health mapping: model -> host IP
|
||||
model_hosts = {
|
||||
"gpu-dense": "192.168.68.8", # RTX 3090
|
||||
"gpu-vision": "192.168.68.110", # RTX 5070
|
||||
"strix-moe": "192.168.68.15" # Strix Halo
|
||||
}
|
||||
|
||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||
host_ip = model_hosts[model]
|
||||
# Single-host aliases: 30s initial timeout, retry once at 90s on failure
|
||||
# Worst-case prefill ~76s, so 90s retry ensures we cover it
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
|
||||
first_kind = None
|
||||
if code == 000 and failure_kind:
|
||||
first_kind = failure_kind
|
||||
time.sleep(1)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=90)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Both attempts failed - check host health to distinguish busy from down
|
||||
host_healthy, host_detail = check_host_health(host_ip)
|
||||
if host_healthy:
|
||||
results.append((model, False, "busy (completion timed out after retry; " + host_detail + ")"))
|
||||
else:
|
||||
# Host unreachable - report both kinds
|
||||
if first_kind:
|
||||
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
|
||||
else:
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
|
||||
elif code == 200:
|
||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
elif code in (401, 403):
|
||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||
timeout=10)
|
||||
alias = "monitor-20260813"
|
||||
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||
else:
|
||||
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
|
||||
# Pool alias (syslog-auto): 60s timeout, retry once on 000
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Retry once with same timeout
|
||||
time.sleep(1)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
if code == 000 and failure_kind:
|
||||
results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)"))
|
||||
elif code in (401, 403):
|
||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||
timeout=10)
|
||||
alias = "monitor-20260813"
|
||||
results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||
else:
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
else:
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
|
||||
return results
|
||||
|
||||
def check_admin_key_list():
|
||||
"""Step 8: Admin API key list - use two-step approach"""
|
||||
# Step 1: Get master key
|
||||
mk_cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"docker exec harness-litellm printenv LITELLM_MASTER_KEY\""
|
||||
mk_rc, mk_stdout, mk_stderr = run_command(mk_cmd)
|
||||
|
||||
if mk_rc != 0:
|
||||
return "Admin Key List", False, "credential-missing (ssh failed: " + mk_stderr + ")"
|
||||
|
||||
mk = mk_stdout
|
||||
if not mk or "NO-CURL" in mk:
|
||||
return "Admin Key List", False, "credential-missing (empty or NO-CURL)"
|
||||
|
||||
# Step 2: Call using the key - use double quotes inside SSH command
|
||||
cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"curl -s -H \\\"Authorization: Bearer " + mk + "\\\" http://127.0.0.1:4000/key/list\""
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Admin Key List", False, "admin-call-failed (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
# Try to parse the response
|
||||
try:
|
||||
data = json.loads(stdout)
|
||||
# Response is a dict with "keys" (paginated list) and "total_count" fields
|
||||
if isinstance(data, dict) and "keys" in data:
|
||||
key_count = data.get("total_count", len(data["keys"]))
|
||||
elif isinstance(data, list):
|
||||
key_count = len(data)
|
||||
else:
|
||||
key_count = 0
|
||||
if key_count == 0:
|
||||
return "Admin Key List", False, "admin-call-failed (empty response)"
|
||||
return "Admin Key List", True, str(key_count) + " total (" + str(len(data.get("keys", []) if isinstance(data, dict) else data)) + " on page 1)" if isinstance(data, dict) else str(key_count) + " total"
|
||||
except Exception as e:
|
||||
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
|
||||
|
||||
def check_github_status():
|
||||
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
|
||||
code, _ = probe_http("https://status.github.com/api/status.json", timeout=15)
|
||||
# GitHub status API returns 301 redirect, which is expected behavior
|
||||
return "GitHub Status", code == 301, str(code)
|
||||
|
||||
def check_prometheus():
|
||||
"""Step 4: Prometheus health"""
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
|
||||
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
|
||||
|
||||
def check_grafana():
|
||||
"""Step 9: Grafana health"""
|
||||
code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
|
||||
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
|
||||
|
||||
def check_docker_stats():
|
||||
"""Step 10: Docker Stats health - fetch from CT 116 host"""
|
||||
# Docker stats is on localhost from CT 116
|
||||
cmd = "ssh -o BatchMode=yes root@192.168.68.116 'curl -s http://127.0.0.1:9324/metrics | head -20'"
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Docker Stats", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
# Check response is non-empty
|
||||
if not stdout or len(stdout) < 100:
|
||||
return "Docker Stats", False, "empty response"
|
||||
|
||||
return "Docker Stats", True, "200 (target: 127.0.0.1:9324/metrics from CT 116)"
|
||||
|
||||
def main():
|
||||
print("🏥 LiteLLM Health Check v1.0.0")
|
||||
print("📍 Backend edge: http://" + BACKEND_HOST)
|
||||
print("")
|
||||
|
||||
all_pass = True
|
||||
degraded = [] # Track degraded (busy) checks
|
||||
|
||||
# Run all checks
|
||||
checks = [
|
||||
check_liveliness(),
|
||||
check_containers(),
|
||||
check_prometheus(),
|
||||
check_grafana(),
|
||||
]
|
||||
|
||||
for result in checks:
|
||||
name, passed, detail = result
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
if not passed:
|
||||
all_pass = False
|
||||
|
||||
# Model probes
|
||||
model_results = check_model_probes()
|
||||
for name, passed, detail in model_results:
|
||||
# Check if this is a busy (degraded) verdict
|
||||
if not passed and detail.startswith("busy "):
|
||||
status = "⚠️"
|
||||
degraded.append(name)
|
||||
else:
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
# Only set all_pass=False for real failures (not busy)
|
||||
if not passed and not detail.startswith("busy "):
|
||||
all_pass = False
|
||||
|
||||
# Admin key list
|
||||
admin_result = check_admin_key_list()
|
||||
status = "✅" if admin_result[1] else "❌"
|
||||
print(" " + status + " Admin Key List: " + admin_result[2])
|
||||
if not admin_result[1]:
|
||||
all_pass = False
|
||||
|
||||
# GitHub status
|
||||
github_result = check_github_status()
|
||||
status = "✅" if github_result[1] else "❌"
|
||||
print(" " + status + " " + github_result[0] + ": " + github_result[2])
|
||||
if not github_result[1]:
|
||||
all_pass = False
|
||||
|
||||
# Docker stats
|
||||
docker_stats_result = check_docker_stats()
|
||||
status = "✅" if docker_stats_result[1] else "❌"
|
||||
print(" " + status + " " + docker_stats_result[0] + ": " + docker_stats_result[2])
|
||||
if not docker_stats_result[1]:
|
||||
all_pass = False
|
||||
|
||||
print("")
|
||||
if all_pass:
|
||||
if degraded:
|
||||
print("✅ All checks passed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||
else:
|
||||
print("✅ All checks passed")
|
||||
return 0
|
||||
else:
|
||||
if degraded:
|
||||
print("❌ Some checks failed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
return 1
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
+8
-10
@@ -11,11 +11,13 @@ set -euo pipefail
|
||||
# ── CT ID → PVE Node mapping (maintained HERE, not in prose contracts) ──
|
||||
declare -A CT_NODES=(
|
||||
# amdpve (192.168.68.15)
|
||||
[111]=amdpve # tdunna
|
||||
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
|
||||
[112]=amdpve # tanko
|
||||
[113]=amdpve # baggy
|
||||
[115]=amdpve # scottdenya
|
||||
[120]=amdpve # adguard2 (added 2026-09-12)
|
||||
# minipve (192.168.68.12)
|
||||
[100]=minipve # abiba (was hwepve)
|
||||
[102]=minipve # adguard (was acerpve)
|
||||
[104]=minipve # authentik
|
||||
[110]=minipve # gitea
|
||||
@@ -25,19 +27,16 @@ declare -A CT_NODES=(
|
||||
[106]=storepve # ra-h-os
|
||||
[107]=storepve # proxmox-backup
|
||||
[108]=storepve # media
|
||||
[111]=storepve # tdunna (was amdpve — corrected 2026-09-12; live per pvesh)
|
||||
[117]=storepve # zulip
|
||||
[118]=storepve # jdownloader
|
||||
# acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
|
||||
# hwepve (192.168.68.4)
|
||||
[100]=hwepve # abiba (was amdpve)
|
||||
[105]=hwepve # kagentz (was amdpve)
|
||||
[114]=hwepve # mumuni (was minipve)
|
||||
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
|
||||
#
|
||||
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
|
||||
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090)
|
||||
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070)
|
||||
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH)
|
||||
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090) [QEMU VM on acerpve]
|
||||
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070) [QEMU VM on ocupve]
|
||||
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH) [QEMU VM on storepve]
|
||||
)
|
||||
|
||||
# Each node must be root-accessible via SSH hostname
|
||||
@@ -50,7 +49,6 @@ declare -A NODE_IPS=(
|
||||
[storepve]=192.168.68.6
|
||||
[acerpve]=192.168.68.9
|
||||
[ocupve]=192.168.68.5
|
||||
[hwepve]=192.168.68.4
|
||||
)
|
||||
|
||||
resolve_node() {
|
||||
@@ -77,7 +75,7 @@ main() {
|
||||
if [[ $# -lt 1 ]]; then
|
||||
echo "Usage: pct-run <CT_ID> [command...]" >&2
|
||||
echo " pct-run 112 cat /etc/hostname" >&2
|
||||
echo " pct-run 114 systemctl status hermes-gateway" >&2
|
||||
echo " pct-run 100 systemctl status hermes-gateway" >&2
|
||||
echo ""
|
||||
echo "Known CTs:" >&2
|
||||
for ct in $(echo "${!CT_NODES[@]}" | tr ' ' '\n' | sort -n); do
|
||||
|
||||
@@ -1,10 +1,11 @@
|
||||
#!/bin/bash
|
||||
# pm2-self-heal — hourly PM2 process check
|
||||
# Part of the pm2-self-heal prose contract
|
||||
# Alerts via Telegram (abiba-zulip decommissioned 2026-07-04)
|
||||
# Alerts via Telegram (primary) and Zulip DM (secondary, abiba-zulip restored 2026-08-03)
|
||||
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||
|
||||
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
||||
LOG="/root/pm2-self-heal.log"
|
||||
TELEGRAM_CHAT_ID="5822977936"
|
||||
|
||||
notify_tg() {
|
||||
@@ -16,8 +17,6 @@ notify_tg() {
|
||||
-d "text=${msg}" \
|
||||
-d "parse_mode=HTML" > /dev/null 2>&1 || true
|
||||
}
|
||||
ALERTS="${ALERTS}$msg"
|
||||
}
|
||||
|
||||
# Log-only mode: replaced by prose contract pm2-self-heal.prose.md
|
||||
# Only alerts Telegram on actual failure (status != online)
|
||||
@@ -33,7 +32,7 @@ TEL_LINE=$(echo "$STATUS" | grep "abiba-telegram")
|
||||
TEL_STATUS=$(echo "$TEL_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
TEL_RESTARTS=$(echo "$TEL_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$TEL_STATUS" != "online" ]; then
|
||||
if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
|
||||
pm2 restart abiba-telegram > /dev/null 2>&1
|
||||
sleep 3
|
||||
TEL_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-telegram")
|
||||
@@ -47,9 +46,75 @@ if [ "$TEL_STATUS" != "online" ]; then
|
||||
fi
|
||||
fi
|
||||
|
||||
# Check abiba-zulip (live Zulip bridge, heartbeating)
|
||||
ZULIP_LINE=$(echo "$STATUS" | grep "abiba-zulip")
|
||||
ZULIP_STATUS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
ZULIP_RESTARTS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$ZULIP_STATUS" != "online" ]; then
|
||||
pm2 restart abiba-zulip > /dev/null 2>&1
|
||||
sleep 3
|
||||
ZULIP_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-zulip")
|
||||
ZULIP_STATUS2=$(echo "$ZULIP_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$ZULIP_STATUS2" = "online" ]; then
|
||||
msg="⚠️ abiba-zulip was **$ZULIP_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 abiba-zulip **failed restart** (was $ZULIP_STATUS, still $ZULIP_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$ZULIP_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ abiba-zulip has **$ZULIP_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Check gitea-runner
|
||||
GITEA_LINE=$(echo "$STATUS" | grep "gitea-runner")
|
||||
GITEA_STATUS=$(echo "$GITEA_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
GITEA_RESTARTS=$(echo "$GITEA_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$GITEA_STATUS" != "online" ]; then
|
||||
pm2 restart gitea-runner > /dev/null 2>&1
|
||||
sleep 3
|
||||
GITEA_LINE2=$(pm2 status --no-color 2>/dev/null | grep "gitea-runner")
|
||||
GITEA_STATUS2=$(echo "$GITEA_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$GITEA_STATUS2" = "online" ]; then
|
||||
msg="⚠️ gitea-runner was **$GITEA_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 gitea-runner **failed restart** (was $GITEA_STATUS, still $GITEA_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$GITEA_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ gitea-runner has **$GITEA_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Check zulip-watchdog
|
||||
WATCHDOG_LINE=$(echo "$STATUS" | grep "zulip-watchdog")
|
||||
WATCHDOG_STATUS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
WATCHDOG_RESTARTS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$WATCHDOG_STATUS" != "online" ]; then
|
||||
pm2 restart zulip-watchdog > /dev/null 2>&1
|
||||
sleep 3
|
||||
WATCHDOG_LINE2=$(pm2 status --no-color 2>/dev/null | grep "zulip-watchdog")
|
||||
WATCHDOG_STATUS2=$(echo "$WATCHDOG_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$WATCHDOG_STATUS2" = "online" ]; then
|
||||
msg="⚠️ zulip-watchdog was **$WATCHDOG_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 zulip-watchdog **failed restart** (was $WATCHDOG_STATUS, still $WATCHDOG_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$WATCHDOG_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ zulip-watchdog has **$WATCHDOG_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Log check
|
||||
{
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS alerts=${ALERTS:+yes}"
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS zulip=$ZULIP_STATUS gitea=$GITEA_STATUS watchdog=$WATCHDOG_STATUS alerts=${ALERTS:+yes}"
|
||||
[ -n "$ALERTS" ] && echo "$ALERTS"
|
||||
} >> "$LOG"
|
||||
|
||||
|
||||
+16
-13
@@ -46,32 +46,35 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
|
||||
|
||||
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
|
||||
|
||||
**Proxmox Cluster "Tabiri" (6 nodes):**
|
||||
- amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
|
||||
- minipve (192.168.68.12): adguard, authentik, gitea, syslog-api, infisical-vault
|
||||
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
|
||||
**Proxmox Cluster "Tabiri" (5 nodes):**
|
||||
- amdpve (192.168.68.15): kagentz, tanko, baggy, scottdenya, adguard2
|
||||
- minipve (192.168.68.12): abiba, adguard, authentik, gitea, syslog-api, infisical-vault
|
||||
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
|
||||
- acerpve (192.168.68.9): llm-gpu
|
||||
- ocupve (192.168.68.5): ocu-llm
|
||||
- hwepve (192.168.68.4): abiba, kagentz, mumuni
|
||||
|
||||
**CT IDs (verified 2026-07-24 against PVE API):**
|
||||
**CT IDs (verified 2026-09-12 against PVE API):**
|
||||
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
|
||||
107:pbs 108:media 110:gitea 111:tdunna 112:tanko
|
||||
113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip
|
||||
118:jdownloader 119:infisical-vault
|
||||
113:baggy 115:scottdenya 116:syslog-api 117:zulip
|
||||
118:jdownloader 119:infisical-vault 120:adguard2
|
||||
|
||||
**CT 111 (tdunna, 192.168.68.129) is REPORT-ONLY — Theo's box; alert only, never garbage-collect.**
|
||||
|
||||
**NO CT 122, CT 123, or .19 exist in the cluster.**
|
||||
|
||||
**CRITICAL RULES (never regress):**
|
||||
1. NO /grafana/ nginx route — it was tried and reverted on 2026-07-02. Grafana is direct LAN at :3001.
|
||||
2. NO .19 IP — Zulip is CT 117 on storepve.
|
||||
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 114.
|
||||
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
|
||||
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
|
||||
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only.
|
||||
5. abiba-zulip PM2 process is ONLINE (verified 2026-09-11) — this rule was stale.
|
||||
|
||||
**Docker on CT 116 (8 containers):**
|
||||
harness-litellm, harness-router, harness-nginx, harness-postgres,
|
||||
harness-redis, harness-dashboard, harness-grafana, harness-prometheus
|
||||
**Docker on CT 116 (11 containers, verified 2026-09-11):**
|
||||
harness-litellm, harness-nginx, harness-postgres, harness-redis,
|
||||
harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager,
|
||||
harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
|
||||
(harness-router was decommissioned 2026-09-11)
|
||||
|
||||
## DIFF TO REVIEW
|
||||
|
||||
|
||||
@@ -12,10 +12,10 @@ echo ""
|
||||
# Authorized agents for restricted contracts
|
||||
# Format: contract_pattern|authorized_agents (comma-separated)
|
||||
declare -A RESTRICTED
|
||||
RESTRICTED["infrastructure-control.prose.md"]="abiba"
|
||||
RESTRICTED["proxmox-monitor.prose.md"]="abiba"
|
||||
RESTRICTED["hermes-config-template.prose.md"]="abiba,mumuni,tanko"
|
||||
RESTRICTED["zulip-health.prose.md"]="abiba,mumuni"
|
||||
RESTRICTED["infrastructure-control.prose.md"]="abiba,abiba-bot,tanko,tanko-bot,mumuni,mumuni-bot"
|
||||
RESTRICTED["proxmox-monitor.prose.md"]="abiba,abiba-bot"
|
||||
RESTRICTED["hermes-config-template.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
|
||||
RESTRICTED["zulip-health.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
|
||||
RESTRICTED["scripts/pm2-self-heal.sh"]="abiba"
|
||||
RESTRICTED["scripts/prose-lint.sh"]="abiba"
|
||||
RESTRICTED["scripts/prose-ai-review.sh"]="abiba"
|
||||
|
||||
+47
-3
@@ -57,7 +57,7 @@ echo "── 2. Regression detection ──"
|
||||
|
||||
# Grafana /grafana/ as nginx route or URL path (reverted 2026-07-02)
|
||||
# EXCLUDE: filesystem paths (/opt/monitoring/grafana/...), directory creation, revert docs
|
||||
GRAFANA_HITS=$(grep -rn '/grafana/' *.prose.md 2>/dev/null \
|
||||
GRAFANA_HITS=$(grep -rn '/grafana/' ./*.prose.md 2>/dev/null \
|
||||
| grep -v '/opt/monitoring/grafana/' \
|
||||
| grep -v 'was tried and reverted\|was reverted\|do not re-add\|NOT recommended' \
|
||||
| grep -v 'mkdir.*grafana\|Create.*grafana' \
|
||||
@@ -71,7 +71,7 @@ else
|
||||
fi
|
||||
|
||||
# Stale CT IDs (CT 122, CT 123 as CT IDs — not IPs .122, .123)
|
||||
CT_STALE=$(grep -rn '\bCT 122\b' *.prose.md 2>/dev/null || true)
|
||||
CT_STALE=$(grep -rn '\bCT 122\b' ./*.prose.md 2>/dev/null || true)
|
||||
if [ -n "$CT_STALE" ]; then
|
||||
echo " ❌ REGRESSION: CT 122 used as CT ID — Tanko is CT 112"
|
||||
echo "$CT_STALE"
|
||||
@@ -84,6 +84,35 @@ fi
|
||||
# .122/.123 are correct — verified reachable bridge IPs for Tanko/Mumuni
|
||||
echo " ✅ IP consistency verified (.19=.122=.123 all reachable)"
|
||||
|
||||
# Report provenance — every contract report must state the absolute path it
|
||||
# executed from, so a stale-consumer report is distinguishable from a real fault
|
||||
# at read time (2026-09-09 probe-drift incident: three false DEGRADED rounds).
|
||||
# Enforced only inside the **Report format** paragraph, and a check-health
|
||||
# contract with no Report format paragraph FAILs rather than being skipped.
|
||||
PROV_FILES=$(grep -rlE '^### check-health|\*\*Report format\*\*' ./*.prose.md 2>/dev/null || true)
|
||||
if [ -z "$PROV_FILES" ]; then
|
||||
echo " ❌ No check-health/report-format contracts found — provenance not enforced"
|
||||
FAILED=1
|
||||
else
|
||||
PROV_BAD=0
|
||||
while IFS= read -r f; do
|
||||
[ -n "$f" ] || continue
|
||||
REPORT_PARA=$(awk '/\*\*Report format\*\*/{found=1} found{print} found && /^[[:space:]]*$/{exit}' "$f")
|
||||
if [ -z "$REPORT_PARA" ]; then
|
||||
echo " ❌ $f: check-health contract has no **Report format** paragraph"
|
||||
PROV_BAD=1
|
||||
elif ! printf '%s\n' "$REPORT_PARA" | grep -qE 'absolute path|pwd -P|executed from'; then
|
||||
echo " ❌ $f: **Report format** lacks execution provenance (absolute path / pwd -P)"
|
||||
PROV_BAD=1
|
||||
fi
|
||||
done <<< "$PROV_FILES"
|
||||
if [ "$PROV_BAD" -eq 1 ]; then
|
||||
FAILED=1
|
||||
else
|
||||
echo " ✅ Report provenance present in all report-format contracts"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── 3. Cross-contract consistency ──
|
||||
echo ""
|
||||
echo "── 3. Cross-contract consistency ──"
|
||||
@@ -106,7 +135,22 @@ fi
|
||||
|
||||
echo " Cross-contract: $WARNINGS total warnings across all checks"
|
||||
|
||||
# ── 4. Summary ──
|
||||
# ── 4. Committed-credential scan ──
|
||||
# The 2026-09-17 purge removed six live credentials that had sat in .md prose
|
||||
# and scripts for weeks. This step makes that class of commit FAIL the gate
|
||||
# instead of printing a warning. Patterns: scripts/secret-patterns.tsv.
|
||||
# Only deliberate synthetic examples may be listed in scripts/secret-allowlist.tsv,
|
||||
# each with a reason. Run `bash scripts/secret-scan.sh --staged` before committing.
|
||||
echo ""
|
||||
echo "── 4. Secret scan (committed credentials) ──"
|
||||
if bash scripts/secret-scan.sh; then
|
||||
echo " ✅ No committed credentials"
|
||||
else
|
||||
echo " ❌ COMMITTED CREDENTIAL DETECTED"
|
||||
FAILED=1
|
||||
fi
|
||||
|
||||
# ── 5. Summary ──
|
||||
echo ""
|
||||
echo "═══════════════════════════════════"
|
||||
if [ $FAILED -eq 1 ]; then
|
||||
|
||||
Executable
+138
@@ -0,0 +1,138 @@
|
||||
#!/bin/bash
|
||||
# proxmox-monitor.sh — Proxmox Cluster + Docker Monitoring Health Check
|
||||
# Implements proxmox-monitor.prose.md (check-health section)
|
||||
#
|
||||
# Legs: Prometheus, Grafana, Docker Stats Exporter, PVE Exporter, PBS GC
|
||||
# All legs must return 200 for healthy status.
|
||||
#
|
||||
# Run: bash scripts/proxmox-monitor.sh
|
||||
# Exits 0 if all probes pass, 1 if any fails.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
CT116_HOST="192.168.68.116"
|
||||
FAILED=()
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
|
||||
echo "=== Proxmox Monitor — $TIMESTAMP ==="
|
||||
echo "Executed from: $(pwd -P)"
|
||||
echo ""
|
||||
|
||||
# 1. Prometheus health (bound to 0.0.0.0:9090 on .116)
|
||||
PROM_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:9090/-/healthy 2>/dev/null)
|
||||
PROM_CODE=$(printf '%s' "$PROM_CODE" | tr -d '[:space:]')
|
||||
[ -n "$PROM_CODE" ] || PROM_CODE="000"
|
||||
|
||||
if [ "$PROM_CODE" = "200" ]; then
|
||||
echo " ✅ Prometheus: alive"
|
||||
else
|
||||
echo " 🔴 Prometheus: probe-failed: ${CT116_HOST}:9090 (expected 200, got ${PROM_CODE})"
|
||||
FAILED+=("prometheus")
|
||||
fi
|
||||
|
||||
# 2. Grafana health (bound to 0.0.0.0:3001 on .116)
|
||||
GRAF_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:3001/api/health 2>/dev/null)
|
||||
GRAF_CODE=$(printf '%s' "$GRAF_CODE" | tr -d '[:space:]')
|
||||
[ -n "$GRAF_CODE" ] || GRAF_CODE="000"
|
||||
|
||||
if [ "$GRAF_CODE" = "200" ]; then
|
||||
echo " ✅ Grafana: alive"
|
||||
else
|
||||
echo " 🔴 Grafana: probe-failed: ${CT116_HOST}:3001 (expected 200, got ${GRAF_CODE})"
|
||||
FAILED+=("grafana")
|
||||
fi
|
||||
|
||||
# 3. Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — probe from .116 localhost)
|
||||
DOCKER_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics" 2>/dev/null)
|
||||
DOCKER_CODE=$(printf '%s' "$DOCKER_CODE" | tr -d '[:space:]')
|
||||
[ -n "$DOCKER_CODE" ] || DOCKER_CODE="000"
|
||||
|
||||
if [ "$DOCKER_CODE" = "200" ]; then
|
||||
echo " ✅ Docker Stats: alive"
|
||||
else
|
||||
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:9324 (expected 200, got ${DOCKER_CODE})"
|
||||
FAILED+=("docker-stats")
|
||||
fi
|
||||
|
||||
# 4. PVE Exporter (bound to 127.0.0.1:9221 on .116 — probe from .116 localhost)
|
||||
PVE_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics" 2>/dev/null)
|
||||
PVE_CODE=$(printf '%s' "$PVE_CODE" | tr -d '[:space:]')
|
||||
[ -n "$PVE_CODE" ] || PVE_CODE="000"
|
||||
|
||||
if [ "$PVE_CODE" = "200" ]; then
|
||||
echo " ✅ PVE Exporter: alive"
|
||||
else
|
||||
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:9221 (expected 200, got ${PVE_CODE})"
|
||||
FAILED+=("pve-exporter")
|
||||
fi
|
||||
|
||||
# 5. PBS GC liveness (storepve-datastore GC must have run within 48h)
|
||||
PBS_GC_OUTPUT=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192.168.68.6 \
|
||||
"pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null)
|
||||
PBS_GC_OUTPUT=$(printf '%s' "$PBS_GC_OUTPUT" | tr -d '[:space:]')
|
||||
[ -n "$PBS_GC_OUTPUT" ] || PBS_GC_OUTPUT="000"
|
||||
|
||||
if [ "$PBS_GC_OUTPUT" = "000" ]; then
|
||||
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
# Parse the JSON to get storepve-datastore's last-run-endtime and pending-bytes
|
||||
PBS_GC_RESULT=$(echo "$PBS_GC_OUTPUT" | python3 -c "
|
||||
import sys, json
|
||||
try:
|
||||
data = json.load(sys.stdin)
|
||||
for store in data:
|
||||
if store['store'] == 'storepve-datastore':
|
||||
endtime = store.get('last-run-endtime')
|
||||
pending = store.get('pending-bytes', 0)
|
||||
if endtime is None or endtime == 0:
|
||||
print('never-run')
|
||||
else:
|
||||
print(f'{endtime}|{pending}')
|
||||
break
|
||||
else:
|
||||
print('absent')
|
||||
except json.JSONDecodeError:
|
||||
print('unparseable')
|
||||
" 2>/dev/null)
|
||||
|
||||
if [ -z "$PBS_GC_RESULT" ] || [ "$PBS_GC_RESULT" = "unparseable" ]; then
|
||||
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)"
|
||||
FAILED+=("pbs-gc")
|
||||
elif [ "$PBS_GC_RESULT" = "absent" ]; then
|
||||
echo " 🔴 PBS GC: never-run (storepve-datastore not found in GC list)"
|
||||
FAILED+=("pbs-gc")
|
||||
elif [ "$PBS_GC_RESULT" = "never-run" ]; then
|
||||
echo " 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
# Parse the endtime|pending format
|
||||
LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1)
|
||||
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||
|
||||
# Convert epoch to age in hours
|
||||
NOW_EPOCH=$(date -u +%s)
|
||||
AGE_HOURS=$(( (NOW_EPOCH - LAST_RUN_ENDTIME) / 3600 ))
|
||||
|
||||
if [ $AGE_HOURS -gt 48 ]; then
|
||||
echo " 🔴 PBS GC: stale (last run ${AGE_HOURS}h ago, pending-bytes: ${PENDING_BYTES} B)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
echo " ✅ PBS GC: healthy (last run ${AGE_HOURS}h ago, pending-bytes: ${PENDING_BYTES} B)"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||
echo " ✅ All legs OK"
|
||||
exit 0
|
||||
else
|
||||
for f in "${FAILED[@]}"; do
|
||||
echo " 🔴 FAILED: $f"
|
||||
done
|
||||
exit 1
|
||||
fi
|
||||
@@ -0,0 +1,55 @@
|
||||
# secret-allowlist.tsv — exceptions for scripts/secret-scan.sh, every entry with a reason.
|
||||
#
|
||||
# Format: <rule-id|*><TAB><path-glob><TAB><literal-substring><TAB><reason>
|
||||
# Blank lines and lines whose first field starts with '#' are ignored.
|
||||
# A finding is suppressed only when ALL THREE of rule, path and literal match:
|
||||
# * the rule id equals the finding's rule id, or is '*'
|
||||
# * the finding's repo-relative path matches <path-glob> (bash glob)
|
||||
# * the finding's line contains <literal-substring> verbatim
|
||||
# An entry whose reason is empty is a hard error (exit 2) — no silent exceptions.
|
||||
#
|
||||
# RULE: never allowlist a live credential, and never broaden an entry (rule '*',
|
||||
# a wide path glob, or a short generic literal) just to silence a finding.
|
||||
# If the finding is real, remove the credential from the file.
|
||||
#
|
||||
# Entries are one per deliberate synthetic example, so the file reads as an
|
||||
# audit trail of reviewed exceptions rather than a list of things to ignore.
|
||||
# Rule '*' is used only where the same literal is matched by more than one rule.
|
||||
#
|
||||
# ── The 2026-09-17 purge placeholders ─────────────────────────────────────
|
||||
# PR #112 replaced six live credentials with `«vault: <project>/<env> <SECRET>»`
|
||||
# references. Those references are safe by construction (they name where the
|
||||
# secret is read from), but they are listed here explicitly rather than being
|
||||
# filtered by a general "vault" rule, so a new occurrence still needs a
|
||||
# deliberate, reasoned entry.
|
||||
secret-assign litellm-api-keys.prose.md MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY» 2026-09-17 purge: replaced the live Mumuni LiteLLM key with its vault reference; no literal credential.
|
||||
secret-assign litellm-api-keys.prose.md MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY» 2026-09-17 purge: replaced the live Mumuni Zulip key with its vault reference; no literal credential.
|
||||
* infrastructure-control.prose.md PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
|
||||
cred-prose infrastructure-control.prose.md Admin credentials: 2026-09-17 purge: the Stirling admin user/password are two `«vault: ...»` references; no literal credential.
|
||||
* scripts/daily-infra-report.py PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
|
||||
secret-assign stirling-pdf-agent-access.prose.md «vault: infrastructure/production STIRLING_API_KEY» 2026-09-17 purge: Stirling PDF API key is read from the vault; the curl example only names the vault path.
|
||||
bearer-token agent-zero-fix-summary.md «vault: agents/production OPENROUTER_API_KEY» 2026-09-17 purge: OpenRouter key is read from the vault; the example curl only names the vault path.
|
||||
# ── Deliberate synthetic examples in contracts (not from the purge) ───────
|
||||
# These exist to teach the rule they illustrate. They are listed here so the
|
||||
# guard is never taught to skip the words "synthetic"/"example" — a fabricated
|
||||
# example is always an explicit exception, never a pattern-level exemption.
|
||||
* hermes-key-enforcement.prose.md sk-synthetic-external-example Rule 15 illustration of a hardcoded external key that is tolerated; fabricated, never a live key.
|
||||
* hermes-key-enforcement.prose.md sk-synthetic-example-12345 Rule 15 illustration of a forbidden hardcoded key; fabricated, never a live key.
|
||||
openai-key hermes-key-enforcement.prose.md sk-synthetic-litellm- Fabricated key name inside a `grep 'LITELLM_API_KEY=...'` example; not a live key.
|
||||
secret-assign hermes-key-enforcement.prose.md sk-NEW_KEY Placeholder standing for the rotated key in an `infisical secrets set` command; not a literal key.
|
||||
openrouter-key agent-zero-openrouter-key.prose.md sk-or-v1-synthetic Synthetic key prefix in the contract's example response; the real key is read from the vault.
|
||||
openai-key litellm-api-keys.prose.md sk-synthetic-tanko-example Fabricated key name in migration history prose; not a live key.
|
||||
openai-key litellm-self-heal.prose.md sk-syslog-local-master-key Deprecated local LiteLLM master key name documented as no-live-usage; kept for history, not a usable credential.
|
||||
# ── Redacted evidence, not a credential ──────────────────────────────────
|
||||
secret-assign docs/probe-drift-round2-evidence.md =sk-... Probe evidence records redacted key trailers (`sk-...x6uw`); the usable part of the key is not present.
|
||||
# ── tests/test_secret_scan.sh fixtures ───────────────────────────────────
|
||||
# The self-test plants these fabricated values into a TEMP tree, whose path no
|
||||
# entry here covers, so each still fails the guard when planted (see the test's
|
||||
# "... fails the guard" cases). They are listed only so the repo-wide scan of
|
||||
# the test file itself stays quiet.
|
||||
* tests/test_secret_scan.sh sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef Self-test fixture: fabricated OpenRouter-shaped key written to a temp tree; the guard must fail on it there.
|
||||
bearer-token tests/test_secret_scan.sh Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb Self-test fixture: fabricated Bearer token written to a temp tree; the guard must fail on it there.
|
||||
proxmox-token tests/test_secret_scan.sh PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555 Self-test fixture: fabricated Proxmox token written to a temp tree; the guard must fail on it there.
|
||||
private-key tests/test_secret_scan.sh -----BEGIN OPENSSH PRIVATE KEY----- Self-test fixture: fabricated PEM banner written to a temp tree; the guard must fail on it there.
|
||||
cred-prose tests/test_secret_scan.sh Admin credentials: Self-test fixture: fabricated prose credential line written to a temp tree; the guard must fail on it there.
|
||||
secret-assign tests/test_secret_scan.sh DB_PASSWORD=correct-horse-battery-staple Self-test fixture: fabricated password assignment written to a temp tree; the guard must fail on it there.
|
||||
|
Can't render this file because it contains an unexpected character in line 23 and column 25.
|
@@ -0,0 +1,22 @@
|
||||
# secret-patterns.tsv — checked-in pattern list for scripts/secret-scan.sh
|
||||
#
|
||||
# Format: <rule-id><TAB><POSIX ERE><TAB><description><TAB><check>
|
||||
# Blank lines and lines whose first field starts with '#' are ignored.
|
||||
# <check> is optional; the only value today is "value", which tells the scanner
|
||||
# to run the matched value through its inert-value classifier (see
|
||||
# value_is_inert in secret-scan.sh) so bare identifiers, env refs and dotted
|
||||
# code access are not reported as credentials. Omit the column to report every
|
||||
# regex hit.
|
||||
# Matching is case-insensitive, so `API_KEY` and `api_key` both count.
|
||||
#
|
||||
# Add a rule here, never inline in secret-scan.sh: this file is the single
|
||||
# auditable list of what the guard considers credential-shaped.
|
||||
openai-key \bsk-[A-Za-z0-9_-]{16,} OpenAI/LiteLLM-style "sk-" secret key (also hyphenated sk-proj- keys)
|
||||
openrouter-key \bsk-or-v1-[A-Za-z0-9_-]{8,} OpenRouter API key
|
||||
stripe-live-key \bsk_live_[A-Za-z0-9]{8,} Stripe live secret key
|
||||
proxmox-token PVEAPIToken=[^[:space:]"']+ Proxmox API token literal
|
||||
bearer-token bearer[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) literal Bearer token (http header or prose)
|
||||
auth-header authorization:[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) Authorization header carrying a raw literal value
|
||||
private-key -----BEGIN [A-Z ]*PRIVATE KEY----- PEM private key block
|
||||
cred-prose credentials?[[:space:]]*[:=][[:space:]]*[^[:space:]] prose credential line carrying a value
|
||||
secret-assign (api[_-]?key|apikey|passwd|password|secret|token)s?["']?[[:space:]]*[:=][[:space:]]*["']?(«.{3,}»|[A-Za-z0-9_./+=-]{8,}) credential assignment carrying a literal value value
|
||||
|
Can't render this file because it contains an unexpected character in line 5 and column 48.
|
Executable
+269
@@ -0,0 +1,269 @@
|
||||
#!/usr/bin/env bash
|
||||
# secret-scan.sh — commit-time secret guard. FAILS (exit 1) on a credential-shaped
|
||||
# string, so a build cannot go green with a credential committed to it.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/secret-scan.sh # scan the whole git-tracked tree (default)
|
||||
# scripts/secret-scan.sh --tree
|
||||
# scripts/secret-scan.sh --path DIR # scan an arbitrary directory (git not required)
|
||||
# scripts/secret-scan.sh --staged # scan added lines in the index (pre-commit)
|
||||
# scripts/secret-scan.sh --diff REF # scan added lines since REF (e.g. origin/master)
|
||||
# --quiet only print the verdict and findings, no per-mode banner
|
||||
#
|
||||
# Exit codes: 0 clean, 1 credential found, 2 usage/config error.
|
||||
#
|
||||
# Patterns live in scripts/secret-patterns.tsv
|
||||
# Exceptions live in scripts/secret-allowlist.tsv (every entry carries a reason;
|
||||
# a missing reason is a hard error, so the guard fails closed).
|
||||
#
|
||||
# Dependencies are deliberately bash + coreutils + grep + sed/awk + git. The
|
||||
# Gitea Actions runner executes job steps INSIDE the runner container, which
|
||||
# has no node and no python by default: keep this script free of both.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SELF_DIR=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||
ROOT=$(cd -- "$SELF_DIR/.." && pwd)
|
||||
PATTERNS_FILE="$SELF_DIR/secret-patterns.tsv"
|
||||
ALLOWLIST_FILE="$SELF_DIR/secret-allowlist.tsv"
|
||||
|
||||
# The guard's own definition files are not scannable content: the pattern list
|
||||
# necessarily contains the pattern text, and the allowlist necessarily contains
|
||||
# the allowed literals. Narrow, exact-path exclusion — not a wildcard.
|
||||
SELF_FILES=(
|
||||
"scripts/secret-scan.sh"
|
||||
"scripts/secret-patterns.tsv"
|
||||
"scripts/secret-allowlist.tsv"
|
||||
)
|
||||
|
||||
MODE="tree"
|
||||
PATH_DIR=""
|
||||
DIFF_REF=""
|
||||
QUIET=0
|
||||
|
||||
usage() {
|
||||
sed -n '2,20p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'
|
||||
exit 2
|
||||
}
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
--tree) MODE="tree" ;;
|
||||
--path) MODE="path"; PATH_DIR="${2:-}"; shift ;;
|
||||
--staged) MODE="staged" ;;
|
||||
--diff) MODE="diff"; DIFF_REF="${2:-}"; shift ;;
|
||||
--quiet) QUIET=1 ;;
|
||||
-h|--help) usage ;;
|
||||
*) echo "secret-scan: unknown argument '$1'" >&2; usage ;;
|
||||
esac
|
||||
shift
|
||||
done
|
||||
|
||||
[ -f "$PATTERNS_FILE" ] || { echo "secret-scan: missing $PATTERNS_FILE" >&2; exit 2; }
|
||||
[ -f "$ALLOWLIST_FILE" ] || { echo "secret-scan: missing $ALLOWLIST_FILE" >&2; exit 2; }
|
||||
if [ "$MODE" = "path" ] && [ -z "$PATH_DIR" ]; then
|
||||
echo "secret-scan: --path needs a directory" >&2; exit 2
|
||||
fi
|
||||
if [ "$MODE" = "diff" ] && [ -z "$DIFF_REF" ]; then
|
||||
echo "secret-scan: --diff needs a base ref" >&2; exit 2
|
||||
fi
|
||||
|
||||
# ── Load patterns ──────────────────────────────────────────────────────────
|
||||
RULE_IDS=()
|
||||
RULE_RES=()
|
||||
RULE_DESCS=()
|
||||
RULE_CHECKS=()
|
||||
COMBINED=""
|
||||
while IFS=$'\t' read -r id re desc check; do
|
||||
case "$id" in ''|'#'*) continue ;; esac
|
||||
[ -n "$re" ] || continue
|
||||
RULE_IDS+=("$id"); RULE_RES+=("$re"); RULE_DESCS+=("$desc"); RULE_CHECKS+=("${check:-}")
|
||||
if [ -z "$COMBINED" ]; then COMBINED="($re)"; else COMBINED="$COMBINED|($re)"; fi
|
||||
done < "$PATTERNS_FILE"
|
||||
if [ "${#RULE_IDS[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: no patterns loaded from $PATTERNS_FILE" >&2; exit 2
|
||||
fi
|
||||
|
||||
# ── Load allowlist (fails closed on a missing reason) ──────────────────────
|
||||
AL_RULES=()
|
||||
AL_GLOBS=()
|
||||
AL_LITS=()
|
||||
AL_REASONS=()
|
||||
AL_LINENO=0
|
||||
while IFS=$'\t' read -r rule glob lit reason; do
|
||||
AL_LINENO=$((AL_LINENO + 1))
|
||||
case "$rule" in ''|'#'*) continue ;; esac
|
||||
if [ -z "$glob" ] || [ -z "$lit" ] || [ -z "$reason" ]; then
|
||||
echo "secret-scan: ❌ $ALLOWLIST_FILE:$AL_LINENO — allowlist entry needs <rule> <path-glob> <literal> <reason>; reason-based exceptions only, refusing to run" >&2
|
||||
exit 2
|
||||
fi
|
||||
AL_RULES+=("$rule"); AL_GLOBS+=("$glob"); AL_LITS+=("$lit"); AL_REASONS+=("$reason")
|
||||
done < "$ALLOWLIST_FILE"
|
||||
|
||||
# nocasematch is toggled only around the regex test; path globs must stay
|
||||
# case-sensitive, so it is never left on.
|
||||
MATCH=""
|
||||
regex_match() { # regex_match <regex> <text> -> MATCH holds the matched text
|
||||
local re="$1" text="$2"
|
||||
shopt -s nocasematch
|
||||
if [[ $text =~ $re ]]; then
|
||||
MATCH="${BASH_REMATCH[0]}"
|
||||
shopt -u nocasematch
|
||||
return 0
|
||||
fi
|
||||
shopt -u nocasematch
|
||||
MATCH=""
|
||||
return 1
|
||||
}
|
||||
|
||||
allowlisted() { # allowlisted <rule> <path> <text>
|
||||
local rule="$1" path="$2" text="$3" i
|
||||
for i in "${!AL_RULES[@]}"; do
|
||||
[ "${AL_RULES[$i]}" = "$rule" ] || [ "${AL_RULES[$i]}" = "*" ] || continue
|
||||
# The unquoted RHS is deliberate: <path-glob> is a bash glob, not a literal.
|
||||
# shellcheck disable=SC2053
|
||||
[[ $path == ${AL_GLOBS[$i]} ]] || continue
|
||||
[[ $text == *"${AL_LITS[$i]}"* ]] || continue
|
||||
return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
mask_value() { # mask_value <text> <match> — never echo a credential to logs.
|
||||
# Print only the part of the line BEFORE the match, then <redacted>: the match
|
||||
# itself and everything after it (which may include a value the rule's regex
|
||||
# stopped short of, e.g. `credentials:` followed by a backticked password) is
|
||||
# never written to stdout.
|
||||
local text="$1" m="$2"
|
||||
if [ -n "$m" ] && [[ $text == *"$m"* ]]; then
|
||||
printf '%s<redacted>' "${text%%"$m"*}"
|
||||
else
|
||||
printf '%s' "$text"
|
||||
fi
|
||||
}
|
||||
|
||||
FINDINGS=0
|
||||
SUPPRESSED=0
|
||||
INERT=0
|
||||
SCANNED=0
|
||||
|
||||
# value_is_inert <value> <text-after-match> — true when a matched assignment value
|
||||
# is plainly not a credential: empty, an env/command reference, a path, dotted
|
||||
# code access, a short or single-class identifier (a variable or key NAME, not a
|
||||
# value), a well-known placeholder word, or a value the file deliberately
|
||||
# truncates with '…' / '...' (a redacted prefix is not a usable credential).
|
||||
# Deliberately does NOT know the words "synthetic" or "example": a fabricated
|
||||
# example must be an explicit allowlist entry.
|
||||
value_is_inert() {
|
||||
local v="$1" rest="$2"
|
||||
case "$rest" in '…'*|'...'*) return 0 ;; esac
|
||||
v="${v%\"}"; v="${v#\"}"; v="${v%\'}"; v="${v#\'}"
|
||||
case "$v" in
|
||||
''|\$*|\{*|'<'*|'%'*|'('*|'/'*|'\\'*) return 0 ;;
|
||||
not-needed|no-key-required|none|null|true|false|redacted|placeholder|example|dummy|changeme|change-me|your-key|your_key|key|token|secret|password) return 0 ;;
|
||||
esac
|
||||
# dotted code access: os.environ.get / process.env.ZULIP_API_KEY / cfg.a
|
||||
if [[ $v =~ ^[a-z_][a-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)+$ ]]; then return 0; fi
|
||||
# bare identifier (no punctuation beyond _): a NAME, not a value. A real
|
||||
# secret in this shape is long and mixes letters with digits.
|
||||
if [[ $v =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
|
||||
[ "${#v}" -lt 20 ] && return 0
|
||||
[[ $v =~ [0-9] ]] || return 0
|
||||
return 1
|
||||
fi
|
||||
return 1
|
||||
}
|
||||
|
||||
report_finding() { # report_finding <path> <line> <text>
|
||||
local path="$1" line="$2" text="$3" i val
|
||||
for i in "${!RULE_IDS[@]}"; do
|
||||
regex_match "${RULE_RES[$i]}" "$text" || continue
|
||||
SCANNED=$((SCANNED + 1))
|
||||
if [ "${RULE_CHECKS[$i]}" = "value" ]; then
|
||||
val="${MATCH#*[:=]}"
|
||||
val="${val# }"
|
||||
if value_is_inert "$val" "${text#*"$MATCH"}"; then
|
||||
INERT=$((INERT + 1))
|
||||
continue
|
||||
fi
|
||||
fi
|
||||
if allowlisted "${RULE_IDS[$i]}" "$path" "$text"; then
|
||||
SUPPRESSED=$((SUPPRESSED + 1))
|
||||
continue
|
||||
fi
|
||||
FINDINGS=$((FINDINGS + 1))
|
||||
printf ' ❌ %s:%s [%s] %s\n' "$path" "$line" "${RULE_IDS[$i]}" "${RULE_DESCS[$i]}"
|
||||
printf ' | %s\n' "$(mask_value "$text" "$MATCH")"
|
||||
done
|
||||
}
|
||||
|
||||
self_excluded() { # self_excluded <repo-relative-path>
|
||||
local p="$1" s
|
||||
for s in "${SELF_FILES[@]}"; do
|
||||
[ "$p" = "$s" ] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
# ── Collect candidate lines and scan them ─────────────────────────────────
|
||||
if [ "$MODE" = "tree" ] || [ "$MODE" = "path" ]; then
|
||||
if [ "$MODE" = "tree" ]; then
|
||||
BASE="$ROOT"
|
||||
git -C "$BASE" rev-parse --git-dir >/dev/null 2>&1 || { echo "secret-scan: --tree needs a git checkout (use --path DIR)" >&2; exit 2; }
|
||||
mapfile -d '' candidate < <(git -C "$BASE" ls-files -z 2>/dev/null)
|
||||
if [ "${#candidate[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: ❌ no tracked files — refusing to report clean" >&2; exit 2
|
||||
fi
|
||||
else
|
||||
BASE=$(cd -- "$PATH_DIR" 2>/dev/null && pwd) || { echo "secret-scan: --path '$PATH_DIR' is not a directory" >&2; exit 2; }
|
||||
mapfile -t candidate < <(cd -- "$BASE" && find . -type f -not -path './.git/*' | sed 's|^\./||')
|
||||
if [ "${#candidate[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: ❌ no files under $BASE — refusing to report clean" >&2; exit 2
|
||||
fi
|
||||
fi
|
||||
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan ($MODE): ${#candidate[@]} files under $BASE ──"
|
||||
for rel in "${candidate[@]}"; do
|
||||
[ -f "$BASE/$rel" ] || continue
|
||||
self_excluded "$rel" && continue
|
||||
while IFS= read -r hit; do
|
||||
[ -n "$hit" ] || continue
|
||||
report_finding "$rel" "${hit%%:*}" "${hit#*:}"
|
||||
done < <(grep -nEIi -e "$COMBINED" "$BASE/$rel" 2>/dev/null || true)
|
||||
done
|
||||
else
|
||||
# --staged / --diff: only ADDED lines, with the post-change line number.
|
||||
if [ "$MODE" = "staged" ]; then
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines in the index ──"
|
||||
DIFF_TEXT=$(git -C "$ROOT" diff --cached --unified=0 --no-color -- . 2>/dev/null)
|
||||
else
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines since $DIFF_REF ──"
|
||||
DIFF_TEXT=$(git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"...HEAD 2>/dev/null \
|
||||
|| git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"..HEAD 2>/dev/null)
|
||||
fi
|
||||
if [ -z "$DIFF_TEXT" ]; then
|
||||
[ "$QUIET" -eq 1 ] || echo " (no added lines)"
|
||||
fi
|
||||
while IFS=$'\t' read -r rel line text; do
|
||||
[ -n "$rel" ] || continue
|
||||
self_excluded "$rel" && continue
|
||||
report_finding "$rel" "$line" "$text"
|
||||
done < <(printf '%s\n' "$DIFF_TEXT" | awk '
|
||||
/^\+\+\+ / { f=$2; sub(/^b\//,"",f); next }
|
||||
/^@@ / { if (match($0, /\+[0-9]+/)) ln=substr($0, RSTART+1, RLENGTH-1)+0; next }
|
||||
(/^\+/ && !/^\+\+\+/) { print f "\t" ln "\t" substr($0,2); ln++; next }
|
||||
')
|
||||
fi
|
||||
|
||||
# ── Verdict ────────────────────────────────────────────────────────────────
|
||||
if [ "$FINDINGS" -gt 0 ]; then
|
||||
echo ""
|
||||
echo "❌ SECRET SCAN FAILED — $FINDINGS credential-shaped string(s) in ${MODE} content."
|
||||
echo " Fix: remove the credential and read it from the vault/env."
|
||||
echo " Only a deliberate synthetic example may be added to scripts/secret-allowlist.tsv,"
|
||||
echo " one entry per file/rule/literal, with a reason. Never allowlist a live credential."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "✅ secret scan clean (${MODE}; ${SUPPRESSED} allowlisted exception(s), ${INERT} inert value(s) ignored)"
|
||||
exit 0
|
||||
Executable
+228
@@ -0,0 +1,228 @@
|
||||
#!/bin/bash
|
||||
# test_infra_monitoring.sh — Asserts that probe calls use the documented targets.
|
||||
#
|
||||
# Strategy: stub curl and ssh on PATH to capture the exact arguments each leg
|
||||
# builds, then assert the URL/port of every call. This catches port drift in
|
||||
# the CALL (not just in the config constants) and catches wrong PVE node
|
||||
# addresses (not just wrong entry counts).
|
||||
#
|
||||
# Run: bash scripts/test_infra_monitoring.sh
|
||||
# Exits 0 if all assertions pass, 1 otherwise.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
SCRIPT="${SCRIPT_DIR}/infra-monitoring.sh"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
assert() {
|
||||
local desc="$1" condition="$2"
|
||||
if eval "$condition"; then
|
||||
echo " ✅ $desc"
|
||||
PASS=$((PASS+1))
|
||||
else
|
||||
echo " 🔴 $desc"
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
}
|
||||
|
||||
echo "=== test_infra_monitoring.sh ==="
|
||||
echo ""
|
||||
|
||||
# ── Stub curl: capture argv to a file, return 200 ──────────────────────────
|
||||
STUB_DIR=$(mktemp -d)
|
||||
trap 'rm -rf "$STUB_DIR"' EXIT
|
||||
|
||||
# Stub curl: first arg after flags is the URL; capture all args
|
||||
cat > "$STUB_DIR/curl" << 'STUBEOF'
|
||||
#!/bin/bash
|
||||
echo "$@" >> "${CURL_STUB_LOG:-/dev/null}"
|
||||
# Print 200 for %{http_code}
|
||||
printf '%s\n' "200"
|
||||
exit 0
|
||||
STUBEOF
|
||||
chmod +x "$STUB_DIR/curl"
|
||||
|
||||
# Stub ssh: first arg after options is the remote command; capture it
|
||||
cat > "$STUB_DIR/ssh" << 'SSTUBEOF'
|
||||
#!/bin/bash
|
||||
echo "SSH $@" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||
# The last arg is the remote command — extract and log curl args
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == curl* ]]; then
|
||||
echo "$arg" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||
fi
|
||||
done
|
||||
printf '%s\n' "200"
|
||||
exit 0
|
||||
SSTUBEOF
|
||||
chmod +x "$STUB_DIR/ssh"
|
||||
|
||||
# ── Run the monitor with stubs ────────────────────────────────────────────
|
||||
CURL_LOG="$STUB_DIR/curl_calls.log"
|
||||
SSH_LOG="$STUB_DIR/ssh_calls.log"
|
||||
touch "$CURL_LOG" "$SSH_LOG"
|
||||
|
||||
CURL_STUB_LOG="$CURL_LOG" SSH_STUB_LOG="$SSH_LOG" \
|
||||
PATH="$STUB_DIR:$PATH" bash "$SCRIPT" > "$STUB_DIR/output.txt" 2>&1
|
||||
|
||||
# ── 1. Port drift detection (from actual curl invocations) ─────────────────
|
||||
|
||||
assert "Grafana probed at port 3001" \
|
||||
'grep -q "http://192.168.68.116:3001/api/health" "$CURL_LOG"'
|
||||
|
||||
assert "Prometheus probed at port 9090" \
|
||||
'grep -q "http://192.168.68.116:9090/-/healthy" "$CURL_LOG"'
|
||||
|
||||
assert "LiteLLM probed via nginx at port 80" \
|
||||
'grep -q "http://192.168.68.116:80/litellm/health" "$CURL_LOG"'
|
||||
|
||||
assert "PVE API probed at port 8006" \
|
||||
'grep -q ":8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "GPU exporter probed at port 9400" \
|
||||
'grep -q ":9400/metrics" "$CURL_LOG"'
|
||||
|
||||
# ── 2. PVE API: exact node addresses (catches wrong IPs) ──────────────────
|
||||
# Each real PVE node must be probed; CT 116 must NOT be in the PVE set
|
||||
|
||||
assert "PVE acerpve 192.168.68.9 probed" \
|
||||
'grep -q "https://192.168.68.9:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE minipve 192.168.68.12 probed" \
|
||||
'grep -q "https://192.168.68.12:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE storepve 192.168.68.6 probed" \
|
||||
'grep -q "https://192.168.68.6:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE amdpve 192.168.68.15 probed" \
|
||||
'grep -q "https://192.168.68.15:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE ocupve 192.168.68.5 probed" \
|
||||
'grep -q "https://192.168.68.5:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
# CT 116 (.116) must NOT appear as a PVE API target
|
||||
assert "CT 116 (.116) NOT probed as PVE API node" \
|
||||
'! grep -q "https://192.168.68.116:8006" "$CURL_LOG"'
|
||||
|
||||
# ── 3. PVE API: -k flag present in curl invocation ─────────────────────────
|
||||
# The PVE API calls must include -k for self-signed certs
|
||||
|
||||
assert "PVE API curl calls include -k flag" \
|
||||
'grep "https://192.168.68.9:8006" "$CURL_LOG" | grep -q -- "-k"'
|
||||
|
||||
# ── 4. Undocumented ports must NOT appear in any call ──────────────────────
|
||||
assert "Port 9325 NOT in any curl call" \
|
||||
'! grep -q ":9325" "$CURL_LOG"'
|
||||
|
||||
assert "Port 9405 NOT in any curl call" \
|
||||
'! grep -q ":9405" "$CURL_LOG"'
|
||||
|
||||
# ── 5. Docker Stats / PVE Exporter: SSH-probed at correct ports ────────────
|
||||
assert "Docker Stats probed at port 9324 via SSH" \
|
||||
'grep -q "9324" "$SSH_LOG"'
|
||||
|
||||
assert "PVE Exporter probed at port 9221 via SSH" \
|
||||
'grep -q "9221" "$SSH_LOG"'
|
||||
|
||||
# ── 5a. Per-leg assertions (proves which leg owns which port) ──────────────
|
||||
# The script source must show Docker Stats using $DOCKER_STATS_PORT and
|
||||
# PVE Exporter using $PVE_EXPORTER_PORT in the correct leg sections
|
||||
assert "Docker Stats leg uses DOCKER_STATS_PORT constant" \
|
||||
'grep -A 3 "# 6. Docker Stats" "$SCRIPT" | grep -q "\$DOCKER_STATS_PORT"'
|
||||
|
||||
assert "PVE Exporter leg uses PVE_EXPORTER_PORT constant" \
|
||||
'grep -A 3 "# 7. PVE Exporter" "$SCRIPT" | grep -q "\$PVE_EXPORTER_PORT"'
|
||||
|
||||
# Verify the constants themselves are set to the correct values
|
||||
assert "DOCKER_STATS_PORT constant set to 9324" \
|
||||
'grep -q "^DOCKER_STATS_PORT=\"9324\"" "$SCRIPT"'
|
||||
|
||||
assert "PVE_EXPORTER_PORT constant set to 9221" \
|
||||
'grep -q "^PVE_EXPORTER_PORT=\"9221\"" "$SCRIPT"'
|
||||
|
||||
# ── 5b. Stale port 9323 (dockerd) NOT probed ──────────────────────────────
|
||||
assert "Port 9323 (dockerd) NOT in SSH log" \
|
||||
'! grep -q "9323" "$SSH_LOG"'
|
||||
|
||||
# ── 6. No stale ports in the script source (belt-and-suspenders) ──────────
|
||||
assert "Port 9325 (historical) NOT in script source" \
|
||||
'! grep -q "9325" "$SCRIPT"'
|
||||
|
||||
assert "Port 9405 (historical) NOT in script source" \
|
||||
'! grep -q "9405" "$SCRIPT"'
|
||||
|
||||
# ── 7. Failure-line content includes non-empty kind ────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||
|
||||
# Test 7a: Unexpected status (500) → kind should be unexpected:500
|
||||
cat > "$TMP_DIR/curl" << 'EOF'
|
||||
#!/bin/bash
|
||||
# Stub: return 500 for Grafana port, 200 otherwise
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == *":3001"* ]]; then
|
||||
echo "500"
|
||||
exit 0
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/curl"
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||
assert "Grafana failure line exists (unexpected status)" \
|
||||
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||
assert "Grafana failure kind is non-empty (unexpected status)" \
|
||||
'[[ -n "$KIND" ]]'
|
||||
|
||||
# Test 7b: TLS error (000 + exit 60) → kind should be tls
|
||||
cat > "$TMP_DIR/curl" << 'EOF'
|
||||
#!/bin/bash
|
||||
# Stub: return 000 and exit 60 for Grafana port (TLS error) on ALL invocations
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == *":3001"* ]]; then
|
||||
echo "000"
|
||||
exit 60
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/curl"
|
||||
# Also stub ssh to return 000 + exit 60 for the retry
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == curl* ]]; then
|
||||
echo "000"
|
||||
exit 60
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
export PATH="$TMP_DIR:$PATH"
|
||||
OUT=$(bash "$SCRIPT" 2>&1)
|
||||
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||
assert "Grafana failure line exists (TLS error)" \
|
||||
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||
assert "Grafana failure kind is tls" \
|
||||
'[[ "$KIND" == "tls" ]]'
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||
if [ $FAIL -gt 0 ]; then
|
||||
echo " 🔴 TESTS FAILED"
|
||||
exit 1
|
||||
else
|
||||
echo " ✅ ALL TESTS PASSED"
|
||||
exit 0
|
||||
fi
|
||||
Executable
+125
@@ -0,0 +1,125 @@
|
||||
#!/bin/bash
|
||||
# test_proxmox_monitor.sh — Stub-driven tests for PBS GC liveness leg
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT="$(cd "$(dirname "$0")" && pwd)/proxmox-monitor.sh"
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
# ── Helpers ────────────────────────────────────────────────────────────────
|
||||
assert() {
|
||||
local desc="$1" cond="$2"
|
||||
if eval "$cond" 2>/dev/null; then
|
||||
echo " ✅ $desc"
|
||||
PASS=$((PASS+1))
|
||||
else
|
||||
echo " 🔴 $desc"
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
}
|
||||
|
||||
# ── 1. Healthy: fresh GC, 0 B pending ─────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
FRESH_ENDTIME=$(( $(date -u +%s) - (1 * 3600) )) # 1 hour ago
|
||||
cat > "$TMP_DIR/ssh" << EOF
|
||||
#!/bin/bash
|
||||
# Stub: return valid JSON with fresh endtime
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":$FRESH_ENDTIME,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Healthy: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Healthy: shows healthy verdict" '[[ "$PBS_LINE" == *"healthy"* ]]'
|
||||
assert "Healthy: shows pending-bytes 0 B" '[[ "$PBS_LINE" == *"pending-bytes: 0 B"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 2. Stale: GC >48h old ─────────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
OLD_ENDTIME=$(( $(date -u +%s) - (49 * 3600) )) # 49 hours ago
|
||||
cat > "$TMP_DIR/ssh" << EOF
|
||||
#!/bin/bash
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":$OLD_ENDTIME,"pending-bytes":1048576}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Stale: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Stale: shows stale verdict" '[[ "$PBS_LINE" == *"stale"* ]]'
|
||||
assert "Stale: shows pending-bytes 1048576 B" '[[ "$PBS_LINE" == *"pending-bytes: 1048576 B"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 3. Probe-failed: empty output ─────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Probe-failed (empty): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Probe-failed (empty): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 4. Probe-failed: unparseable output ───────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo "proxmox-backup-manager: command not found"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Probe-failed (unparseable): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Probe-failed (unparseable): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 5. Null endtime: never-run ────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":null,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Null endtime: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Null endtime: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 6. Datastore absent ───────────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo '[{"store":"s3-archive","last-run-endtime":null,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Datastore absent: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Datastore absent: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── Summary ────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||
if [ $FAIL -gt 0 ]; then
|
||||
echo " 🔴 TESTS FAILED"
|
||||
exit 1
|
||||
else
|
||||
echo " ✅ ALL TESTS PASSED"
|
||||
exit 0
|
||||
fi
|
||||
+197
-111
@@ -1,57 +1,66 @@
|
||||
#!/bin/bash
|
||||
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
|
||||
# Implements zulip-health.prose.md v2
|
||||
# Runs every 15 min via cron. Alerts via Telegram.
|
||||
# Implements zulip-health.prose.md v3
|
||||
# Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'.
|
||||
# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B
|
||||
# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes
|
||||
# agent leg is retired — see the note after the Tanko leg.
|
||||
set -euo pipefail
|
||||
|
||||
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
|
||||
# Never fall back to a literal key.
|
||||
# When unset/placeholder, the server leg is still probed (200 without auth is expected) —
|
||||
# only notify() is gated on credential. The pi/Tanko/kagentz
|
||||
# legs do not need the Zulip API key. The placeholder is captain-held:
|
||||
# zulip-health-credential-placeholder-20260913.
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||
ZULIP_SITE="https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
OWNER_ZULIP_ID="9"
|
||||
|
||||
# Email config
|
||||
GMAIL_USER="jtabiri@gmail.com"
|
||||
GMAIL_PASS="rgbuomwcydxwbszd"
|
||||
EMAIL_TO="jerome@sysloggh.com"
|
||||
# Track whether the Zulip API credential is usable
|
||||
ZULIP_CRED_OK=1
|
||||
if [ -z "$ZULIP_API_KEY" ] || [[ "$ZULIP_API_KEY" == *"placeholder"* ]] || [[ "$ZULIP_API_KEY" == *"REDACTED"* ]]; then
|
||||
ZULIP_CRED_OK=0
|
||||
fi
|
||||
|
||||
LOG="/root/zulip-health-monitor.log"
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
ISSUES=0
|
||||
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
||||
|
||||
notify() {
|
||||
local severity="$1" msg="$2"
|
||||
echo "[$severity] $msg"
|
||||
|
||||
# Zulip DM to owner
|
||||
local content="${severity} Zulip Monitor: ${msg}"
|
||||
local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
|
||||
# Email alert
|
||||
local subject="${severity} Zulip Monitor Alert"
|
||||
python3 -c "
|
||||
import smtplib
|
||||
from email.mime.text import MIMEText
|
||||
m = MIMEText('''${msg}''')
|
||||
m['From'] = 'abiba@sysloggh.com'
|
||||
m['To'] = '${EMAIL_TO}'
|
||||
m['Subject'] = '${subject}'
|
||||
s = smtplib.SMTP('smtp.gmail.com', 587)
|
||||
s.starttls()
|
||||
s.login('${GMAIL_USER}', '${GMAIL_PASS}')
|
||||
s.sendmail('abiba@sysloggh.com', ['${EMAIL_TO}'], m.as_string())
|
||||
s.quit()
|
||||
" 2>/dev/null || true
|
||||
# Zulip DM to owner (skip if no credential)
|
||||
if [ "$ZULIP_CRED_OK" -eq 1 ]; then
|
||||
local content="${severity} Zulip Monitor: ${msg}"
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||
> /dev/null 2>&1 \
|
||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)">> "$LOG"
|
||||
else
|
||||
echo " ALERT SUPPRESSED (no credential): ${severity} ${msg}" >> "$LOG"
|
||||
fi
|
||||
}
|
||||
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
ISSUES=0
|
||||
LOG="/root/zulip-health-monitor.log"
|
||||
|
||||
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
||||
|
||||
# ── Global: Zulip Server ──
|
||||
# F3: Always probe server regardless of credential — 200 without auth is expected
|
||||
# (verified live: server_settings returns 200 with no credential or wrong key).
|
||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||
https://chat.sysloggh.net/api/v1/server_settings \
|
||||
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null || echo "000")
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
|
||||
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
|
||||
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
|
||||
if [ "$SERVER_CODE" != "200" ]; then
|
||||
notify "🔴" "Zulip server returned HTTP $SERVER_CODE"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
@@ -60,96 +69,173 @@ else
|
||||
fi
|
||||
|
||||
# ── Platform A: pi (Abiba) ──
|
||||
PI_HEALTH=$(curl -sf --connect-timeout 5 http://localhost:9200/health 2>/dev/null || echo "{}")
|
||||
PI_CONNECTED=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('connected',False))" 2>/dev/null)
|
||||
PI_ERROR=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('last_error') or '')" 2>/dev/null)
|
||||
PI_RETRIES=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('retry_count',0))" 2>/dev/null)
|
||||
# Probes the pi Zulip extension health endpoint (:9200/health, served by the
|
||||
# extension's startHealthServer; shape documented in zulip-health.prose.md).
|
||||
# FAIL-SAFE contract (pinned by tests/zulip-monitor-abiba.sh): connection state
|
||||
# lives NESTED at zulip.connected / zulip.last_error — there is no top-level
|
||||
# `connected` and no retry counter in the payload. A fetch error, non-2xx
|
||||
# response, empty/unparseable body, or payload missing a boolean
|
||||
# zulip.connected is a PROBE FAILURE: it alerts and NEVER calls pm2 restart.
|
||||
# pm2 restart runs ONLY on affirmative zulip.connected=false.
|
||||
# -- abiba-leg-start (verbatim-extracted by tests/zulip-monitor-abiba.sh)
|
||||
PI_HTTP=$(curl -s -o /dev/null --connect-timeout 5 --max-time 10 -w '%{http_code}' http://localhost:9200/health 2>/dev/null) || PI_HTTP="000"
|
||||
PI_HTTP=$(printf '%s' "$PI_HTTP" | tr -d '[:space:]')
|
||||
[ -n "$PI_HTTP" ] || PI_HTTP="000"
|
||||
PI_BODY=$(curl -s --connect-timeout 5 --max-time 10 http://localhost:9200/health 2>/dev/null || true)
|
||||
PI_STATE=$(printf '%s' "$PI_BODY" | python3 -c '
|
||||
import sys, json
|
||||
code = sys.argv[1]
|
||||
body = sys.stdin.read()
|
||||
try:
|
||||
d = json.loads(body)
|
||||
except Exception:
|
||||
sys.stdout.write("probe-failed|unparseable body")
|
||||
sys.exit(0)
|
||||
if not code.startswith("2"):
|
||||
sys.stdout.write("probe-failed|HTTP %s" % code)
|
||||
sys.exit(0)
|
||||
if not isinstance(d, dict) or not isinstance(d.get("zulip"), dict):
|
||||
sys.stdout.write("probe-failed|missing zulip.connected")
|
||||
sys.exit(0)
|
||||
z = d["zulip"]
|
||||
if "connected" not in z or not isinstance(z["connected"], bool):
|
||||
sys.stdout.write("probe-failed|missing or non-boolean zulip.connected")
|
||||
sys.exit(0)
|
||||
err = z.get("last_error") or ""
|
||||
if z["connected"]:
|
||||
if err:
|
||||
sys.stdout.write("degraded|%s" % err)
|
||||
else:
|
||||
sys.stdout.write("healthy|%s" % z.get("messages_processed", 0))
|
||||
else:
|
||||
sys.stdout.write("disconnected|")
|
||||
' "$PI_HTTP" 2>/dev/null) || PI_STATE="probe-failed|python error"
|
||||
PI_VERDICT=${PI_STATE%%|*}
|
||||
PI_DETAIL=${PI_STATE#*|}
|
||||
|
||||
if [ "$PI_CONNECTED" != "True" ]; then
|
||||
notify "🔴" "Abiba pi extension DISCONNECTED — restarting"
|
||||
pm2 restart abiba-zulip 2>/dev/null || true
|
||||
case "$PI_VERDICT" in
|
||||
healthy)
|
||||
echo " Abiba: ✅ Connected (processed=$PI_DETAIL)" >> "$LOG" ;;
|
||||
degraded)
|
||||
notify "🟡" "Abiba pi extension error: ${PI_DETAIL:0:100}"
|
||||
echo " Abiba: 🟡 Error: ${PI_DETAIL:0:100}" >> "$LOG" ;;
|
||||
disconnected)
|
||||
notify "🔴" "Abiba pi extension DISCONNECTED — restarting"
|
||||
pm2 restart abiba-zulip 2>/dev/null || true
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Abiba: ❌ Disconnected — restarted" >> "$LOG" ;;
|
||||
probe-failed)
|
||||
notify "🟠" "Abiba pi extension health probe FAILED (${PI_DETAIL}; HTTP $PI_HTTP) — NOT restarting, manual check needed"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Abiba: ⚠️ Probe failed (${PI_DETAIL}; HTTP $PI_HTTP) — NOT restarted" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟠" "Abiba pi extension health probe returned unexpected verdict (${PI_STATE}) — NOT restarting, manual check needed"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Abiba: ⚠️ Unexpected probe verdict (${PI_STATE}) — NOT restarted" >> "$LOG" ;;
|
||||
esac
|
||||
# -- abiba-leg-end
|
||||
|
||||
# ── Platform B: Tanko (DSH dsh-web on amdpve CT 112) ──
|
||||
# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker
|
||||
# key availability varies — so probes run from the amdpve vantage via `pct exec`.
|
||||
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on amdpve
|
||||
# (192.168.68.15). The gateway binds 127.0.0.1:3080 loopback-only by design — a
|
||||
# remote :3080 probe is refused and is NOT a fault.
|
||||
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
||||
"pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true)
|
||||
[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown"
|
||||
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
||||
"pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true)
|
||||
[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000"
|
||||
|
||||
if [ "$TANKO_SVC" != "active" ]; then
|
||||
notify "🔴" "Tanko (DSH dsh-web) service state: $TANKO_SVC — needs restart"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Abiba: ❌ Disconnected — restarted" >> "$LOG"
|
||||
elif [ -n "$PI_ERROR" ]; then
|
||||
notify "🟡" "Abiba pi extension error: ${PI_ERROR:0:100}"
|
||||
echo " Abiba: 🟡 Error: ${PI_ERROR:0:100}" >> "$LOG"
|
||||
elif [ "$PI_RETRIES" -ge 3 ]; then
|
||||
notify "🟡" "Abiba pi extension: $PI_RETRIES retries — restarting"
|
||||
pm2 restart abiba-zulip 2>/dev/null || true
|
||||
echo " Abiba: 🟡 $PI_RETRIES retries — restarted" >> "$LOG"
|
||||
echo " Tanko: ❌ service=$TANKO_SVC" >> "$LOG"
|
||||
elif [ "$TANKO_HTTP" = "000" ]; then
|
||||
notify "🔴" "Tanko (DSH dsh-web) HTTP :3080 connection refused/timeout — needs restart"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Tanko: ❌ http=000 (refused/timeout)" >> "$LOG"
|
||||
else
|
||||
echo " Abiba: ✅ Connected (processed=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('messages_processed',0))" 2>/dev/null))" >> "$LOG"
|
||||
case "$TANKO_HTTP" in
|
||||
200|301|302|307|308|401|403)
|
||||
echo " Tanko: ✅ service=active http=$TANKO_HTTP" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "Tanko (DSH dsh-web) HTTP :3080 answered $TANKO_HTTP — running, unexpected status"
|
||||
echo " Tanko: 🟡 service=active http=$TANKO_HTTP (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# ── Platform B: Hermes (Tanko) ──
|
||||
TANKO_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 jerome@192.168.68.122 \
|
||||
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
|
||||
TANKO_ZULIP=$(echo "$TANKO_STATE" | python3 -c "
|
||||
import sys,json
|
||||
d=json.load(sys.stdin)
|
||||
p=d.get('platforms',{}).get('zulip',{})
|
||||
print(p.get('state','unknown'))
|
||||
" 2>/dev/null)
|
||||
|
||||
if [ "$TANKO_ZULIP" != "connected" ]; then
|
||||
notify "🔴" "Tanko (Hermes) Zulip state: $TANKO_ZULIP — needs restart"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Tanko: ❌ state=$TANKO_ZULIP" >> "$LOG"
|
||||
else
|
||||
echo " Tanko: ✅ Zulip connected" >> "$LOG"
|
||||
fi
|
||||
|
||||
# ── Platform B: Hermes (Mumuni) ──
|
||||
MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \
|
||||
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
|
||||
MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c "
|
||||
import sys,json
|
||||
d=json.load(sys.stdin)
|
||||
p=d.get('platforms',{}).get('zulip',{})
|
||||
print(p.get('state','unknown'))
|
||||
" 2>/dev/null)
|
||||
|
||||
if [ "$MUMUNI_ZULIP" != "connected" ]; then
|
||||
notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG"
|
||||
else
|
||||
echo " Mumuni: ✅ Zulip connected" >> "$LOG"
|
||||
fi
|
||||
# ── Removed: the former "Platform B: Hermes" agent leg ──
|
||||
# Captain ruling 2026-09-10: that agent moved off this host onto her own
|
||||
# container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now
|
||||
# monitored on her side — see the out-of-scope note in zulip-health.prose.md.
|
||||
# The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway
|
||||
# state, which reported "unknown" on every run and posted a false 🔴 DM plus an
|
||||
# #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must
|
||||
# never contact her former host.
|
||||
|
||||
# ── Platform C: Agent Zero (kagentz) ──
|
||||
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json 2>/dev/null" 2>/dev/null || echo "")
|
||||
AZ_ALIVE=$(echo "$AZ_A2A" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('name',''))" 2>/dev/null)
|
||||
# C1: A2A liveness (no credential needed) — probes the container's internal :80/a2a/
|
||||
# C2: A2A response verification (needs LITELLM_KEY) — probes POST /a2a with auth
|
||||
# C3: Public access path (no credential needed) — probes https://kagentz.sysloggh.net/
|
||||
|
||||
if [ "$AZ_ALIVE" != "kagentz" ]; then
|
||||
notify "🔴" "kagentz A2A server DOWN — restarting"
|
||||
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero bash -c 'pkill -9 -f a2a_agent; sleep 1; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &'" 2>/dev/null || true
|
||||
# C1: A2A liveness (container-internal probe)
|
||||
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
|
||||
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||
[ -n "$AZ_A2A_CODE" ] || AZ_A2A_CODE="000"
|
||||
|
||||
if [ "$AZ_A2A_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz A2A server DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: ❌ A2A down — restarted" >> "$LOG"
|
||||
echo " kagentz C1: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||
else
|
||||
echo " kagentz: ✅ A2A alive" >> "$LOG"
|
||||
|
||||
# Check adapter process
|
||||
AZ_ADAPTER=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero ps aux 2>/dev/null | grep adapter | grep -v grep | wc -l" 2>/dev/null || echo "0")
|
||||
if [ "$AZ_ADAPTER" -lt 1 ]; then
|
||||
notify "🔴" "kagentz Zulip adapter DOWN — restarting"
|
||||
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero bash -c 'cd /a0/usr/kagentz-zulip && ZULIP_SITE=https://chat.sysloggh.net ZULIP_EMAIL=kagentz-bot@chat.sysloggh.net ZULIP_API_KEY=E9q9PXJTxftPYBkb5pBDWupDO7KK21ty ZULIP_AGENT_NAME=kagentz A2A_URL=http://localhost:8001/a2a A2A_TOKEN=8zNgdOEXzYxjQvTl /opt/venv-a0/bin/python3 -u adapter.py > /tmp/zulip-adapter.log 2>&1 &'" 2>/dev/null || true
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: ❌ Adapter down — restarted" >> "$LOG"
|
||||
else
|
||||
echo " kagentz: ✅ Adapter running" >> "$LOG"
|
||||
fi
|
||||
case "$AZ_A2A_CODE" in
|
||||
200|401)
|
||||
echo " kagentz C1: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C1: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# C3: Public access path (the captain's point of view)
|
||||
# Probes the public URL that NetBird proxies to the container. 200/302/401 = alive,
|
||||
# 502 = proxy's "upstream refused" page (incident), connection failed = incident.
|
||||
# Never restarts anything — the contract forbids restarting the platform.
|
||||
KAGENTZ_PUBLIC_CODE=$(curl -s -o /dev/null --connect-timeout 10 --max-time 15 \
|
||||
-w '%{http_code}' https://kagentz.sysloggh.net/ 2>/dev/null) || KAGENTZ_PUBLIC_CODE="000"
|
||||
KAGENTZ_PUBLIC_CODE=$(printf '%s' "$KAGENTZ_PUBLIC_CODE" | tr -d '[:space:]')
|
||||
[ -n "$KAGENTZ_PUBLIC_CODE" ] || KAGENTZ_PUBLIC_CODE="000"
|
||||
|
||||
if [ "$KAGENTZ_PUBLIC_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz public URL DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: ❌ public URL down (HTTP 000)" >> "$LOG"
|
||||
elif [ "$KAGENTZ_PUBLIC_CODE" = "502" ]; then
|
||||
notify "🔴" "kagentz public URL 502 (upstream refused)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: ❌ public URL 502 (upstream refused)" >> "$LOG"
|
||||
else
|
||||
case "$KAGENTZ_PUBLIC_CODE" in
|
||||
200|302|401)
|
||||
echo " kagentz C3: ✅ public URL alive (HTTP $KAGENTZ_PUBLIC_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz public URL answered HTTP $KAGENTZ_PUBLIC_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: 🟡 public URL unexpected http=$KAGENTZ_PUBLIC_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# ── Summary ──
|
||||
# The run verdict is non-optimistic: when ISSUES > 0, the run is an INCIDENT.
|
||||
# The lane must quote this Result line verbatim in its status report.
|
||||
if [ "$ISSUES" -eq 0 ]; then
|
||||
echo " Result: ✅ All healthy" >> "$LOG"
|
||||
echo " Result: ✅ 0 issues (all healthy)" >> "$LOG"
|
||||
else
|
||||
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
||||
echo " Result: 🔴 INCIDENT — $ISSUES issue(s) found" >> "$LOG"
|
||||
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
||||
fi
|
||||
|
||||
|
||||
@@ -27,7 +27,7 @@ Agent (Hermes/pi) ──curl + X-API-Key──► Stirling-PDF (:8989) ──►
|
||||
| Base URL | `http://192.168.68.7:8989` |
|
||||
| Auth Method | API Key (header) |
|
||||
| Header Name | `X-API-Key` |
|
||||
| API Key | `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88` |
|
||||
| API Key | `«vault: infrastructure/production STIRLING_API_KEY»` |
|
||||
| Key Source | `SECURITY_CUSTOMGLOBALAPIKEY` in `/opt/home_stack/docker-compose.yml` |
|
||||
| Swagger | `http://192.168.68.7:8989/swagger-ui.html` |
|
||||
| Health | `http://192.168.68.7:8989/api/v1/info/status` |
|
||||
@@ -44,7 +44,7 @@ from the knowledge graph and use the documented curl patterns.
|
||||
Direct bash invocations:
|
||||
```bash
|
||||
curl -X POST "http://192.168.68.7:8989/api/v1/split-pdf" \
|
||||
-H "X-API-Key: adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88" \
|
||||
-H "X-API-Key: «vault: infrastructure/production STIRLING_API_KEY»" \
|
||||
-F "fileInput=@/path/to/file.pdf" \
|
||||
-F "pageNumbers=1,2,3" \
|
||||
-o /tmp/output.zip
|
||||
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
{
|
||||
"status": "ok",
|
||||
"platform": "pi",
|
||||
"agent": "abiba",
|
||||
"zulip": {
|
||||
"connected": true,
|
||||
"site": "https://chat.sysloggh.net",
|
||||
"email": "abiba-bot@chat.sysloggh.net",
|
||||
"queue_id": "ee7f8b6d-9d53-48a7-ad58-f6e999771001",
|
||||
"bot_user_id": 21,
|
||||
"messages_processed": 0,
|
||||
"skipped": 0,
|
||||
"last_error": null
|
||||
},
|
||||
"circuit_breaker": {
|
||||
"state": "CLOSED",
|
||||
"failures": 0,
|
||||
"successes": 5,
|
||||
"totalRequests": 5,
|
||||
"failureRate": "0.000",
|
||||
"openedAt": null
|
||||
},
|
||||
"workers": [],
|
||||
"worker_count": 0
|
||||
}
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
{
|
||||
"status": "down",
|
||||
"platform": "pi",
|
||||
"agent": "abiba",
|
||||
"zulip": {
|
||||
"connected": false,
|
||||
"site": "https://chat.sysloggh.net",
|
||||
"email": "abiba-bot@chat.sysloggh.net",
|
||||
"queue_id": null,
|
||||
"bot_user_id": null,
|
||||
"messages_processed": 0,
|
||||
"skipped": 0,
|
||||
"last_error": "Zulip API error 401: queue registration failed"
|
||||
},
|
||||
"circuit_breaker": {
|
||||
"state": "CLOSED",
|
||||
"failures": 0,
|
||||
"successes": 0,
|
||||
"totalRequests": 0,
|
||||
"failureRate": "0.000",
|
||||
"openedAt": null
|
||||
},
|
||||
"workers": [],
|
||||
"worker_count": 0
|
||||
}
|
||||
@@ -0,0 +1,261 @@
|
||||
"""Regression tests for the 2026-09-12 retired-alias sweep in audit-hermes-config.py.
|
||||
|
||||
WHY THIS FILE EXISTS: the executable audit pushed agent configs toward a DEAD alias.
|
||||
Rule 8 required `auxiliary.vision.model == "gpu-light"` and
|
||||
`auxiliary.web_extract.model == "gpu-light"`, but `gpu-light` (and its raw predecessor
|
||||
`gemma-4-12b`) were retired on 2026-09-12 and now return 400 `Invalid model name`; the
|
||||
live RTX 5070 alias is `gpu-vision`. A config that adopted the correct canonical alias
|
||||
therefore FAILED our own audit, so the audit was actively enforcing a broken config.
|
||||
|
||||
These tests execute the real CLI (`python3 audit-hermes-config.py <config>`) and assert
|
||||
observable behaviour — exit code and the emitted rule message — for the live alias and
|
||||
for both retired names. No network, vault, or SSH access is required.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parent.parent
|
||||
AUDIT = ROOT / "audit-hermes-config.py"
|
||||
|
||||
BASE = """
|
||||
model:
|
||||
api_key: ""
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
max_tokens: 4096
|
||||
default: syslog-auto
|
||||
provider: harness
|
||||
fallback_providers:
|
||||
provider: deepseek
|
||||
model: deepseek-v4-flash
|
||||
api_key_env: DEEPSEEK_API_KEY
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
threshold: 0.65
|
||||
max_context_window: 131072
|
||||
auxiliary:
|
||||
vision:
|
||||
model: {alias}
|
||||
provider: harness
|
||||
web_extract:
|
||||
model: {alias}
|
||||
provider: harness
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
delegation:
|
||||
provider: harness
|
||||
custom_providers:
|
||||
- name: harness
|
||||
key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
"""
|
||||
|
||||
|
||||
def _run_config(tmp_path, name, text):
|
||||
cfg = tmp_path / name
|
||||
cfg.write_text(text)
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(AUDIT), str(cfg)],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
return proc.returncode, proc.stdout
|
||||
|
||||
|
||||
def _run(tmp_path, alias):
|
||||
return _run_config(tmp_path, f"{alias}.yaml", BASE.format(alias=alias))
|
||||
|
||||
|
||||
def test_live_canonical_alias_passes(tmp_path):
|
||||
"""The RTX 5070 alias that actually resolves must satisfy Rule 8."""
|
||||
code, out = _run(tmp_path, "gpu-vision")
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_retired_gpu_light_is_rejected(tmp_path):
|
||||
"""A config pinned to the retired alias must fail, not pass."""
|
||||
code, out = _run(tmp_path, "gpu-light")
|
||||
assert code == 1, out
|
||||
assert "auxiliary.vision.model must be gpu-vision" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_gemma_is_rejected(tmp_path):
|
||||
"""The retired raw model name must fail Rule 8 as well."""
|
||||
code, out = _run(tmp_path, "gemma-4-12b")
|
||||
assert code == 1, out
|
||||
assert "auxiliary.vision.model must be gpu-vision" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_corrected_compression_example_passes(tmp_path):
|
||||
"""The corrected workaround (vision=gpu-vision, compression=syslog-auto) must PASS."""
|
||||
code, out = _run(tmp_path, "gpu-vision")
|
||||
assert code == 0, out
|
||||
assert "[Rule 7] compression.model must be syslog-auto (got 'syslog-auto')" in out
|
||||
assert "[Rule 7] auxiliary.compression.model must be syslog-auto (got 'syslog-auto')" in out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_delegation_is_rejected(tmp_path):
|
||||
"""delegation.model has no dedicated value rule, so a retired name there used to PASS."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"delegation-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
"delegation:\n provider: harness",
|
||||
"delegation:\n provider: harness\n model: gpu-light",
|
||||
),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "delegation.model = 'gpu-light' is retired" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_custom_providers_is_rejected(tmp_path):
|
||||
"""custom_providers[*].model is model-bearing; a retired name there must fail."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"custom-provider-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
" - name: harness\n key_env: LITELLM_API_KEY",
|
||||
" - name: harness\n model: gpu-light\n key_env: LITELLM_API_KEY",
|
||||
),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "custom_providers[0].model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_raw_name_fails(tmp_path):
|
||||
"""Retired raw names no longer resolve (400), so they fail; the 2026-09-12 registry change moved qwen3.6-27B-code from raw-but-live to non-resolving."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"retired-qwen.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
"delegation:\n provider: harness",
|
||||
"delegation:\n provider: harness\n model: qwen3.6-27B-code",
|
||||
),
|
||||
)
|
||||
assert code != 0, out
|
||||
assert "delegation.model = 'qwen3.6-27B-code' is retired and no longer resolves" in out
|
||||
assert "use gpu-dense" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_fallback_providers_is_rejected(tmp_path):
|
||||
"""fallback_providers.model is model-bearing; a retired name there must fail."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"fallback-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(" model: deepseek-v4-flash", " model: gpu-light"),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "fallback_providers.model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_x_search_is_rejected(tmp_path):
|
||||
"""x_search.model was previously not enumerated; the derivation must catch it."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"x-search-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
"delegation:\n provider: harness",
|
||||
"delegation:\n provider: harness\nx_search:\n model: gpu-light",
|
||||
),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "x_search.model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
|
||||
"""A nested auxiliary sub-block outside the named three must still be derived."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"nested-aux-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
" compression:\n model: syslog-auto\n provider: harness\ndelegation:",
|
||||
" compression:\n model: syslog-auto\n provider: harness\n"
|
||||
" tasks:\n summarize:\n model: gpu-light\ndelegation:",
|
||||
),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_canonical_internal_path_passes(tmp_path):
|
||||
"""Rule 5 must accept the canonical internal base from hermes-key-enforcement.prose.md."""
|
||||
code, out = _run(
|
||||
tmp_path,
|
||||
"gpu-vision",
|
||||
)
|
||||
# Override the base_url in the config
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "canonical-internal.yaml", cfg_text)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
# Verify the correct message is shown
|
||||
assert "model.base_url is canonical" in out
|
||||
|
||||
|
||||
def test_wrong_base_url_fails(tmp_path):
|
||||
"""Rule 5 must reject paths outside the allowed list."""
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
" base_url: http://192.168.68.116/litellm/v1/responses",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "wrong-base.yaml", cfg_text)
|
||||
assert code == 1, out
|
||||
assert "RESULT: FAIL" in out
|
||||
assert "model.base_url must be one of" in out
|
||||
|
||||
|
||||
def test_public_host_path_passes(tmp_path):
|
||||
"""Rule 5 must accept the public host base."""
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
" base_url: https://litellm.sysloggh.net/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "public-host.yaml", cfg_text)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_old_rule5_check_would_fail_canonical(tmp_path):
|
||||
"""
|
||||
Proof that the OLD Rule 5 check would fail the canonical internal path.
|
||||
This proves the bug existed before the fix.
|
||||
"""
|
||||
# OLD check expected /v1, so the canonical /litellm/v1 would have failed
|
||||
canonical_cfg = BASE.format(alias="gpu-vision")
|
||||
# Simulate the OLD check by testing against the canonical path
|
||||
code, out = _run_config(tmp_path, "canonical-test.yaml", canonical_cfg)
|
||||
# NEW check: canonical /litellm/v1 SHOULD pass
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
# OLD check expected /v1, so the internal /v1 would have passed
|
||||
# NEW check: internal /v1 is non-canonical but working (WARN not FAIL)
|
||||
old_cfg = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "old-check-test.yaml", old_cfg)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
# NEW check should pass
|
||||
new_cfg = BASE.format(alias="gpu-vision")
|
||||
code, out = _run_config(tmp_path, "new-check-test.yaml", new_cfg)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
@@ -0,0 +1,138 @@
|
||||
"""
|
||||
Regression tests for daily-infra-report.py fixes (PR #64).
|
||||
|
||||
Tests:
|
||||
(a) Asserts the nested zulip read feeds the agent-card fields
|
||||
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
|
||||
"""
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch, MagicMock
|
||||
|
||||
# Add scripts to path
|
||||
sys.path.insert(0, str(Path(__file__).parent.parent / "scripts"))
|
||||
import importlib.util
|
||||
|
||||
def load_script():
|
||||
"""Load the daily-infra-report script as a module."""
|
||||
script_path = Path(__file__).parent.parent / "scripts" / "daily-infra-report.py"
|
||||
spec = importlib.util.spec_from_file_location("daily_infra_report", script_path)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
def test_nested_zulip_read_feeds_agent_card():
|
||||
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
|
||||
# Mock the http_get_body response with nested structure
|
||||
mock_health_response = json.dumps({
|
||||
"status": "ok",
|
||||
"platform": "pi",
|
||||
"agent": "abiba",
|
||||
"zulip": {
|
||||
"connected": True,
|
||||
"queue_id": "test-queue-id",
|
||||
"messages_processed": 42,
|
||||
"skipped": 5,
|
||||
"last_error": None
|
||||
}
|
||||
})
|
||||
|
||||
# Import and patch
|
||||
report_mod = load_script()
|
||||
|
||||
with patch.object(report_mod, 'http_get_body', return_value=mock_health_response):
|
||||
# Simulate the collect() function's Zulip section
|
||||
zulip_health = json.loads(report_mod.http_get_body("http://localhost:9200/health"))
|
||||
zulip_state = zulip_health.get("zulip", {})
|
||||
|
||||
# Assert the nested key is read correctly
|
||||
assert zulip_state.get("connected") == True, "Zulip connected should be True from nested key"
|
||||
assert zulip_state.get("messages_processed") == 42, "messages_processed should be 42 from nested key"
|
||||
assert zulip_state.get("queue_id") == "test-queue-id", "queue_id should be read from nested key"
|
||||
|
||||
# Simulate the agent card field population
|
||||
agent_card = {
|
||||
"zulip_connected": zulip_state.get("connected", False),
|
||||
"zulip_processed": zulip_state.get("messages_processed", 0),
|
||||
}
|
||||
|
||||
assert agent_card["zulip_connected"] == True, "Agent card should show Zulip connected"
|
||||
assert agent_card["zulip_processed"] == 42, "Agent card should show 42 processed messages"
|
||||
|
||||
|
||||
def test_unreachable_pve_get_renders_labelled_unreachable():
|
||||
"""Test that an unreachable PVE API renders 'unreachable' instead of '0/0'."""
|
||||
# Import and patch
|
||||
report_mod = load_script()
|
||||
|
||||
# Test pve_get returns None on error
|
||||
with patch.object(report_mod.subprocess, 'run') as mock_run:
|
||||
mock_run.return_value.returncode = 7 # Connection failure
|
||||
result = report_mod.pve_get("/api2/json/nodes")
|
||||
assert result is None, "pve_get should return None on connection failure"
|
||||
|
||||
# Test the render logic
|
||||
report = {
|
||||
"nodes": {},
|
||||
"node_count": 0,
|
||||
"nodes_online": 0,
|
||||
"pve_probe_status": "unreachable",
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
}
|
||||
|
||||
# The render should show "unreachable" not "0/0"
|
||||
pve_status_label = "unreachable" if report.get('pve_probe_status') == 'unreachable' else f"{report['nodes_online']}/{report['node_count']}"
|
||||
|
||||
assert pve_status_label == "unreachable", "PVE status should show 'unreachable' when probe fails, not '0/0'"
|
||||
|
||||
|
||||
def test_unreachable_resources_renders_labelled_unreachable():
|
||||
"""Test that unreachable resources probe renders 'unreachable' instead of '0/0'."""
|
||||
report_mod = load_script()
|
||||
|
||||
# Test resources probe returns None
|
||||
with patch.object(report_mod.subprocess, 'run') as mock_run:
|
||||
mock_run.return_value.returncode = 7
|
||||
result = report_mod.pve_get("/api2/json/cluster/resources")
|
||||
assert result is None, "pve_get for resources should return None on connection failure"
|
||||
|
||||
# Test the render logic
|
||||
report = {
|
||||
"resources_probe_status": "unreachable",
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
}
|
||||
|
||||
resources_label = "unreachable" if report.get('resources_probe_status') == 'unreachable' else f"{report['running_vms']}/{report['total_vms']}"
|
||||
|
||||
assert resources_label == "unreachable", "Resources status should show 'unreachable' when probe fails, not '0/0'"
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print("Running tests...")
|
||||
try:
|
||||
test_nested_zulip_read_feeds_agent_card()
|
||||
print("✓ test_nested_zulip_read_feeds_agent_card passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_nested_zulip_read_feeds_agent_card failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_unreachable_pve_get_renders_labelled_unreachable()
|
||||
print("✓ test_unreachable_pve_get_renders_labelled_unreachable passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_unreachable_pve_get_renders_labelled_unreachable failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_unreachable_resources_renders_labelled_unreachable()
|
||||
print("✓ test_unreachable_resources_renders_labelled_unreachable passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
print("All tests passed!")
|
||||
@@ -0,0 +1,185 @@
|
||||
"""Regression tests for the disk-gc report-only gate (CT 111 / tdunna / .129).
|
||||
|
||||
WHY THIS FILE EXISTS: `disk-gc-threat-response.prose.md` defined AMBER as "GC scheduled
|
||||
for next run" and its Execution loop called `gc-executor` for EVERY threat, with no
|
||||
guest-level exclusion. CT 111 (tdunna, 192.168.68.129) belongs to Theo and is
|
||||
report-only per the captain (2026-08-17, re-confirmed 2026-09-10) — so a single AMBER
|
||||
reading on that guest would have scheduled GC commands (apt clean, journal vacuum,
|
||||
log/tmp deletion, snap removal) against someone else's box. The only marker was
|
||||
frontmatter `report_only_agents`, which names an AGENT while the scan unit is a GUEST.
|
||||
|
||||
These tests execute the real planner (`scripts/disk-gc-plan.py`) and assert observable
|
||||
behaviour: an excluded guest never produces a `gc-executor` action at any level, while
|
||||
our own guests still do.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parent.parent
|
||||
PLAN = ROOT / "scripts" / "disk-gc-plan.py"
|
||||
|
||||
|
||||
def _plan(scan, tmp_path):
|
||||
scan_file = tmp_path / "scan.json"
|
||||
scan_file.write_text(json.dumps(scan))
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(PLAN), "--scan", str(scan_file), "--json"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
return json.loads(proc.stdout)
|
||||
|
||||
|
||||
def _actions_for(plan, target):
|
||||
return [row for row in plan if str(row["target"]) == str(target)]
|
||||
|
||||
|
||||
def test_excluded_guest_never_gets_gc_at_any_level(tmp_path):
|
||||
"""CT 111 at AMBER, RED and CRITICAL — always report-only, never gc-executor."""
|
||||
for pct, level in ((84, "AMBER"), (90, "RED"), (97, "CRITICAL")):
|
||||
plan = _plan([{"id": 111, "hostname": "tdunna", "ip": "192.168.68.129",
|
||||
"usage_pct": pct}], tmp_path)
|
||||
rows = _actions_for(plan, 111)
|
||||
assert rows, f"CT 111 must still be reported at {level}"
|
||||
assert rows[0]["level"] == level
|
||||
assert rows[0]["action"] == "report-only", rows
|
||||
assert not any(r["action"] == "gc-executor" for r in rows)
|
||||
|
||||
|
||||
def test_exclusion_matches_on_any_identity_key(tmp_path):
|
||||
"""The gate is keyed on guest/host, so id, ct, ctid, hostname or IP all match."""
|
||||
for entry in ({"id": 111, "usage_pct": 95},
|
||||
{"ct": 111, "usage_pct": 95},
|
||||
{"ctid": 111, "usage_pct": 95},
|
||||
{"hostname": "tdunna", "usage_pct": 95},
|
||||
{"ip": "192.168.68.129", "usage_pct": 95}):
|
||||
plan = _plan([entry], tmp_path)
|
||||
assert all(r["action"] == "report-only" for r in plan), (entry, plan)
|
||||
|
||||
|
||||
def test_exclusion_matches_encoded_identities(tmp_path):
|
||||
"""A differently-encoded CT 111 id must not slip past the gate to gc-executor."""
|
||||
encodings = ({"id": 111.0},
|
||||
{"id": "111.0"},
|
||||
{"id": "lxc/111"},
|
||||
{"id": "qemu/111"},
|
||||
{"id": "0111"},
|
||||
{"id": " 111 "})
|
||||
for alias in encodings:
|
||||
plan = _plan([{**alias, "usage_pct": 97}], tmp_path)
|
||||
assert plan, alias
|
||||
assert plan[0]["action"] == "report-only", (alias, plan)
|
||||
|
||||
|
||||
def test_our_own_guests_still_get_gc(tmp_path):
|
||||
"""acerpve .9 and amdpve .15 are ours — they must still be acted on."""
|
||||
plan = _plan([{"hostname": "acerpve", "ip": "192.168.68.9", "usage_pct": 77},
|
||||
{"hostname": "amdpve", "ip": "192.168.68.15", "usage_pct": 76}], tmp_path)
|
||||
assert len(plan) == 2
|
||||
assert all(r["action"] == "gc-executor" for r in plan), plan
|
||||
|
||||
|
||||
def test_below_threshold_emits_nothing(tmp_path):
|
||||
"""GREEN guests produce no action at all."""
|
||||
assert _plan([{"id": 111, "usage_pct": 40}], tmp_path) == []
|
||||
|
||||
|
||||
def test_agent_name_alone_does_not_gate_a_guest(tmp_path):
|
||||
"""An agent-name marker must not be the gate: an unrelated guest still gets GC."""
|
||||
plan = _plan([{"id": 999, "hostname": "koby", "usage_pct": 95}], tmp_path)
|
||||
assert plan and plan[0]["action"] == "gc-executor"
|
||||
|
||||
|
||||
def test_unidentified_threat_fails_closed(tmp_path):
|
||||
"""A threshold-crossing entry with no recognized identity must not schedule GC."""
|
||||
plan = _plan([{"usage_pct": 97}], tmp_path)
|
||||
assert plan, "an unidentified threat must still be reported"
|
||||
assert plan[0]["action"] == "report-only", plan
|
||||
assert plan[0]["reason"], plan
|
||||
|
||||
|
||||
def test_exclusion_entry_without_identity_fails_closed(tmp_path):
|
||||
"""A mis-typed exclusion entry must break the run, never silently disable the gate."""
|
||||
contract = tmp_path / "broken.prose.md"
|
||||
contract.write_text(
|
||||
"```yaml\n"
|
||||
"report_only_guests:\n"
|
||||
" - node: storepve\n"
|
||||
" reason: \"typo - no guest identity\"\n"
|
||||
"```\n"
|
||||
)
|
||||
scan_file = tmp_path / "scan.json"
|
||||
scan_file.write_text(json.dumps([{"ct": 111, "usage_pct": 97}]))
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(PLAN), "--scan", str(scan_file),
|
||||
"--contract", str(contract), "--json"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
assert proc.returncode != 0, proc.stdout
|
||||
assert "gc-executor" not in proc.stdout
|
||||
assert "identity" in proc.stderr.lower(), proc.stderr
|
||||
|
||||
|
||||
def test_empty_exclusion_block_fails_closed(tmp_path):
|
||||
"""An emptied report_only_guests list must break the run, not disable the gate."""
|
||||
contract = tmp_path / "empty.prose.md"
|
||||
contract.write_text("```yaml\nreport_only_guests: []\n```\n")
|
||||
scan_file = tmp_path / "scan.json"
|
||||
scan_file.write_text(json.dumps([{"ct": 111, "usage_pct": 97}]))
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(PLAN), "--scan", str(scan_file),
|
||||
"--contract", str(contract), "--json"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
assert proc.returncode != 0, proc.stdout
|
||||
assert "gc-executor" not in proc.stdout
|
||||
assert "report-only gate" in proc.stderr.lower(), proc.stderr
|
||||
|
||||
|
||||
def _plan_with_contract(scan, contract_text, tmp_path, name):
|
||||
contract = tmp_path / name
|
||||
contract.write_text(contract_text)
|
||||
scan_file = tmp_path / f"scan-{name}.json"
|
||||
scan_file.write_text(json.dumps(scan))
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(PLAN), "--scan", str(scan_file),
|
||||
"--contract", str(contract), "--json"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
return json.loads(proc.stdout)
|
||||
|
||||
|
||||
def test_gate_is_read_from_the_contract_block(tmp_path):
|
||||
"""The gate is data-driven by the contract block: the planner excludes the guest
|
||||
when the block names it and acts on it when the block does not. Executes the real
|
||||
planner interface against both fixtures so the behaviour change is observable."""
|
||||
scan = [{"id": 111, "hostname": "tdunna", "ip": "192.168.68.129", "usage_pct": 95}]
|
||||
with_gate = (
|
||||
"```yaml\n"
|
||||
"report_only_guests:\n"
|
||||
" - guest: 111\n"
|
||||
" hostname: tdunna\n"
|
||||
" ip: 192.168.68.129\n"
|
||||
" reason: \"fixture reason\"\n"
|
||||
"```\n"
|
||||
)
|
||||
without_gate = (
|
||||
"```yaml\n"
|
||||
"report_only_guests:\n"
|
||||
" - guest: 999\n"
|
||||
" reason: \"fixture excludes a different guest\"\n"
|
||||
"```\n"
|
||||
)
|
||||
|
||||
gated = _plan_with_contract(scan, with_gate, tmp_path, "gated.prose.md")
|
||||
assert gated[0]["action"] == "report-only", gated
|
||||
assert gated[0]["reason"] == "fixture reason", gated
|
||||
|
||||
ungated = _plan_with_contract(scan, without_gate, tmp_path, "ungated.prose.md")
|
||||
assert ungated[0]["action"] == "gc-executor", ungated
|
||||
assert ungated[0]["action"] != gated[0]["action"]
|
||||
@@ -0,0 +1,381 @@
|
||||
"""Regression tests for the 2026-09-10 retirement of the Mumuni monitoring leg.
|
||||
|
||||
WHY THIS FILE EXISTS: captain ruling 2026-09-10 — Mumuni moved off this host
|
||||
onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated
|
||||
`hermes` user) and is monitored from her side. The monitor nevertheless kept
|
||||
ssh'ing to root@192.168.68.24 for `~/.hermes/gateway_state.json` on the
|
||||
decommissioned deployment, read "unknown" on every run, and posted a false 🔴
|
||||
"Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert to the
|
||||
captain. The daily infra digest published a matching `mumuni:unknown` row.
|
||||
|
||||
CONTRACT UNDER TEST:
|
||||
* `scripts/zulip-monitor.sh` carries NO Mumuni probe and NO 192.168.68.24
|
||||
reference; it never ssh'es .24, and even on a failing run it emits no Mumuni
|
||||
notify (stdout alert, Zulip payload, or log line).
|
||||
* The Abiba (pi — the Zulip bridge), Tanko (DSH) and Agent Zero (kagentz) legs
|
||||
still work: deleting the Mumuni leg must not have gutted the rest.
|
||||
* `scripts/daily-infra-report.py` no longer probes .24 for a Hermes gateway
|
||||
state and no longer emits a `mumuni` agent entry.
|
||||
* `scripts/agent-health-check.py`'s AGENTS roster has no mumuni entry. This is
|
||||
a pin, not a behavior change — verify the probe was already gone.
|
||||
* `zulip-health.prose.md` retires the Mumuni-only steps and says explicitly
|
||||
that Mumuni is not monitored from this host.
|
||||
|
||||
HOW: behavioral execution plus one named deliverable-text contract. The sandbox
|
||||
copies the shipped monitor verbatim and rewrites only its LOG constant, then
|
||||
runs it with stub ssh/curl on PATH; the ssh stub records every host it is asked
|
||||
to reach, so "never probes .24" and "no Mumuni notify" are asserted from
|
||||
observed behavior. The daily digest is pinned by importing it and exercising
|
||||
collect() and build_html() directly. The single source-text assertion is the
|
||||
deliverable-text contract the captain acceptance names for the shipped monitor.
|
||||
|
||||
Usage: python3 -m pytest tests/test_mumuni_monitor_removal.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
import os
|
||||
import pathlib
|
||||
import stat
|
||||
import subprocess
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||
DAILY_REPORT = ROOT / "scripts" / "daily-infra-report.py"
|
||||
AHC = ROOT / "scripts" / "agent-health-check.py"
|
||||
HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
|
||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||
|
||||
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
|
||||
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||
|
||||
|
||||
# ── scripts/zulip-monitor.sh: deliverable-text contract ─────────────
|
||||
|
||||
def test_zulip_monitor_deliverable_text_contract():
|
||||
"""Owned deliverable-text contract for scripts/zulip-monitor.sh.
|
||||
|
||||
Captain acceptance requires the shipped monitor to contain no Mumuni probe
|
||||
identifier and no 192.168.68.24 literal. Behavioral proof that the monitor
|
||||
never contacts that host and never emits a Mumuni notify lives in the
|
||||
sandbox tests below; this only pins the named text contract.
|
||||
"""
|
||||
text = ZULIP_MONITOR.read_text()
|
||||
assert "mumuni" not in text.lower()
|
||||
assert MUMUNI_IP not in text
|
||||
|
||||
|
||||
# ── scripts/zulip-monitor.sh: behavioral sandbox ─────────────────────
|
||||
|
||||
SSH_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub ssh: record the target host, then answer by host + remote command.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||
host=""
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
*@192.168.*) host="${a##*@}" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||
cmd="${*: -1}"
|
||||
case "$host" in
|
||||
192.168.68.15)
|
||||
case "$cmd" in
|
||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||
esac ;;
|
||||
192.168.68.14)
|
||||
case "$cmd" in
|
||||
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "$AZ_A2A_EXIT" ;;
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||
# kagentz C3 public URL, and record every call (including notify) payloads.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
case " $* " in
|
||||
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
*kagentz.sysloggh.net*)
|
||||
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
path.write_text(body)
|
||||
path.chmod(path.stat().st_mode
|
||||
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a_code="401", az_a2a_exit=0,
|
||||
kagentz_public_code="302"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
Everything else — legs, labels, notify logic — is the shipped script.
|
||||
"""
|
||||
sandbox = tmp_path / "sandbox"
|
||||
bindir = sandbox / "bin"
|
||||
record = sandbox / "record"
|
||||
bindir.mkdir(parents=True)
|
||||
record.mkdir()
|
||||
|
||||
_write_exec(bindir / "ssh", SSH_STUB)
|
||||
_write_exec(bindir / "curl", CURL_STUB)
|
||||
|
||||
source = ZULIP_MONITOR.read_text()
|
||||
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||
log_path = sandbox / "zulip-health-monitor.log"
|
||||
script = sandbox / "zulip-monitor.sh"
|
||||
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||
|
||||
env = dict(os.environ)
|
||||
env.update({
|
||||
"PATH": f"{bindir}:{env['PATH']}",
|
||||
"RECORD_DIR": str(record),
|
||||
"TANKO_SVC": tanko_svc,
|
||||
"TANKO_HTTP": tanko_http,
|
||||
"AZ_A2A_CODE": az_a2a_code,
|
||||
"AZ_A2A_EXIT": str(az_a2a_exit),
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
return proc, record, log_path
|
||||
|
||||
|
||||
def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
||||
proc, record, log_path = _run_monitor(tmp_path)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# Every retained leg actually ran and passed.
|
||||
assert "Server: ✅ HTTP 200" in log
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "Tanko: ✅ service=active http=200" in log
|
||||
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||
|
||||
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||
assert proc.stdout == ""
|
||||
assert "Mumuni" not in log
|
||||
assert "🔴" not in log
|
||||
|
||||
# Observed behavior: .24 is never resolved, only Tanko's vantage and the
|
||||
# Agent Zero host are contacted.
|
||||
hosts = record.joinpath("ssh.hosts").read_text().split()
|
||||
assert MUMUNI_IP not in hosts
|
||||
assert set(hosts) == {TANKO_VANTAGE, AGENT_ZERO_HOST}
|
||||
assert not record.joinpath("unexpected-ssh").exists()
|
||||
|
||||
|
||||
def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
||||
# Failure path: exercises notify() end to end so "no Mumuni notify" is
|
||||
# proven on the alert path, not only on the quiet healthy path.
|
||||
proc, record, log_path = _run_monitor(tmp_path, tanko_svc="inactive",
|
||||
tanko_http="000")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
|
||||
alerts = proc.stdout
|
||||
assert "Tanko (DSH dsh-web) service state: inactive" in alerts
|
||||
assert "1 issue(s) found" in alerts
|
||||
|
||||
# No Mumuni text in stdout, the log, or any Zulip DM/stream payload.
|
||||
assert "Mumuni" not in alerts
|
||||
assert "Mumuni" not in log_path.read_text()
|
||||
assert MUMUNI_IP not in alerts + log_path.read_text()
|
||||
payloads = record.joinpath("curl.calls").read_text()
|
||||
assert "Mumuni" not in payloads
|
||||
assert MUMUNI_IP not in payloads
|
||||
|
||||
# The rest of the monitor still ran alongside the failing Tanko leg.
|
||||
log = log_path.read_text()
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "kagentz C1: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
|
||||
|
||||
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||
proc, record, log_path = _run_monitor(tmp_path, az_a2a_code="500")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz C1: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||
assert "kagentz C1: ✅ A2A alive" not in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
assert "kagentz A2A server answered HTTP 500" in proc.stdout
|
||||
|
||||
|
||||
def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
|
||||
# curl prints the http_code before failing, so the ssh stub exits non-zero
|
||||
# with "000" on stdout — exercising the real outage path.
|
||||
proc, record, log_path = _run_monitor(tmp_path, az_a2a_code="000",
|
||||
az_a2a_exit=7)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||
assert "kagentz C1: ✅ A2A alive" not in log
|
||||
assert "unexpected" not in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
|
||||
|
||||
|
||||
# ── scripts/daily-infra-report.py: behavioral digest checks ──────────
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def daily():
|
||||
spec = importlib.util.spec_from_file_location("daily_infra_report", DAILY_REPORT)
|
||||
assert spec and spec.loader
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
DAILY_AGENTS = {
|
||||
"abiba": {
|
||||
"platform": "pi", "ct": 100, "ip": MUMUNI_IP,
|
||||
"zulip_connected": True, "zulip_processed": 5,
|
||||
"pm2_status": "online", "pm2_restarts": "0", "pm2_uptime": "1h",
|
||||
},
|
||||
"tanko": {
|
||||
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
|
||||
"gateway_state": "n/a (DSH)", "zulip_state": "connected",
|
||||
"telegram_state": "unknown", "gateway_pid": None, "updated_at": "",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def _fabricated_report(agents):
|
||||
return {
|
||||
"nodes": {},
|
||||
"node_count": 1,
|
||||
"nodes_online": 1,
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
"stopped_vms": [],
|
||||
"vms_by_node": {n: [] for n in
|
||||
["amdpve", "minipve", "storepve", "acerpve", "ocupve"]},
|
||||
"storage": [],
|
||||
"docker_vm": {"total": 0, "running": 0, "unhealthy": [],
|
||||
"containers": [], "reclaimable": "", "disk_used": "1%"},
|
||||
"docker_syslog": {"total": 0, "running": 0, "containers": []},
|
||||
"docker_netbird": {"total": 0, "running": 0, "containers": []},
|
||||
"endpoints": [],
|
||||
"litellm": {"checks": []},
|
||||
"nfs": [],
|
||||
"zulip_ext": {
|
||||
"connected": True, "queue_id": "queue", "last_error": None,
|
||||
"messages_processed": 0, "retry_count": 0, "pm2": {},
|
||||
"pm2_healthy": True, "bot_skipped_15min": 0, "finalized_1h": 0,
|
||||
"failed_finalize_1h": 0, "finalize_fail_pct": 0,
|
||||
"server_status": "200",
|
||||
},
|
||||
"agents": agents,
|
||||
}
|
||||
|
||||
|
||||
def _agent_status_card(html):
|
||||
start = html.index("🤖 Agent Status")
|
||||
end = html.index("💬 Zulip Extension")
|
||||
return html[start:end]
|
||||
|
||||
|
||||
def test_daily_report_renders_only_abiba_and_tanko_agents(daily):
|
||||
"""build_html() over a Mumuni-free agent set must render no Mumuni row and
|
||||
no Mumuni gateway-unknown issue, while abiba and tanko rows still render."""
|
||||
html = daily.build_html(_fabricated_report(dict(DAILY_AGENTS)))
|
||||
card = _agent_status_card(html)
|
||||
assert "mumuni" not in card.lower()
|
||||
assert "abiba" in card
|
||||
assert "tanko" in card
|
||||
assert "mumuni" not in html.lower()
|
||||
|
||||
|
||||
def test_daily_report_collect_never_probes_mumuni(monkeypatch, daily):
|
||||
"""collect() with ssh stubbed must add no mumuni agent and must never ssh
|
||||
its decommissioned .24 host."""
|
||||
probed = []
|
||||
|
||||
class _NoSubprocess:
|
||||
@staticmethod
|
||||
def check_output(*args, **kwargs):
|
||||
return b""
|
||||
|
||||
def fake_ssh(host, cmd):
|
||||
probed.append(host)
|
||||
return ""
|
||||
|
||||
monkeypatch.setattr(daily, "pve_get", lambda path: [])
|
||||
monkeypatch.setattr(daily, "ssh_jerome", lambda host, cmd: "")
|
||||
monkeypatch.setattr(daily, "ssh", fake_ssh)
|
||||
monkeypatch.setattr(daily, "http_get",
|
||||
lambda url, auth=None, timeout=10: "200")
|
||||
monkeypatch.setattr(daily, "http_get_body",
|
||||
lambda url, auth=None, timeout=10: "")
|
||||
monkeypatch.setattr(daily, "count_in_log", lambda *a, **k: 0)
|
||||
monkeypatch.setattr(daily, "subprocess", _NoSubprocess)
|
||||
|
||||
report = daily.collect()
|
||||
assert "mumuni" not in report["agents"]
|
||||
assert MUMUNI_IP not in probed
|
||||
|
||||
|
||||
# ── scripts/agent-health-check.py: roster pin ───────────────────────
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def ahc():
|
||||
spec = importlib.util.spec_from_file_location("agent_health_check_roster", AHC)
|
||||
assert spec and spec.loader
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
def test_agent_health_roster_has_no_mumuni_entry(ahc):
|
||||
assert "mumuni" not in ahc.AGENTS
|
||||
|
||||
|
||||
# ── zulip-health.prose.md: contract reconciliation ──────────────────
|
||||
|
||||
def test_health_contract_retires_mumuni_only_steps():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
assert MUMUNI_IP not in text
|
||||
for step in ("**B4: Gateway Process**", "**B5: Heartbeat Verification**",
|
||||
"**B6: Response Delivery**"):
|
||||
assert step not in text
|
||||
|
||||
|
||||
def test_health_contract_states_mumuni_is_not_monitored_from_this_host():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
assert "Mumuni is NOT monitored from this host" in text
|
||||
assert "monitored on her side" in text
|
||||
assert "her own container" in text
|
||||
|
||||
|
||||
def test_health_contract_keeps_tanko_agent_zero_and_bridge_steps():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
for marker in ("**B1:", "**B2:", "**B3:", "Step 4: Platform C",
|
||||
"Step 2: Platform A", "Step 1: Zulip Server Liveness"):
|
||||
assert marker in text, marker
|
||||
@@ -0,0 +1,466 @@
|
||||
"""Regression tests for the 2026-09-09/10 probe-drift corrections.
|
||||
|
||||
WHY THIS FILE EXISTS: the monitoring contracts kept emitting false alarms from
|
||||
stale expectations rather than live faults.
|
||||
|
||||
* agent-health-check v3 reported 6 failures that were all stale expectations:
|
||||
abiba (pi-only since the harness purge) was tested as a Hermes host, koby
|
||||
(report-only per the captain's 2026-08-17 ruling) was counted as repairable,
|
||||
koby's CT 111 was probed on amdpve where it does not exist (it runs on
|
||||
storepve .6), and the wrapper infisical check had two bugs — it read only
|
||||
the first 20 lines, so koonimo's wrapper (which references /usr/bin/infisical
|
||||
past line 20) false-failed, and it treated koby's genuine no-infisical
|
||||
(~/.hermes/.env) wrapper as broken.
|
||||
* gpu-monitor emitted "DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000" three
|
||||
times from probing bare port 80 on GPU hosts while :8080 answered 200.
|
||||
* infrastructure-monitoring probed CT 116 for the PVE API (no pveproxy ->
|
||||
000) instead of the five real cluster nodes, which answer 401 = alive.
|
||||
|
||||
These tests execute the health script (with SSH/vault stubbed) and the real
|
||||
provenance consumer (scripts/prose-lint.sh), and parse the contracts' executable
|
||||
check-health probe blocks into normalized probe sets. No live network, vault, or
|
||||
SSH access is required.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
import json
|
||||
import os
|
||||
import pathlib
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
import textwrap
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
AHC = ROOT / "scripts" / "agent-health-check.py"
|
||||
LINT = ROOT / "scripts" / "prose-lint.sh"
|
||||
GPU = ROOT / "gpu-monitor.prose.md"
|
||||
INFRA = ROOT / "infrastructure-monitoring.prose.md"
|
||||
|
||||
PVE_NODE_IPS = {
|
||||
"192.168.68.9",
|
||||
"192.168.68.5",
|
||||
"192.168.68.15",
|
||||
"192.168.68.6",
|
||||
"192.168.68.12",
|
||||
}
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def ahc():
|
||||
"""Import agent-health-check.py without live network/SSH side effects."""
|
||||
spec = importlib.util.spec_from_file_location("agent_health_check", AHC)
|
||||
assert spec and spec.loader
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
# ── helpers: execute the health script with SSH/vault stubbed ─────────
|
||||
|
||||
def _run_main(ahc, monkeypatch, capsys, argv, ssh_result=None):
|
||||
ahc.FAIL.clear()
|
||||
ahc.REPORT_ONLY.clear()
|
||||
monkeypatch.setattr(ahc, "load_agent_keys", lambda: None)
|
||||
monkeypatch.setattr(ahc, "ssh", lambda *a, **k: ssh_result)
|
||||
monkeypatch.setattr(sys, "argv", ["agent-health-check.py", "--no-deploy", *argv])
|
||||
with pytest.raises(SystemExit) as exc:
|
||||
ahc.main()
|
||||
return exc.value.code, capsys.readouterr().out
|
||||
|
||||
|
||||
def _json_payload(out):
|
||||
for line in reversed(out.splitlines()):
|
||||
if line.startswith('{"timestamp"'):
|
||||
return json.loads(line)
|
||||
raise AssertionError(f"no JSON payload in output:\n{out}")
|
||||
|
||||
|
||||
# ── agent-health-check: stale-expectation legs ───────────────────────
|
||||
|
||||
def test_import_does_not_contact_vault(ahc):
|
||||
# Keys are loaded in main() via load_agent_keys(); importing must stay inert.
|
||||
assert callable(ahc.load_agent_keys)
|
||||
assert all(agent.get("key") is None for agent in ahc.AGENTS.values())
|
||||
|
||||
|
||||
def test_abiba_is_pi_only_runtime(ahc):
|
||||
# .24 has run pi-only since the harness purge: no Hermes gateway, config, or
|
||||
# wrapper. Probing those legs produced false failures.
|
||||
assert ahc.AGENTS["abiba"]["runtime"] == "pi"
|
||||
|
||||
|
||||
def test_koby_is_report_only(ahc):
|
||||
# Captain's 2026-08-17 ruling (Rule 17): detect and report, never repair.
|
||||
assert ahc.AGENTS["koby"]["report_only"] is True
|
||||
|
||||
|
||||
def test_koby_ct111_is_on_storepve(ahc):
|
||||
# Live-verified 2026-09-10: `pct status 111` = running on storepve (.6);
|
||||
# amdpve has no lxc/111.conf, which is what false-failed before.
|
||||
assert ahc.AGENTS["koby"]["pve"] == "storepve"
|
||||
|
||||
|
||||
def test_report_only_legs_never_count_as_failures(ahc):
|
||||
for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)):
|
||||
ahc.FAIL.clear()
|
||||
ahc.REPORT_ONLY.clear()
|
||||
ahc._fail(f"probe:{agent}", agent)
|
||||
if report_only:
|
||||
assert ahc.FAIL == []
|
||||
assert ahc.REPORT_ONLY == [f"probe:{agent}"]
|
||||
else:
|
||||
assert ahc.FAIL == [f"probe:{agent}"]
|
||||
assert ahc.REPORT_ONLY == []
|
||||
ahc.FAIL.clear()
|
||||
ahc.REPORT_ONLY.clear()
|
||||
|
||||
|
||||
def test_failure_recording_accepts_agentless_keys(ahc):
|
||||
ahc.FAIL.clear()
|
||||
try:
|
||||
ahc._fail("gpu-no-port:gpu-rtx3090 (.8)")
|
||||
assert ahc.FAIL == ["gpu-no-port:gpu-rtx3090 (.8)"]
|
||||
finally:
|
||||
ahc.FAIL.clear()
|
||||
|
||||
|
||||
def test_json_reports_absolute_execution_provenance(ahc, monkeypatch, capsys):
|
||||
code, out = _run_main(ahc, monkeypatch, capsys, ["--json"])
|
||||
payload = _json_payload(out)
|
||||
assert payload["execution_path"] == os.path.abspath(str(AHC))
|
||||
assert payload["cwd"] == os.getcwd()
|
||||
assert code == 1 # stubbed SSH fails every leg, but provenance is still emitted
|
||||
|
||||
|
||||
def test_quiet_run_still_carries_provenance_on_the_alert_path(ahc, monkeypatch, capsys):
|
||||
# The cron runs --quiet; a failure report must still carry provenance. The
|
||||
# header line is suppressed in quiet mode, so the ALERT line is the carrier.
|
||||
code, out = _run_main(ahc, monkeypatch, capsys, ["--quiet"])
|
||||
assert code == 1
|
||||
assert "📍 executed from:" not in out
|
||||
alerts = [ln for ln in out.splitlines() if ln.startswith("ALERT agent-health:")]
|
||||
assert alerts, out
|
||||
assert f"script={os.path.abspath(str(AHC))}" in alerts[0]
|
||||
assert f"cwd={os.getcwd()}" in alerts[0]
|
||||
|
||||
|
||||
def test_quiet_healthy_run_emits_no_stdout(ahc, monkeypatch, capsys):
|
||||
# --quiet is documented as "only output on failure": a run with no fleet
|
||||
# failures must produce no stdout at all (the production cron runs --quiet).
|
||||
for name in ("check_keys", "check_gpu_ports", "check_agents", "check_ct_liveness",
|
||||
"check_config_integrity", "check_wrapper_integrity", "check_vault_secrets"):
|
||||
monkeypatch.setattr(ahc, name, lambda: None)
|
||||
code, out = _run_main(ahc, monkeypatch, capsys, ["--quiet"])
|
||||
assert code == 0
|
||||
assert out == ""
|
||||
|
||||
|
||||
def test_json_surfaces_report_only_findings_separately(ahc, monkeypatch, capsys):
|
||||
# Koby's down legs are reported but must not count as fleet failures; the
|
||||
# --json payload exposes them in their own array (item 1 + f8).
|
||||
_, out = _run_main(ahc, monkeypatch, capsys, ["--json"])
|
||||
payload = _json_payload(out)
|
||||
assert isinstance(payload["report_only"], list)
|
||||
assert any(key.startswith(("gateway-down:koby", "ct-unreachable:koby"))
|
||||
for key in payload["report_only"])
|
||||
assert not any("koby" in key for key in payload["failures"])
|
||||
|
||||
|
||||
# ── agent-health-check: wrapper infisical behavior (f3) ───────────────
|
||||
|
||||
def _stub_wrapper_ssh(ahc, monkeypatch, wrapper_body, test_x_result="OK", command_v="/usr/local/bin/infisical"):
|
||||
def fake_ssh(host, cmd, user="root"):
|
||||
if cmd.startswith("cat /root/.local/bin/hermes"):
|
||||
return wrapper_body
|
||||
if cmd.startswith("ls -la /root/.local/bin/hermes "):
|
||||
return "-rwxr-xr-x 1 root root 0 Jan 1 00:00 /root/.local/bin/hermes"
|
||||
if cmd.startswith("ls -la /root/.local/bin/hermes-real") or "venv/bin/hermes" in cmd:
|
||||
return "-rwxr-xr-x 1 root root 0 Jan 1 00:00 /root/.local/bin/hermes-real"
|
||||
if cmd.startswith("grep -c 'LITELLM_API_KEY'"):
|
||||
return "1"
|
||||
if cmd.startswith("test -x "):
|
||||
path = cmd[len("test -x "):].split()[0]
|
||||
if isinstance(test_x_result, dict):
|
||||
return test_x_result.get(path, "MISS")
|
||||
return test_x_result
|
||||
if cmd.startswith("command -v infisical"):
|
||||
return command_v
|
||||
return None
|
||||
|
||||
monkeypatch.setattr(ahc, "ssh", fake_ssh)
|
||||
monkeypatch.setattr(ahc, "AGENTS", {"koonimo": dict(ahc.AGENTS["koonimo"])})
|
||||
ahc.FAIL.clear()
|
||||
ahc.REPORT_ONLY.clear()
|
||||
|
||||
|
||||
def test_env_based_wrapper_without_infisical_is_not_failed(ahc, monkeypatch, capsys):
|
||||
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||
"#!/bin/bash\nsource ~/.hermes/.env\nexec hermes-real \"$@\"\n")
|
||||
ahc.check_wrapper_integrity()
|
||||
out = capsys.readouterr().out
|
||||
assert ahc.FAIL == []
|
||||
assert "wrapper resolves creds without infisical" in out
|
||||
|
||||
|
||||
def test_dangling_absolute_infisical_path_is_failed(ahc, monkeypatch, capsys):
|
||||
# Wrapper hardcodes /usr/bin/infisical, which is absent, while PATH resolves
|
||||
# infisical to /usr/local/bin/infisical. The literal path must be verified,
|
||||
# not inferred from PATH resolution.
|
||||
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||
"#!/bin/bash\n/usr/bin/infisical run -- hermes-real \"$@\"\n",
|
||||
test_x_result="MISS", command_v="/usr/local/bin/infisical")
|
||||
ahc.check_wrapper_integrity()
|
||||
assert "wrapper-infisical-path:koonimo" in ahc.FAIL
|
||||
|
||||
|
||||
def test_existing_absolute_infisical_path_passes(ahc, monkeypatch, capsys):
|
||||
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||
"#!/bin/bash\n/usr/bin/infisical run -- hermes-real \"$@\"\n",
|
||||
test_x_result="OK")
|
||||
ahc.check_wrapper_integrity()
|
||||
out = capsys.readouterr().out
|
||||
assert ahc.FAIL == []
|
||||
assert "wrapper infisical path OK" in out
|
||||
|
||||
|
||||
def test_comment_mentioning_removed_infisical_path_is_not_failed(ahc, monkeypatch, capsys):
|
||||
# litellm-api-keys.prose.md documents `rm -f /usr/local/bin/infisical`; a
|
||||
# wrapper comment about that migration must not manufacture a dangling path
|
||||
# when the real invocation (/usr/bin/infisical) is present and executable.
|
||||
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||
"#!/bin/bash\n# migrated from /usr/local/bin/infisical\n"
|
||||
"exec /usr/bin/infisical run -- hermes-real \"$@\"\n",
|
||||
test_x_result={"/usr/bin/infisical": "OK",
|
||||
"/usr/local/bin/infisical": "MISS"})
|
||||
ahc.check_wrapper_integrity()
|
||||
out = capsys.readouterr().out
|
||||
assert ahc.FAIL == []
|
||||
assert "wrapper infisical path OK" in out
|
||||
|
||||
|
||||
def test_comment_only_infisical_mention_does_not_reach_path_check(ahc, monkeypatch, capsys):
|
||||
# A comment-only mention of a removed infisical path on a healthy .env-based
|
||||
# wrapper is not an invocation: it must not fall through to the `command -v`
|
||||
# PATH check and false-FAIL `wrapper-no-infisical`.
|
||||
_stub_wrapper_ssh(ahc, monkeypatch,
|
||||
"#!/bin/bash\n# migrated from /usr/local/bin/infisical\n"
|
||||
"source ~/.hermes/.env\nexec hermes-real \"$@\"\n",
|
||||
test_x_result="MISS", command_v=None)
|
||||
ahc.check_wrapper_integrity()
|
||||
out = capsys.readouterr().out
|
||||
assert ahc.FAIL == []
|
||||
assert "wrapper resolves creds without infisical" in out
|
||||
|
||||
|
||||
# ── item 4: prose-lint enforces report provenance (real consumer) ─────
|
||||
|
||||
GOOD_CONTRACT = textwrap.dedent("""\
|
||||
---
|
||||
kind: function
|
||||
name: good
|
||||
description: fixture with provenance
|
||||
---
|
||||
|
||||
## Parameters
|
||||
|
||||
- x: y
|
||||
|
||||
## Returns
|
||||
|
||||
ok
|
||||
|
||||
### check-health
|
||||
|
||||
```bash
|
||||
pwd -P
|
||||
```
|
||||
|
||||
**Report format**: Begin with the absolute path the probe executed from.
|
||||
""")
|
||||
|
||||
DECOY_CONTRACT = textwrap.dedent("""\
|
||||
---
|
||||
kind: function
|
||||
name: decoy
|
||||
description: fixture with provenance only outside the report format
|
||||
---
|
||||
|
||||
## Parameters
|
||||
|
||||
- x: y
|
||||
|
||||
## Returns
|
||||
|
||||
ok
|
||||
|
||||
The absolute path of the config is /etc/foo.
|
||||
|
||||
### check-health
|
||||
|
||||
```bash
|
||||
true
|
||||
```
|
||||
|
||||
**Report format**: Summarize actual results from each probe.
|
||||
""")
|
||||
|
||||
MISSING_CONTRACT = textwrap.dedent("""\
|
||||
---
|
||||
kind: function
|
||||
name: missing
|
||||
description: check-health contract with no report format
|
||||
---
|
||||
|
||||
## Parameters
|
||||
|
||||
- x: y
|
||||
|
||||
## Returns
|
||||
|
||||
ok
|
||||
|
||||
### check-health
|
||||
|
||||
```bash
|
||||
pwd -P
|
||||
```
|
||||
""")
|
||||
|
||||
|
||||
def _run_lint(tmp_path, text, name):
|
||||
(tmp_path / name).write_text(text)
|
||||
return subprocess.run(["bash", str(LINT)], cwd=tmp_path,
|
||||
capture_output=True, text=True)
|
||||
|
||||
|
||||
def test_prose_lint_accepts_report_format_with_provenance(tmp_path):
|
||||
result = _run_lint(tmp_path, GOOD_CONTRACT, "good.prose.md")
|
||||
assert result.returncode == 0, result.stdout + result.stderr
|
||||
|
||||
|
||||
def test_prose_lint_rejects_report_format_without_provenance(tmp_path):
|
||||
result = _run_lint(tmp_path, DECOY_CONTRACT, "decoy.prose.md")
|
||||
assert result.returncode == 1, result.stdout
|
||||
assert "lacks execution provenance" in result.stdout
|
||||
|
||||
|
||||
def test_prose_lint_requires_report_format_on_check_health_contract(tmp_path):
|
||||
result = _run_lint(tmp_path, MISSING_CONTRACT, "missing.prose.md")
|
||||
assert result.returncode == 1, result.stdout
|
||||
assert "no **Report format** paragraph" in result.stdout
|
||||
|
||||
|
||||
# ── contracts: parse the executable check-health probe block ─────────
|
||||
|
||||
def _check_health_block(contract):
|
||||
"""Extract the bash probe block under ### check-health (the probe interface)."""
|
||||
text = contract.read_text()
|
||||
marker = "### check-health"
|
||||
assert marker in text, f"{contract.name} has no {marker}"
|
||||
after = text.split(marker, 1)[1]
|
||||
match = re.search(r"```bash\n(.*?)```", after, re.S)
|
||||
assert match, f"{contract.name} check-health has no bash probe block"
|
||||
return match.group(1)
|
||||
|
||||
|
||||
def _loop_nodes(block):
|
||||
nodes = []
|
||||
for line in block.splitlines():
|
||||
match = re.match(r"\s*for\s+\w+\s+in\s+(.+?);?\s*do\b", line)
|
||||
if match:
|
||||
nodes = match.group(1).split()
|
||||
return nodes
|
||||
|
||||
|
||||
def _record(url):
|
||||
"""Normalize a URL into a probe record: host, port, path, expected status."""
|
||||
match = re.match(r"https?://([^/\s\"')]+)(/[^\s\"')]*)?", url)
|
||||
assert match, f"unparseable probe URL: {url}"
|
||||
hostport = match.group(1)
|
||||
if "@" in hostport:
|
||||
hostport = hostport.split("@", 1)[1]
|
||||
if hostport.startswith("["):
|
||||
host, port = hostport[1:hostport.index("]")], None
|
||||
elif ":" in hostport:
|
||||
host, raw_port = hostport.rsplit(":", 1)
|
||||
port = int(raw_port) if raw_port.isdigit() else None
|
||||
else:
|
||||
host, port = hostport, None
|
||||
return {"host": host, "port": port, "path": match.group(2) or "/",
|
||||
"expected": None}
|
||||
|
||||
|
||||
def _probes(block):
|
||||
"""Parse the executable check-health bash block into a normalized probe model.
|
||||
|
||||
Comments are not probes; an `# Expected: <status>` comment annotates the
|
||||
preceding probe. URLs using the block's shell-loop variable `$node` are
|
||||
expanded over the loop's node list.
|
||||
"""
|
||||
loop_nodes = _loop_nodes(block)
|
||||
probes = []
|
||||
last = None
|
||||
for raw in block.splitlines():
|
||||
stripped = raw.strip()
|
||||
if stripped.startswith("#"):
|
||||
expected = re.search(r"Expected:\s*(\d{3})", stripped, re.I)
|
||||
if expected and last is not None:
|
||||
last["expected"] = int(expected.group(1))
|
||||
continue
|
||||
for url in re.findall(r"https?://[^\s\"')]+", raw):
|
||||
hosts = loop_nodes if "$node" in url else [None]
|
||||
for node in hosts:
|
||||
record = _record(url.replace("$node", node) if node else url)
|
||||
probes.append(record)
|
||||
last = record
|
||||
return probes
|
||||
|
||||
|
||||
def test_gpu_monitor_probes_every_gpu_health_on_8080():
|
||||
probes = _probes(_check_health_block(GPU))
|
||||
targets = {(p["host"], p["port"], p["path"]) for p in probes}
|
||||
assert ("192.168.68.8", 8080, "/health") in targets
|
||||
assert ("192.168.68.110", 8080, "/health") in targets
|
||||
|
||||
|
||||
def test_gpu_monitor_never_probes_bare_port_80_on_gpu_hosts():
|
||||
probes = _probes(_check_health_block(GPU))
|
||||
gpu_hosts = {"192.168.68.8", "192.168.68.110", "192.168.68.15"}
|
||||
offenders = [p for p in probes
|
||||
if p["host"] in gpu_hosts and p["port"] in (None, 80)]
|
||||
assert offenders == []
|
||||
|
||||
|
||||
def test_probe_model_flags_explicit_port_80_on_gpu_host():
|
||||
# Regression: a bare-port probe may be spelled with an explicit :80.
|
||||
block = ("curl -s -o /dev/null -w '%{http_code}' "
|
||||
"http://192.168.68.8:80/health\n")
|
||||
gpu_hosts = {"192.168.68.8", "192.168.68.110", "192.168.68.15"}
|
||||
offenders = [p for p in _probes(block)
|
||||
if p["host"] in gpu_hosts and p["port"] in (None, 80)]
|
||||
assert offenders and offenders[0]["port"] == 80
|
||||
|
||||
|
||||
def test_gpu_monitor_treats_router_301_as_alive():
|
||||
probes = _probes(_check_health_block(GPU))
|
||||
unified = [p for p in probes
|
||||
if p["host"] == "192.168.68.116" and p["path"] == "/health/unified"]
|
||||
assert unified, "router /health/unified probe missing"
|
||||
assert unified[0]["expected"] == 301
|
||||
|
||||
|
||||
def test_infra_monitoring_probes_every_real_pve_node():
|
||||
probes = _probes(_check_health_block(INFRA))
|
||||
pve = {(p["host"], p["port"], p["path"]) for p in probes if p["port"] == 8006}
|
||||
assert {host for host, _, _ in pve} == PVE_NODE_IPS
|
||||
assert {path for _, _, path in pve} == {"/api2/json/version"}
|
||||
|
||||
|
||||
def test_infra_monitoring_does_not_probe_ct116_for_pve_api():
|
||||
probes = _probes(_check_health_block(INFRA))
|
||||
assert not any(p["host"] == "192.168.68.116" and p["port"] == 8006
|
||||
for p in probes)
|
||||
Executable
+153
@@ -0,0 +1,153 @@
|
||||
#!/usr/bin/env bash
|
||||
# test_secret_scan.sh — self-test for the commit-time secret guard.
|
||||
#
|
||||
# Run: bash tests/test_secret_scan.sh
|
||||
# Exit: 0 all cases passed, 1 a case failed.
|
||||
#
|
||||
# WHY THIS FILE EXISTS: a scanner that is never observed to fail is not a guard.
|
||||
# Every fixture below is fabricated and pattern-shaped; the test writes it to a
|
||||
# temp tree (a path no allowlist entry covers) and asserts the guard FAILS. The
|
||||
# same fixtures are deliberately listed in scripts/secret-allowlist.tsv, so the
|
||||
# repo-wide tree scan stays quiet while a planted copy still bites — that is the
|
||||
# difference between an explicit, reasoned exception and a guard trained to
|
||||
# ignore a word.
|
||||
#
|
||||
# Only bash + coreutils + grep. No python/node: the Gitea runner executes job
|
||||
# steps inside the runner container, which has neither.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
HERE=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||
ROOT=$(cd -- "$HERE/.." && pwd)
|
||||
SCAN="$ROOT/scripts/secret-scan.sh"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
LAST_OUT=""
|
||||
|
||||
ok() { PASS=$((PASS + 1)); echo " ✅ $1"; }
|
||||
bad() { FAIL=$((FAIL + 1)); echo " ❌ $1"; }
|
||||
|
||||
expect_exit() { # expect_exit <want-code> <label> <cmd...>
|
||||
local want="$1" label="$2"; shift 2
|
||||
local rc
|
||||
LAST_OUT=$("$@" 2>&1); rc=$?
|
||||
if [ "$rc" -eq "$want" ]; then ok "$label (exit $rc)"; else
|
||||
bad "$label (wanted exit $want, got $rc)"
|
||||
printf '%s\n' "$LAST_OUT" | sed 's/^/ /' | head -8
|
||||
fi
|
||||
}
|
||||
|
||||
expect_contains() { # expect_contains <label> <needle>
|
||||
if printf '%s' "$LAST_OUT" | grep -qF -- "$2"; then ok "$1"; else
|
||||
bad "$1 (output did not mention: $2)"
|
||||
fi
|
||||
}
|
||||
|
||||
TMPROOT=$(mktemp -d)
|
||||
trap 'rm -rf "$TMPROOT"' EXIT
|
||||
|
||||
echo "── secret-scan self-test ──"
|
||||
|
||||
# ── 1. Guard syntax ───────────────────────────────────────────────────────
|
||||
expect_exit 0 "scanner parses with bash -n" bash -n "$SCAN"
|
||||
|
||||
# ── 2. Guard FAILS on planted, pattern-matching fixtures ──────────────────
|
||||
mkdir -p "$TMPROOT/planted"
|
||||
cat > "$TMPROOT/planted/ops.env" <<'EOF'
|
||||
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
|
||||
EOF
|
||||
expect_exit 1 "planted sk-or-v1 key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted sk-or-v1 key names the openrouter-key rule" "[openrouter-key]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/curl.sh" <<'EOF'
|
||||
curl -s -H "Authorization: Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb" http://example.invalid/
|
||||
EOF
|
||||
expect_exit 1 "planted literal Bearer token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted Bearer token names the bearer-token rule" "[bearer-token]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/pve.sh" <<'EOF'
|
||||
AUTH="Authorization: PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555"
|
||||
EOF
|
||||
expect_exit 1 "planted Proxmox token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted Proxmox token names the proxmox-token rule" "[proxmox-token]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/deploy-key.pem" <<'EOF'
|
||||
-----BEGIN OPENSSH PRIVATE KEY-----
|
||||
b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW
|
||||
-----END OPENSSH PRIVATE KEY-----
|
||||
EOF
|
||||
expect_exit 1 "planted PEM private key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted PEM key names the private-key rule" "[private-key]"
|
||||
|
||||
# Prose is scanned exactly like code — the original exposures were in .md files.
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/handover.md" <<'EOF'
|
||||
- Admin credentials: `admin` / `correct-horse-battery-staple`
|
||||
EOF
|
||||
expect_exit 1 "planted prose credential line fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted prose line names the cred-prose rule" "[cred-prose]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/config.env" <<'EOF'
|
||||
DB_PASSWORD=correct-horse-battery-staple
|
||||
EOF
|
||||
expect_exit 1 "planted password assignment fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted password assignment names the secret-assign rule" "[secret-assign]"
|
||||
|
||||
# ── 3. Guard stays QUIET on inert values and on the real tree ─────────────
|
||||
mkdir -p "$TMPROOT/inert"
|
||||
cat > "$TMPROOT/inert/config.yaml" <<'EOF'
|
||||
api_key: not-needed
|
||||
bearer_token=monitor_key
|
||||
api_key: $LITELLM_API_KEY
|
||||
EOF
|
||||
expect_exit 0 "env refs, sentinels and variable names are not credentials" bash "$SCAN" --path "$TMPROOT/inert" --quiet
|
||||
|
||||
expect_exit 0 "current repo tree passes the guard" bash "$SCAN" --tree
|
||||
expect_contains "tree run reports the allowlisted exceptions it applied" "allowlisted exception(s)"
|
||||
|
||||
# ── 4. Allowlist entries are path-explicit, not word-based ────────────────
|
||||
# This exact line is allowlisted in infrastructure-control.prose.md; the same
|
||||
# text at an unlisted path must still fail, proving the exception is per-file
|
||||
# and reviewed, not a blanket "ignore the word vault".
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/unlisted.md" <<'EOF'
|
||||
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||
EOF
|
||||
expect_exit 1 "allowlisted text at an unlisted path still fails" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
|
||||
# ── 5. Commit-time mode: the guard blocks a STAGED credential ─────────────
|
||||
# A throwaway git repo with its own copy of the scanner, so this exercises the
|
||||
# real pre-commit path (--staged) without touching this repo's index.
|
||||
mkdir -p "$TMPROOT/repo/scripts"
|
||||
cp "$SCAN" "$TMPROOT/repo/scripts/secret-scan.sh"
|
||||
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/repo/scripts/secret-patterns.tsv"
|
||||
cp "$ROOT/scripts/secret-allowlist.tsv" "$TMPROOT/repo/scripts/secret-allowlist.tsv"
|
||||
git -C "$TMPROOT/repo" init -q
|
||||
git -C "$TMPROOT/repo" -c user.email=t@example.invalid -c user.name=test commit -q --allow-empty -m base
|
||||
cat > "$TMPROOT/repo/planted.env" <<'EOF'
|
||||
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
|
||||
EOF
|
||||
git -C "$TMPROOT/repo" add planted.env
|
||||
expect_exit 1 "staged credential fails at commit time (--staged)" bash "$TMPROOT/repo/scripts/secret-scan.sh" --staged --quiet
|
||||
expect_contains "staged credential names the openrouter-key rule" "[openrouter-key]"
|
||||
|
||||
# ── 6. Fail closed: an allowlist entry without a reason is a hard error ───
|
||||
mkdir -p "$TMPROOT/scanner" "$TMPROOT/clean"
|
||||
cp "$SCAN" "$TMPROOT/scanner/secret-scan.sh"
|
||||
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/scanner/secret-patterns.tsv"
|
||||
printf '*\t*.md\twhatever\n' > "$TMPROOT/scanner/secret-allowlist.tsv"
|
||||
echo "placeholder" > "$TMPROOT/clean/ok.md"
|
||||
expect_exit 2 "allowlist entry with no reason fails closed" bash "$TMPROOT/scanner/secret-scan.sh" --path "$TMPROOT/clean" --quiet
|
||||
|
||||
# ── Verdict ───────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
if [ "$FAIL" -gt 0 ]; then
|
||||
echo "❌ secret-scan self-test FAILED — $PASS passed, $FAIL failed"
|
||||
exit 1
|
||||
fi
|
||||
echo "✅ secret-scan self-test passed ($PASS cases)"
|
||||
@@ -0,0 +1,209 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Behavioural tests for zulip-monitor.sh kagentz C1/C3 legs and the Result verdict.
|
||||
|
||||
WHY THIS FILE EXISTS: the 2026-09-19 kagentz A2A outage was correctly detected
|
||||
by the monitor (C1 returned 000, the run logged an issue) but the lane's own
|
||||
summarization in ops.status wrote "OK" with the note "A2A server DOWN — expected
|
||||
(no credentials configured)". The optimistic verdict came from the lane, not the
|
||||
script. The fix adds a C3 public-access-path leg and makes the Result line say
|
||||
"INCIDENT" when issues are found, so the lane can quote it verbatim.
|
||||
|
||||
CONTRACT UNDER TEST:
|
||||
* C1 (A2A liveness, no credential): 000 → INCIDENT.
|
||||
* C3 (public access path, no credential): 502 → INCIDENT, 000 → INCIDENT,
|
||||
200/302/401 → alive.
|
||||
* Result verdict: when ISSUES > 0, the log's Result line says "INCIDENT",
|
||||
not just "issues found".
|
||||
* Healthy control: C1 401 + C3 302 → 0 issues, "all healthy".
|
||||
|
||||
HOW: behavioural execution using the sandbox pattern already in this repo
|
||||
(tests/test_mumuni_monitor_removal.py). The sandbox copies the shipped monitor
|
||||
verbatim, rewrites only its LOG constant, and runs it with stub ssh/curl on
|
||||
PATH. Each test asserts from the run's own log/verdict, not from file text.
|
||||
|
||||
Usage: python3 -m pytest tests/test_zulip_kagentz_legs.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import pathlib
|
||||
import stat
|
||||
import subprocess
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||
|
||||
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||
|
||||
|
||||
# ── Stub ssh: answers Tanko and Agent Zero probes by env vars ──────────
|
||||
|
||||
SSH_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub ssh: record the target host, then answer by host + remote command.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||
host=""
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
*@192.168.*) host="${a##*@}" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||
cmd="${*: -1}"
|
||||
case "$host" in
|
||||
192.168.68.15)
|
||||
case "$cmd" in
|
||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||
esac ;;
|
||||
192.168.68.14)
|
||||
case "$cmd" in
|
||||
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "${AZ_A2A_EXIT:-0}" ;;
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
# ── Stub curl: serves Zulip server, Abiba health, and C3 public URL ───
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||
# C3 public URL probe (https://kagentz.sysloggh.net/). Record every call.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
case " $* " in
|
||||
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
*kagentz.sysloggh.net*)
|
||||
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
path.write_text(body)
|
||||
path.chmod(path.stat().st_mode
|
||||
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a_code="401", az_a2a_exit=0,
|
||||
kagentz_public_code="302"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
Everything else — legs, labels, notify logic — is the shipped script.
|
||||
"""
|
||||
sandbox = tmp_path / "sandbox"
|
||||
bindir = sandbox / "bin"
|
||||
record = sandbox / "record"
|
||||
bindir.mkdir(parents=True)
|
||||
record.mkdir()
|
||||
|
||||
_write_exec(bindir / "ssh", SSH_STUB)
|
||||
_write_exec(bindir / "curl", CURL_STUB)
|
||||
|
||||
source = ZULIP_MONITOR.read_text()
|
||||
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||
log_path = sandbox / "zulip-health-monitor.log"
|
||||
script = sandbox / "zulip-monitor.sh"
|
||||
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||
|
||||
env = dict(os.environ)
|
||||
env.update({
|
||||
"PATH": f"{bindir}:{env['PATH']}",
|
||||
"RECORD_DIR": str(record),
|
||||
"TANKO_SVC": tanko_svc,
|
||||
"TANKO_HTTP": tanko_http,
|
||||
"AZ_A2A_CODE": az_a2a_code,
|
||||
"AZ_A2A_EXIT": str(az_a2a_exit),
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
return proc, record, log_path
|
||||
|
||||
|
||||
# ── Required behavioural cases ─────────────────────────────────────────
|
||||
|
||||
def test_c3_502_is_incident(tmp_path):
|
||||
"""C3 public leg returns 502 → the run's verdict is an INCIDENT,
|
||||
the C3 line names the 502, and the run is not summarised as healthy."""
|
||||
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="502")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C3 line names the 502.
|
||||
assert "kagentz C3: ❌ public URL 502 (upstream refused)" in log
|
||||
# The verdict is an INCIDENT, not healthy.
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
# The run is not summarised as healthy.
|
||||
assert "✅ 0 issues" not in log
|
||||
|
||||
|
||||
def test_c3_000_is_incident(tmp_path):
|
||||
"""C3 public leg returns 000 → INCIDENT."""
|
||||
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="000")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C3 line reports the connection failure.
|
||||
assert "kagentz C3: ❌ public URL down (HTTP 000)" in log
|
||||
# The verdict is an INCIDENT.
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
assert "✅ 0 issues" not in log
|
||||
|
||||
|
||||
def test_healthy_control_c1_401_c3_302(tmp_path):
|
||||
"""Healthy control: C1 401 plus C3 302 → 0 issues and a healthy verdict,
|
||||
proving the new leg cannot cry wolf."""
|
||||
proc, record, log_path = _run_monitor(tmp_path,
|
||||
az_a2a_code="401",
|
||||
kagentz_public_code="302")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# Both legs report alive.
|
||||
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||
# Zero issues, healthy verdict.
|
||||
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||
# No INCIDENT.
|
||||
assert "INCIDENT" not in log
|
||||
# No notify fired for kagentz.
|
||||
assert "kagentz public URL" not in proc.stdout
|
||||
assert "kagentz A2A server" not in proc.stdout
|
||||
|
||||
|
||||
def test_c1_000_is_incident(tmp_path):
|
||||
"""C1 returns 000 → INCIDENT, keeping the leg that actually caught this
|
||||
outage covered behaviourally."""
|
||||
proc, record, log_path = _run_monitor(tmp_path,
|
||||
az_a2a_code="000",
|
||||
az_a2a_exit=7,
|
||||
kagentz_public_code="302")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C1 line reports the A2A down.
|
||||
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||
# The verdict is an INCIDENT (even though C3 is healthy).
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
# The notify fired for the A2A down.
|
||||
assert "kagentz A2A server DOWN" in proc.stdout
|
||||
Executable
+211
@@ -0,0 +1,211 @@
|
||||
#!/bin/bash
|
||||
# tests/zulip-monitor-abiba.sh — regression test pinning the producer→consumer
|
||||
# contract between the pi Zulip extension's :9200/health payload and the Abiba
|
||||
# leg of scripts/zulip-monitor.sh.
|
||||
#
|
||||
# WHY THIS TEST EXISTS: 2026-09-09 live incident. The monitor parsed the health
|
||||
# payload at the WRONG nesting level (d.get('connected') at top level, while the
|
||||
# extension serves zulip.connected) so PI_CONNECTED was always False and every
|
||||
# monitor run restarted a healthy bot: pm2 showed restarts=8 with the process
|
||||
# created 2026-09-09T09:35:09Z, the monitor log recorded four ❌ Abiba verdicts
|
||||
# (04:23, 05:35, 06:55, 09:35 UTC) and zero ✅, while the Zulip server answered
|
||||
# HTTP 200 and the bot logged a clean connect plus continuing heartbeats. The
|
||||
# watchdog was the fault, not the connection. This test makes that class of
|
||||
# regression fail loudly instead of silently restarting healthy services.
|
||||
#
|
||||
# CONTRACT UNDER TEST (must hold for scripts/zulip-monitor.sh):
|
||||
# * Connection state is NESTED: zulip.connected (boolean) and zulip.last_error
|
||||
# live inside the `zulip` object. There is NO top-level `connected` and NO
|
||||
# retry counter anywhere in the payload (verified against the extension's
|
||||
# startHealthServer handler) — the old retry_count branch was dropped.
|
||||
# * zulip.connected=true -> log "✅ Connected", NO pm2 restart.
|
||||
# * zulip.connected=false -> alert, pm2 restart abiba-zulip.
|
||||
# * fetch error / non-2xx / empty body / unparseable body / missing or
|
||||
# non-boolean zulip.connected -> "⚠️ Probe failed" alert with a
|
||||
# "NOT restarting" label, NO pm2 restart. A parse miss must never kill a
|
||||
# healthy service.
|
||||
# * zulip.connected=true with last_error -> degraded 🟡 warning, no restart.
|
||||
#
|
||||
# HOW: the Abiba leg of the shipped script sits between the
|
||||
# `# -- abiba-leg-start` / `# -- abiba-leg-end` marker comments. This runner
|
||||
# extracts that block verbatim and executes it with a stubbed curl (fixture body
|
||||
# + HTTP code), recorded notify()/pm2 shims, and a temp $LOG. If the markers
|
||||
# disappear (fix reverted or renamed) extraction yields nothing and the suite
|
||||
# fails — the bug cannot return silently.
|
||||
#
|
||||
# Usage: bash tests/zulip-monitor-abiba.sh [path/to/zulip-monitor.sh]
|
||||
# Exit 0 iff every check passes.
|
||||
#
|
||||
# shellcheck disable=SC2034,SC2329,SC1090
|
||||
# LOG/ISSUES and the notify/pm2/curl stubs below are consumed at runtime by
|
||||
# the leg extracted between the marker comments and `source`d in each case;
|
||||
# the static analyzer cannot see across that dynamic source, so it flags them.
|
||||
set -uo pipefail
|
||||
|
||||
ROOT=$(cd "$(dirname "$0")/.." && pwd)
|
||||
SCRIPT=${1:-"$ROOT/scripts/zulip-monitor.sh"}
|
||||
FIXTURES="$ROOT/tests/fixtures"
|
||||
TMP=$(mktemp -d)
|
||||
trap 'rm -rf "$TMP"' EXIT
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
ok() { PASS=$((PASS + 1)); printf ' \033[32m✔\033[0m %s\n' "$1"; }
|
||||
bad() { FAIL=$((FAIL + 1)); printf ' \033[31m✘\033[0m %s\n' "$1"; }
|
||||
|
||||
echo "== tests/zulip-monitor-abiba.sh — Abiba leg vs :9200/health producer contract =="
|
||||
echo "target script: $SCRIPT"
|
||||
|
||||
# --- structural guards -------------------------------------------------------
|
||||
if ! grep -q '^# -- abiba-leg-start' "$SCRIPT"; then
|
||||
echo "✘ FATAL: $SCRIPT has no '# -- abiba-leg-start' marker — the fix has been reverted or renamed."
|
||||
exit 1
|
||||
fi
|
||||
if ! grep -q '^# -- abiba-leg-end' "$SCRIPT"; then
|
||||
echo "✘ FATAL: $SCRIPT has no '# -- abiba-leg-end' marker."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
LEG="$TMP/leg.sh"
|
||||
awk '/^# -- abiba-leg-start/{f=1; next}
|
||||
/^# -- abiba-leg-end/{f=0; next}
|
||||
f' "$SCRIPT" > "$LEG"
|
||||
if [ ! -s "$LEG" ]; then
|
||||
echo "✘ FATAL: extracted Abiba leg is empty."
|
||||
exit 1
|
||||
fi
|
||||
echo "== structural =="
|
||||
if bash -n "$SCRIPT"; then ok "syntax: bash -n $SCRIPT"; else bad "syntax: bash -n $SCRIPT failed"; fi
|
||||
if bash -n "$LEG"; then ok "syntax: extracted leg parses (bash -n)"; else bad "syntax: extracted leg fails bash -n"; fi
|
||||
|
||||
# --- per-case harness ---------------------------------------------------------
|
||||
CURRENT_NAME=""
|
||||
CURRENT_DIR=""
|
||||
|
||||
# $1 case name, $2 http-code, $3 body (file path or literal)
|
||||
run_case() {
|
||||
local name="$1" http="$2" body_src="$3" body
|
||||
CURRENT_NAME="$name"
|
||||
CURRENT_DIR=$(mktemp -d "$TMP/case.XXXXXX")
|
||||
if [ -f "$body_src" ]; then
|
||||
body=$(cat "$body_src")
|
||||
else
|
||||
body="$body_src"
|
||||
fi
|
||||
(
|
||||
LOG="$CURRENT_DIR/log"; ISSUES=0
|
||||
notify() { printf 'ALERT [%s] %s\n' "$1" "$2" >> "$CURRENT_DIR/alerts"; }
|
||||
pm2() { printf 'PM2 %s\n' "$*" >> "$CURRENT_DIR/pm2"; }
|
||||
curl() {
|
||||
local url=""
|
||||
for a in "$@"; do case "$a" in http*) url="$a";; esac; done
|
||||
case "$url" in
|
||||
*:9200/health*)
|
||||
case " $* " in
|
||||
*"-w"*) printf '%s' "$http" ;; # -w '%{http_code}' code probe
|
||||
*) printf '%s' "$body" ;; # body probe
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-CURL %s\n' "$*" >> "$CURRENT_DIR/unexpected-curl"
|
||||
return 7 ;;
|
||||
esac
|
||||
return 0
|
||||
}
|
||||
source "$LEG"
|
||||
)
|
||||
}
|
||||
|
||||
assert_log_has() {
|
||||
if grep -qF -- "$1" "$CURRENT_DIR/log"; then ok "$CURRENT_NAME — log has: $1"; else bad "$CURRENT_NAME — log MISSING: $1"; fi
|
||||
}
|
||||
assert_log_lacks() {
|
||||
if grep -qF -- "$1" "$CURRENT_DIR/log"; then bad "$CURRENT_NAME — log must NOT contain: $1"; else ok "$CURRENT_NAME — log correctly lacks: $1"; fi
|
||||
}
|
||||
assert_alert_has() {
|
||||
if grep -qF -- "$1" "$CURRENT_DIR/alerts"; then ok "$CURRENT_NAME — alert sent: $1"; else bad "$CURRENT_NAME — alert MISSING: $1"; fi
|
||||
}
|
||||
assert_alert_empty() {
|
||||
if [ ! -s "$CURRENT_DIR/alerts" ]; then ok "$CURRENT_NAME — no alert sent (quiet healthy path)"; else bad "$CURRENT_NAME — unexpected alert: $(cat "$CURRENT_DIR/alerts")"; fi
|
||||
}
|
||||
assert_pm2_restarted() {
|
||||
if grep -qF "PM2 restart abiba-zulip" "$CURRENT_DIR/pm2"; then ok "$CURRENT_NAME — pm2 restart abiba-zulip was called"; else bad "$CURRENT_NAME — expected pm2 restart abiba-zulip, pm2 log: $(cat "$CURRENT_DIR/pm2" 2>/dev/null)"; fi
|
||||
}
|
||||
assert_no_restart() {
|
||||
if [ ! -s "$CURRENT_DIR/pm2" ]; then ok "$CURRENT_NAME — NO pm2 restart (fail-safe holds)"; else bad "$CURRENT_NAME — pm2 was called but must NOT be: $(cat "$CURRENT_DIR/pm2")"; fi
|
||||
}
|
||||
assert_no_unexpected_curl() {
|
||||
if [ ! -s "$CURRENT_DIR/unexpected-curl" ]; then ok "$CURRENT_NAME — only :9200/health was probed"; else bad "$CURRENT_NAME — unexpected curl: $(cat "$CURRENT_DIR/unexpected-curl")"; fi
|
||||
}
|
||||
|
||||
# --- case 1: real payload shape, zulip.connected=true -> healthy, no restart --
|
||||
echo "== case 1: connected (real producer payload: nested zulip.connected=true) =="
|
||||
run_case "connected" 200 "$FIXTURES/zulip-health-connected.json"
|
||||
assert_log_has "Abiba: ✅ Connected (processed=0)"
|
||||
assert_log_lacks "Disconnected"
|
||||
assert_alert_empty
|
||||
assert_no_restart
|
||||
assert_no_unexpected_curl
|
||||
|
||||
# --- case 2: zulip.connected=false -> disconnected, restart -------------------
|
||||
echo "== case 2: disconnected (nested zulip.connected=false triggers restart) =="
|
||||
run_case "disconnected" 200 "$FIXTURES/zulip-health-disconnected.json"
|
||||
assert_log_has "Abiba: ❌ Disconnected — restarted"
|
||||
assert_alert_has "DISCONNECTED — restarting"
|
||||
assert_pm2_restarted
|
||||
assert_no_unexpected_curl
|
||||
|
||||
# --- cases 3-9: probe failures must alert and MUST NOT restart ----------------
|
||||
echo "== probe-failure cases: alert 'NOT restarting', zero pm2 restarts =="
|
||||
|
||||
run_case "empty body" 200 ""
|
||||
assert_log_has "Abiba: ⚠️ Probe failed"
|
||||
assert_log_lacks "❌ Disconnected"
|
||||
assert_alert_has "NOT restarting"
|
||||
assert_no_restart
|
||||
|
||||
run_case "garbage body" 200 '{not valid json!!'
|
||||
assert_log_has "Abiba: ⚠️ Probe failed"
|
||||
assert_alert_has "NOT restarting"
|
||||
assert_no_restart
|
||||
|
||||
run_case "missing zulip key" 200 '{"status":"ok","platform":"pi","agent":"abiba"}'
|
||||
assert_log_has "Probe failed"
|
||||
assert_alert_has "NOT restarting"
|
||||
assert_no_restart
|
||||
|
||||
run_case "zulip without connected" 200 '{"status":"ok","zulip":{"last_error":null}}'
|
||||
assert_log_has "Probe failed"
|
||||
assert_alert_has "NOT restarting"
|
||||
assert_no_restart
|
||||
|
||||
run_case "non-boolean connected" 200 '{"status":"ok","zulip":{"connected":"true"}}'
|
||||
assert_log_has "Probe failed"
|
||||
assert_alert_has "NOT restarting"
|
||||
assert_no_restart
|
||||
|
||||
run_case "fetch failure http 000" 000 ""
|
||||
assert_log_has "Probe failed"
|
||||
assert_alert_has "NOT restarting"
|
||||
assert_no_restart
|
||||
|
||||
run_case "non-2xx http 500" 500 '{"error":"boom"}'
|
||||
assert_log_has "Probe failed"
|
||||
assert_alert_has "NOT restarting"
|
||||
assert_no_restart
|
||||
|
||||
# --- case 10: connected but last_error set -> degraded 🟡, no restart ---------
|
||||
echo "== case 10: degraded (connected=true but last_error set) warns, no restart =="
|
||||
run_case "degraded" 200 '{"status":"ok","zulip":{"connected":true,"last_error":"transient queue hiccup","messages_processed":3}}'
|
||||
assert_log_has "Abiba: 🟡 Error: transient queue hiccup"
|
||||
assert_log_lacks "❌ Disconnected"
|
||||
assert_no_restart
|
||||
|
||||
# --- summary -------------------------------------------------------------------
|
||||
echo ""
|
||||
if [ "$FAIL" -eq 0 ]; then
|
||||
echo "✅ ALL CHECKS PASSED ($PASS/$PASS) — tests/zulip-monitor-abiba.sh"
|
||||
exit 0
|
||||
else
|
||||
echo "❌ $FAIL CHECK(S) FAILED ($PASS passed) — tests/zulip-monitor-abiba.sh"
|
||||
exit 1
|
||||
fi
|
||||
+321
-77
@@ -1,24 +1,36 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-health
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Hermes agents Tanko/Mumuni), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||
version: 3.0.0
|
||||
version: 3.4.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
---
|
||||
|
||||
# Zulip Mesh Health Monitor
|
||||
|
||||
Monitors ALL Zulip-connected agents across three platforms (pi, Hermes, Agent Zero).
|
||||
Runs every 15 minutes in the background. Also triggers on session start.
|
||||
Monitors the Zulip-connected agents under this host's operational control (pi,
|
||||
DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on
|
||||
session start.
|
||||
|
||||
> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She
|
||||
> moved off this host onto her own container — kagentz CT 105 on minipve
|
||||
> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her
|
||||
> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may
|
||||
> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on
|
||||
> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat,
|
||||
> response delivery) are retired: they always read "unknown" against the
|
||||
> decommissioned deployment and produced a false 🔴 alert on every run.
|
||||
|
||||
## Requires
|
||||
|
||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on hwepve), and Agent Zero Docker host (192.168.68.14)
|
||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||
- **PM2** on localhost for pi process management
|
||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||
- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200`
|
||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||
- **Relay access** via RA-H OS MCP for alert delivery
|
||||
|
||||
@@ -45,11 +57,9 @@ Runs every 15 minutes in the background. Also triggers on session start.
|
||||
"severity": "healthy"
|
||||
},
|
||||
"tanko": {
|
||||
"platform": "hermes",
|
||||
"zulip_state": "connected",
|
||||
"heartbeat_age_seconds": 45,
|
||||
"gateway_pid": 1234,
|
||||
"edit_fail_rate_pct": 0,
|
||||
"platform": "dsh",
|
||||
"service_state": "active",
|
||||
"http_status": 200,
|
||||
"severity": "healthy"
|
||||
}
|
||||
}
|
||||
@@ -81,13 +91,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
|
||||
## Streaming Support (2026-07-05)
|
||||
|
||||
Zulip agents now support progressive message editing during agent generation.
|
||||
When a Hermes agent (Tanko, Mumuni) processes a message, the response is
|
||||
When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is
|
||||
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
||||
|
||||
- Adapter implements `edit_message()` using `_api_patch()` helper
|
||||
- Gateway stream consumer progressively edits the Zulip message
|
||||
- User sees real-time agent thinking instead of waiting for full response
|
||||
- Verified: Tanko (CT 112) and Mumuni (CT 114) both have streaming active
|
||||
- Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here
|
||||
|
||||
### Verification
|
||||
```bash
|
||||
@@ -117,20 +127,48 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
|
||||
|
||||
## Execution
|
||||
|
||||
### Liveness rule (scoped)
|
||||
|
||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN.
|
||||
|
||||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above).
|
||||
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
||||
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
||||
|
||||
### Step 1: Zulip Server Liveness
|
||||
|
||||
```bash
|
||||
curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \
|
||||
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY'
|
||||
# Probe the Zulip API (authenticated, any HTTP status = ALIVE)
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
|
||||
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code"
|
||||
fi
|
||||
```
|
||||
|
||||
Expected: `200`. If not → mark `zulip_server_status: "down"`, skip per-platform checks, alert.
|
||||
Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault.
|
||||
|
||||
### Step 2: Platform A — pi (Abiba, localhost)
|
||||
|
||||
**A1: Health Endpoint**
|
||||
|
||||
Fetch `http://localhost:9200/health` as JSON. Check:
|
||||
```bash
|
||||
# Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE)
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health)
|
||||
echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Abiba extension http://127.0.0.1:9200/health -> $code"
|
||||
fi
|
||||
```
|
||||
|
||||
Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.**
|
||||
|
||||
Check the JSON payload:
|
||||
|
||||
| Field | Healthy | Critical |
|
||||
|-------|---------|----------|
|
||||
@@ -182,102 +220,302 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
|
||||
| `last_error` set | Log and monitor |
|
||||
| Crash loop >10/h | Alert user |
|
||||
|
||||
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24)
|
||||
|
||||
**B1: Gateway State**
|
||||
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
|
||||
|
||||
Mumuni is out of scope for this host (see the note above): she runs on her own
|
||||
container and is monitored on her side.
|
||||
|
||||
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
|
||||
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
|
||||
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **amdpve**
|
||||
PVE host (**192.168.68.15**). Direct SSH to 192.168.68.122 is not a dependency
|
||||
of this contract — per-worker key availability varies — so CT 112 probes run
|
||||
from the amdpve vantage via `pct exec`:
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
|
||||
ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100
|
||||
ssh root@192.168.68.15 "pct exec 112 -- <command>"
|
||||
```
|
||||
|
||||
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
|
||||
> **By design (verified 2026-09-08):** the `dsh-web` gateway binds
|
||||
> `127.0.0.1:3080` **loopback-only**. A remote probe against
|
||||
> `192.168.68.122:3080` gets connection-refused — that is EXPECTED, NOT a fault,
|
||||
> and must never be raised as Tanko down. Only loopback probes from inside
|
||||
> CT 112 (or the public-URL fallback below) are valid health signals.
|
||||
|
||||
**B2: Agent Process**
|
||||
**B1: Gateway Service State (Tanko)**
|
||||
|
||||
```bash
|
||||
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web"
|
||||
```
|
||||
|
||||
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling):
|
||||
Expected: `active`. Anything else → gateway service down → apply the Tanko heal
|
||||
(restart via DSH service, Platform B Actions table below).
|
||||
|
||||
| Agent | Restart command | Notes |
|
||||
|-------|-----------------|-------|
|
||||
| Mumuni (.24) | `pm2 restart abiba-zulip` | Hermes gateway runs under PM2 as `abiba-zulip` |
|
||||
| Tanko (.122) | `bash /opt/hermes-zulip-plugin/run.sh` (or the agent's systemd/user unit) | Tanko does NOT use PM2 — never run `pm2 restart mumuni-zulip` for Tanko (process does not exist) |
|
||||
|
||||
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
||||
|
||||
**B3: Heartbeat Verification**
|
||||
**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)**
|
||||
|
||||
```bash
|
||||
ssh root@<CT> "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
|
||||
```
|
||||
|
||||
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
|
||||
Silence > 300s → warning. Silence > 600s → critical.
|
||||
Alive = **ANY** HTTP status response from the endpoint — the expected set is
|
||||
`200`/`301`/`302`/`307`/`308`/`401`/`403` (the gateway UI is token-gated and
|
||||
legitimately answers with redirects/auth-challenges, so never require a bare
|
||||
`200`), and any other status, including `404`/`5xx`, also counts alive: a
|
||||
process answering `503` is running and self-heal must NOT restart-loop it.
|
||||
Down = connection refused (`000`) or timeout only. Statuses outside the
|
||||
expected set are logged/reported as a warning — reported, never healed on.
|
||||
|
||||
**B4: Response Delivery**
|
||||
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)**
|
||||
|
||||
```bash
|
||||
ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
||||
curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/
|
||||
```
|
||||
|
||||
> 50% fail rate → critical.
|
||||
Fallback only — used when the monitoring node has no pct/SSH path to amdpve.
|
||||
Alive = **ANY** HTTP status response from the endpoint — healthy signals are
|
||||
`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other
|
||||
status, including `404`/`5xx`, also counts alive: the endpoint is up and
|
||||
answering and must NOT be restart-looped. Down = connection refused (`000`) or
|
||||
timeout only. Never expect a bare `200` — the public URL terminates in the
|
||||
token-gated authentik chain. Statuses outside the healthy set are
|
||||
logged/reported as a warning — reported, never healed on.
|
||||
|
||||
**Platform B Actions**
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
|
||||
| No heartbeat in 10min | Same as above |
|
||||
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
||||
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
||||
| `dsh-web` service not `active` | Restart Tanko via DSH service |
|
||||
| HTTP `:3080` connection refused/timeout (`000`) | Same as above |
|
||||
| HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
|
||||
|
||||
**B4: dsh-web Authentication (Tanko — restart-persistent login)**
|
||||
|
||||
The dsh-web UI is token-gated. On every start the process prints a random
|
||||
launch token to the journal:
|
||||
|
||||
```
|
||||
dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||||
```
|
||||
|
||||
The token only bootstraps an authority-bound, HMAC-signed browser cookie with a
|
||||
30-day lifetime. The signing secret is durable in
|
||||
`/root/.dsh/.credentials.yaml` (key `client-connection/browser-session`), so a
|
||||
cookie minted once keeps working across `dsh-web` restarts; the launch token
|
||||
itself rotates on every restart.
|
||||
|
||||
**Login endpoint (public, Authentik-gated):**
|
||||
`https://tankodhs.sysloggh.net/dsh-web-login`
|
||||
|
||||
It lives inside the Authentik-gated `:80` server block
|
||||
(`/etc/nginx/sites-available/dsh`, symlinked from
|
||||
`/etc/nginx/sites-enabled/dsh`) as `location = /dsh-web-login`, guarded by
|
||||
`auth_request /outpost.goauthentik.io/auth/nginx`. It proxies to dsh-web with
|
||||
`Host: tankodhs.sysloggh.net`, so the minted cookie is bound to the public
|
||||
authority — never to `127.0.0.1:3080`. The token-dependent line is isolated in
|
||||
the generated include `/etc/dsh-web/nginx-login.conf`:
|
||||
|
||||
```
|
||||
proxy_pass http://127.0.0.1:3080/?token=<TOKEN>;
|
||||
```
|
||||
|
||||
**Token refresh (non-disruptive):**
|
||||
`/opt/deepseek-harness/capture-dsh-token.sh` (source:
|
||||
`scripts/capture-dsh-token.sh`) reads candidate launch tokens from the journal
|
||||
**scoped to the service's current systemd invocation**
|
||||
(`systemctl show -p InvocationID` + `_SYSTEMD_INVOCATION_ID=`), re-sampling the
|
||||
invocation on every pass so a restart that lands during the wait switches to the
|
||||
new invocation; a restarted process's stale token is never considered while its
|
||||
new startup banner is still pending and there is no whole-journal or
|
||||
cross-invocation fallback. Each candidate
|
||||
is then functionally verified against dsh-web with `Host: tankodhs.sysloggh.net`,
|
||||
using the first the running process accepts with `303`. It waits up to 120s for
|
||||
a restarted process to accept a token and re-probes every current-invocation
|
||||
candidate on each pass, so a token that briefly returns `000` while the service
|
||||
is still starting is not disqualified. If none is accepted it leaves the include
|
||||
untouched and exits so the timer retries (exiting non-zero when a pending reload
|
||||
is still outstanding). It writes
|
||||
`/etc/dsh-web/launch-token` and regenerates `/etc/dsh-web/nginx-login.conf`,
|
||||
reloading nginx only when the on-disk include differs from the generated one or
|
||||
the applied-state stamp does not match the token (`nginx -t` guards the reload,
|
||||
and the stamp is written only after a successful `nginx -s reload`, so a failed
|
||||
or interrupted reload is retried on the next run). Any failed reload records a
|
||||
pending-reload marker under `/etc/dsh-web/`; the next run attempts the reload
|
||||
before the token wait, independent of token state, and clears the marker only
|
||||
once the reload succeeds, so a disabled legacy `:8081` file can never leave the
|
||||
running nginx unreloaded. The generated include is recreated before any
|
||||
`nginx -t` if it is missing, so a failed run cannot wedge recovery.
|
||||
Runs are serialized with `flock` on `/run/capture-dsh-token.lock`. It **never
|
||||
stops or starts `dsh-web`**.
|
||||
It is triggered by the `dsh-web.service` drop-in
|
||||
`/etc/systemd/system/dsh-web.service.d/20-token-refresh.conf`
|
||||
(`ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service`) and by
|
||||
`dsh-web-token.timer` every 2 minutes for reconciliation.
|
||||
|
||||
<details><summary>Installed systemd wiring (CT 112)</summary>
|
||||
|
||||
```ini
|
||||
# /etc/systemd/system/dsh-web-token.service
|
||||
[Unit]
|
||||
Description=Refresh the dsh-web launch token for the nginx login endpoint
|
||||
After=dsh-web.service
|
||||
[Service]
|
||||
Type=oneshot
|
||||
TimeoutStartSec=180
|
||||
ExecStart=/opt/deepseek-harness/capture-dsh-token.sh
|
||||
|
||||
# /etc/systemd/system/dsh-web-token.timer
|
||||
[Unit]
|
||||
Description=Periodically refresh the dsh-web login token
|
||||
[Timer]
|
||||
OnBootSec=90s
|
||||
OnUnitActiveSec=120s
|
||||
AccuracySec=10s
|
||||
Persistent=true
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
|
||||
# /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf
|
||||
[Service]
|
||||
ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
> **Do NOT reintroduce the `:8081` endpoint.** It listened on `0.0.0.0:8081`
|
||||
> with no `auth_request` and was a full Authentik bypass for anyone on the LAN.
|
||||
> The script now removes `/etc/nginx/sites-enabled/dsh.token` automatically if
|
||||
> it ever reappears.
|
||||
|
||||
**Authentication flow:**
|
||||
1. `GET https://tankodhs.sysloggh.net/dsh-web-login`
|
||||
2. Unauthenticated → Authentik sign-in; once authenticated the request reaches
|
||||
dsh-web with `Host: tankodhs.sysloggh.net`.
|
||||
3. dsh-web accepts the launch token on `GET /`, writes the
|
||||
`dsh-auth-<authority-hash>` cookie (30 days, `HttpOnly`, `SameSite=Strict`)
|
||||
and returns `303` to `/`.
|
||||
4. Every later request through `/` presents that cookie; the token is not needed
|
||||
again until the cookie expires or a new browser is used.
|
||||
|
||||
**Verification** (amdpve vantage):
|
||||
```bash
|
||||
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
|
||||
# Expected: 302
|
||||
|
||||
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
||||
-w '%{http_code}\n' http://192.168.68.122:8081/"
|
||||
# Expected: 000
|
||||
|
||||
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
|
||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the minted dsh-auth-... cookie (authority
|
||||
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
|
||||
|
||||
# 4. Token refresh is non-disruptive and idempotent.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
||||
# Expected: "token unchanged; nginx not reloaded" when nothing changed
|
||||
```
|
||||
|
||||
**Restart durability (acceptance):** after `systemctl restart dsh-web`, (a) the
|
||||
cookie minted before the restart still returns `200` on `/`, and (b) the
|
||||
refreshed `/etc/dsh-web/nginx-login.conf` carries the new token and mints a
|
||||
fresh cookie. Both verified live 2026-09-11.
|
||||
|
||||
```bash
|
||||
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
|
||||
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
|
||||
# until the socket answers (any status but 000) before asserting the cookie.
|
||||
for i in $(seq 1 60); do
|
||||
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
|
||||
[ "$UP" != "000" ] && break
|
||||
sleep 2
|
||||
done
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the pre-restart cookie is still accepted.
|
||||
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
|
||||
# manual run may no-op on the flock, so poll until the include carries a token
|
||||
# the running process accepts (bounded wait) before the mint+reuse check.
|
||||
for i in $(seq 1 60); do
|
||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
||||
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
|
||||
[ "$CODE" = "303" ] && break
|
||||
sleep 2
|
||||
done
|
||||
# Expected: 303 — the include now holds the token the running process accepts.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the refreshed token minted a fresh cookie.
|
||||
```
|
||||
|
||||
|
||||
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
|
||||
|
||||
**C1: A2A Server Health**
|
||||
> **The kagentz Zulip adapter leg is retired (2026-09-12).** Its code
|
||||
> (`/a0/usr/kagentz-zulip/`) no longer exists in the agent-zero container, so
|
||||
> the former adapter-process and heartbeat/queue checks always failed and the
|
||||
> monitor issued a restart for something that could not start, posting a false
|
||||
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
|
||||
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
|
||||
> liveness/response and public-path access only, and a probe must never restart
|
||||
> a platform.
|
||||
|
||||
**C1: A2A Server Health (no credential needed)**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json"
|
||||
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
|
||||
# is auth-gated: an unauthenticated probe gets 401, which means the server is up.
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
|
||||
```
|
||||
|
||||
Expected: `{"name":"kagentz",...}`. Connection refused → A2A server down.
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down (INCIDENT). Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**C2: Adapter Process**
|
||||
**C2: A2A Response Verification (requires LITELLM_KEY)**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.14 "docker exec agent-zero ps aux | grep adapter | grep -v grep"
|
||||
```
|
||||
|
||||
Adapter should be running. Missing → restart inside container.
|
||||
|
||||
**C3: Heartbeat & Queue**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.14 "docker exec agent-zero grep Heartbeat /tmp/zulip-adapter.log | tail -3"
|
||||
```
|
||||
|
||||
Check: `processed=N` incrementing, `silence < 600s`, `reconnects` ≈ 0.
|
||||
|
||||
**C4: A2A Response Verification**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:8001/a2a \
|
||||
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:80/a2a \
|
||||
-H 'Content-Type: application/json' \
|
||||
-H 'Authorization: Bearer $LITELLM_KEY' \
|
||||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||
```
|
||||
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`.
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set (this is a credential issue, NOT a server-down incident).
|
||||
|
||||
**C3: Public Access Path (no credential needed)**
|
||||
|
||||
```bash
|
||||
# Probes the public URL that NetBird proxies to the agent-zero container.
|
||||
# This is the captain's point of view: if the captain can't reach it, it's down.
|
||||
# 200/302/401 = alive, 502 = proxy's "upstream refused" page (INCIDENT),
|
||||
# connection failed (000) = INCIDENT. Never restarts anything.
|
||||
curl -s -o /dev/null --connect-timeout 10 --max-time 15 -w '%{http_code}' https://kagentz.sysloggh.net/
|
||||
```
|
||||
|
||||
Expected: `200` (Agent Zero login page), `302` (redirect), or `401` (auth-gated) = alive. `502` = NetBird proxy's "upstream refused" page (the container's port 80 is not listening) = INCIDENT. `000` (connection failed) = INCIDENT. Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**Platform C Actions**
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| A2A `.well-known/agent.json` fails | `docker exec agent-zero bash -c "pkill -9 -f a2a_agent; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &"` |
|
||||
| Adapter process missing | Restart adapter inside container with env vars |
|
||||
| Silence > 600s | Restart adapter (auto-reconnect handles BAD_EVENT_QUEUE_ID) |
|
||||
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
|
||||
| C1 A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||
| C1 A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||
| C2 A2A returns 401 | Check `LITELLM_KEY` is set (credential issue, NOT server-down) |
|
||||
| C3 public URL returns `502` (upstream refused) | Alert only — the container's port 80 is not listening; investigate the agent-zero container |
|
||||
| C3 public URL returns `000` (connection failed) | Alert only — the public path is down; investigate the NetBird proxy or the container |
|
||||
| C3 public URL returns a status other than `200`/`302`/`401` | Log/report as a warning — reported, never healed on |
|
||||
|
||||
### Step 5: Global Checks
|
||||
|
||||
@@ -285,19 +523,25 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`.
|
||||
|
||||
Check each agent's log for excessive bot-to-bot chatter:
|
||||
- Abiba: `Skipped.*bot msgs` count
|
||||
- Tanko/Mumuni: Repeated DM exchanges between bots
|
||||
- kagentz: Adapter log for bot DMs being processed
|
||||
- Tanko: Repeated DM exchanges between bots
|
||||
|
||||
If any bot processes >50 bot-originated messages in 15min → warning.
|
||||
|
||||
### Step 6: Compile and Report
|
||||
|
||||
1. Compile all platform checks and severity
|
||||
2. Determine `overall_severity` from worst per-agent severity
|
||||
3. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||
4. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||
5. If any agent critical or >2 degraded: send relay message to user
|
||||
6. Update `last_check` timestamp in `### Maintains` snapshot
|
||||
1. Run `scripts/zulip-monitor.sh` and take its final `Result:` line as the
|
||||
authoritative run verdict. The verdict line is either
|
||||
`Result: ✅ 0 issues (all healthy)` or
|
||||
`Result: 🔴 INCIDENT — N issue(s) found`.
|
||||
2. Quote that `Result:` line verbatim in the status report. When it says
|
||||
`INCIDENT`, the run MUST be reported as an incident — never summarised as
|
||||
OK/healthy and never annotated as "expected".
|
||||
3. Compile all platform checks and severity
|
||||
4. Determine `overall_severity` from worst per-agent severity
|
||||
5. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||
6. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||
7. If any agent critical or >2 degraded: send relay message to user
|
||||
8. Update `last_check` timestamp in `### Maintains` snapshot
|
||||
|
||||
### Restart Debounce
|
||||
|
||||
|
||||
@@ -10,8 +10,9 @@ description: >
|
||||
|
||||
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
|
||||
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
|
||||
> monitoring now happens through Telegram. Hermes agents (Tanko, Mumuni) and
|
||||
> Agent Zero (kagentz) continue to use Zulip.
|
||||
> monitoring now happens through Telegram. Mumuni (Hermes) and Tanko (DSH)
|
||||
> continue to use Zulip; Agent Zero's Zulip adapter is retired — see
|
||||
> `zulip-health.prose.md` for current Platform C (Agent Zero) state.
|
||||
|
||||
## Maintains
|
||||
|
||||
|
||||
@@ -420,7 +420,7 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S
|
||||
|-------|----------|-------------|--------------|-------------|
|
||||
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
|
||||
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
|
||||
| **Mumuni** | Hermes (CT 114) | ✅ Connected | No issues found | None needed |
|
||||
| **Mumuni** | Hermes (kagentz CT 105, migrated 2026-08-29) | ✅ Connected | No issues found | None needed |
|
||||
|
||||
### Key Fixes Applied
|
||||
|
||||
|
||||
@@ -13,7 +13,8 @@ triggers:
|
||||
|
||||
> **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code
|
||||
> (`performHealthCheck()`). That code has been removed. Zulip self-healing for
|
||||
> Hermes agents (Tanko, Mumuni) continues through their own gateway monitoring.
|
||||
> agents (Mumuni on Hermes, Tanko on DSH) continues through their own platform
|
||||
> monitoring.
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -65,8 +66,7 @@ triggers:
|
||||
|------|----|------|---------|
|
||||
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
|
||||
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
|
||||
| Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` |
|
||||
| Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` |
|
||||
| Tanko | 192.168.68.122 (CT 112) | jerome | DSH (DeepSeek Harness) — restart via DSH service, not `hermes gateway restart` (no longer a Hermes agent since 2026-08-27) |
|
||||
|
||||
## Debounce
|
||||
|
||||
@@ -75,9 +75,8 @@ Track via `/tmp/zulip-heal-debounce-<agent>` (unix timestamp of last restart).
|
||||
|
||||
## Reporting
|
||||
|
||||
Every cycle produces a knowledge graph node:
|
||||
- Title: `[LEARN] zulip-self-heal: <timestamp>`
|
||||
- metadata: { type: "remediation", status: "fixed" | "escalated" | "healthy" }
|
||||
This contract is RETIRED — health-check logs are NOT knowledge graph content.
|
||||
No graph nodes are created. Logs go to Gitea (SyslogSolution/health-logs).
|
||||
- Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>"
|
||||
- Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention"
|
||||
|
||||
|
||||
Reference in New Issue
Block a user