Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
767bd22d9c | ||
|
|
e97145c88f | ||
|
|
55f1208eb8 | ||
|
|
b9adf353ee | ||
|
|
aeb66ea22d | ||
|
|
288f74cf84 | ||
|
|
c66671dbee | ||
|
|
b80d3142aa | ||
|
|
266fa1f835 | ||
|
|
85ea1f4f3d | ||
|
|
b2a259fa23 | ||
|
|
fb185ed90a | ||
|
|
b001b657d4 | ||
|
|
a1ffeaad34 | ||
|
|
287657a77a | ||
|
|
21f9073e0b | ||
|
|
32fe7c0652 | ||
|
|
25cf2f5eef | ||
|
|
26f2301188 | ||
|
|
a6a459acc0 | ||
|
|
14bed6e916 | ||
|
|
2e70c834cb | ||
|
|
4f59b82404 | ||
|
|
8a4ee4cf5a | ||
|
|
952fca9c92 | ||
|
|
b1462f3e79 | ||
|
|
bfdff13ae7 | ||
|
|
8bf32f6f0f |
@@ -628,18 +628,18 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 3.0.0
|
||||
version: 3.2.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/15 * * * *'
|
||||
description: "Every 15 minutes \u2014 monitors all Zulip-connected agents"
|
||||
description: "Every 15 minutes \u2014 monitors the Zulip-connected agents under this host's control (pi, DSH, Agent Zero)"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 120
|
||||
requires:
|
||||
- Zulip API key for abiba-bot@chat.sysloggh.net
|
||||
- SSH access to all Hermes agents
|
||||
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
|
||||
verification:
|
||||
postconditions:
|
||||
- check: bot registration active
|
||||
|
||||
@@ -3,15 +3,16 @@ kind: responsibility
|
||||
name: infrastructure-update
|
||||
description: >
|
||||
Autonomous system-wide update contract covering all 5 Proxmox nodes,
|
||||
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
|
||||
15+ containers/VMs, and 5 Docker ecosystems (docker-vm .7, CT 116 .116,
|
||||
CT 117, hwpve .11, NetBird VPS 72.61.0.17). Updates apt packages,
|
||||
Docker images, and container stacks in safe waves with health checks
|
||||
and automatic rollback on failure.
|
||||
agent: abiba
|
||||
triggers:
|
||||
- on "infra update" command
|
||||
- weekly (Sunday 03:00 EDT) via cron
|
||||
- weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
|
||||
- on security advisory relay from Mumuni
|
||||
version: 1.2.0
|
||||
version: 1.3.0
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -80,7 +81,13 @@ Before ANY update wave:
|
||||
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
|
||||
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
|
||||
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
|
||||
| CT 117 (zulip, storepve) | Zulip | `docker pull zulip/docker-zulip:latest && docker restart zulip-zulip-1` |
|
||||
| CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) |
|
||||
| CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) |
|
||||
| hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` |
|
||||
| NetBird VPS (72.61.0.17) | NetBird (server, dashboard, proxy, traefik, crowdsec) | `ssh root@72.61.0.17 'cd /root && docker compose pull && docker compose up -d'` |
|
||||
| VM 109 (.7) | Trove test | `cd /opt/trove-test && docker compose pull && docker compose up -d` |
|
||||
| VM 109 (.7) | docker-stats | `cd /opt/docker-stats && docker compose pull && docker compose up -d` |
|
||||
| CT 116 (.116) | Monitoring (Grafana, Prometheus, Alertmanager, PVE exporter) | `cd /opt/monitoring && docker compose pull && docker compose up -d` |
|
||||
|
||||
**Verify after Wave 3:**
|
||||
- All containers healthy: `docker ps` on each host
|
||||
@@ -88,8 +95,12 @@ Before ANY update wave:
|
||||
- MCP integration test: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools (23 RA-H OS + 67 GitHub)
|
||||
- Zulip test: send test message to #agent-hub
|
||||
- Dashboard loading: `curl localhost:3001/` (via CT 116)
|
||||
- Firecrawl test: `curl :3002/`
|
||||
- Firecrawl test: `curl -X POST http://192.168.68.7:3002/v1/search -H 'Content-Type: application/json' -d '{"query":"health","limit":1}'` → `"success":true` (GET `/` returns 200)
|
||||
- Authentik test: `curl http://192.168.68.11:9000/` → 302 redirect to login
|
||||
- NetBird test: `curl -s -o /dev/null -w '%{http_code}' https://netbird.sysloggh.net/` → 200
|
||||
- harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
|
||||
- SearXNG test: `curl :8888`
|
||||
- Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
|
||||
|
||||
## Wave 4: Proxmox Kernel Reboot
|
||||
|
||||
@@ -158,6 +169,8 @@ Before Wave 1, snapshot these files:
|
||||
/opt/search-stack/searxng/docker-compose.yml (VM 109 .7)
|
||||
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
||||
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
||||
/root/compose.yml (hwpve .11 — Authentik server/worker/postgres)
|
||||
/root/docker-compose.yml (NetBird VPS — netbird server/dashboard/proxy, traefik, crowdsec)
|
||||
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
||||
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
|
||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||
@@ -220,7 +233,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
||||
|
||||
- [ ] All 5 PVE nodes updated, no reboot-loop
|
||||
- [ ] All VMs/CTs running post-update
|
||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
|
||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117 + hwpve .11 + NetBird VPS)
|
||||
- [ ] LiteLLM inference passing (syslog-auto test)
|
||||
- [ ] Zulip server + all 3 agents connected
|
||||
- [ ] GPU fleet at full capacity (3/3)
|
||||
|
||||
@@ -113,7 +113,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||
|
||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||
|
||||
## Maintains
|
||||
|
||||
+28
-6
@@ -6,7 +6,7 @@ name: memory-fixer
|
||||
description: >
|
||||
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
||||
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||
version: 2.0.0
|
||||
version: 2.1.0
|
||||
---
|
||||
---
|
||||
|
||||
@@ -69,13 +69,15 @@ FROM nodes
|
||||
WHERE json_extract(metadata, '$.namespace') IS NULL;
|
||||
```
|
||||
|
||||
### 3. Staleness Review Tagging
|
||||
### 3. Staleness Review Tagging (refresh-suggested nodes only)
|
||||
|
||||
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
|
||||
|
||||
**Archive-suggested nodes are NO LONGER tagged — they are archived outright (see Level 1 fix 4).** Tagging with `[REVIEW: refresh]` applies only to living nodes (infrastructure, deployment, system, system-health, business, philosophy, research, learning, investigation, analysis, project, agent, registry, policy).
|
||||
|
||||
**Exclusion Rules:**
|
||||
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
|
||||
- Nodes whose `description` already starts with `[REVIEW:` are NOT re-processed
|
||||
- Nodes whose `description` already starts with `[REVIEW:` or `[ARCHIVED]` are NOT re-processed
|
||||
|
||||
```sql
|
||||
SELECT id, title, json_extract(metadata, '$.type') as node_type,
|
||||
@@ -104,9 +106,28 @@ LIMIT 10;
|
||||
|
||||
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`.
|
||||
|
||||
### 4. Stale-Node Archiving (Level 1 — standing Kwame directive, 2026-09-11)
|
||||
|
||||
**Kwame's standing directive: stale nodes CAN be archived by the fixer. No per-batch escalation, no `[REVIEW: archive]` tagging — archive them.**
|
||||
|
||||
For every node whose suggested action is `archive` (i.e. its type is NOT one of the living types in fix 3), archive it in a **single** `updateNode` call:
|
||||
|
||||
```python
|
||||
updateNode(id, {
|
||||
"description": "[ARCHIVED] " + originalDescriptionWithoutReviewTag,
|
||||
"metadata": {"state": "archived"}
|
||||
})
|
||||
```
|
||||
|
||||
- `state` transitions **DO work through `updateNode`** (`archived`, and back to `active`). The former "state only accepts processed/not_processed, use SSH" claim was wrong — verified 2026-09-11 by archiving 7 nodes (#61, #373, #388, #465, #475, #526, #1476) over the bridge with `updated_at` auto-bumping. **SSH to the bridge host is a fallback, not a requirement**, and it is blocked from kagentz anyway.
|
||||
- Pass `description` and `metadata` in the **same** call, and always keep the `updates` object nested: `{"id": N, "updates": {…}}`.
|
||||
- Archiving is non-destructive: the node stays in the graph, marked `state: archived` + `[ARCHIVED] ` prefix. **Living nodes (refresh-suggested) are NEVER archived** without a specific Kwame decision — they are the cluster/agent/business canon.
|
||||
|
||||
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
|
||||
|
||||
## Level 2 Escalations (Kwame Decision Required)
|
||||
|
||||
1. **Stale nodes** flagged with `[REVIEW: …]` — Archive, refresh, or keep?
|
||||
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
|
||||
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
|
||||
3. **Orphan Nodes >90 days old** — Archive or connect?
|
||||
|
||||
@@ -144,7 +165,7 @@ Reply with:
|
||||
|
||||
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
|
||||
|
||||
> ⚠️ `updateNode` cannot set `state` to non-standard values (restricted to `processed`/`not_processed`) and cannot add metadata keys. For state transitions and `updated_at` bumps, use **direct SSH + SQLite** on the bridge host:
|
||||
> ⚠️ **Corrected 2026-09-11:** `updateNode` DOES accept `state` changes — `{"updates": {"description": …, "metadata": {"state": "archived"}}}` works over the bridge, and `updated_at` bumps automatically. The old "use direct SSH + SQLite for state transitions" instruction was based on a wrong assumption; SSH is a fallback only (and is blocked from kagentz). Use one `updateNode` call for both the tag and the state.
|
||||
> ```bash
|
||||
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
|
||||
> ```
|
||||
@@ -172,8 +193,9 @@ The result must be 0 rows when all decisions are executed. Report what was done.
|
||||
## Checks
|
||||
|
||||
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
||||
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
|
||||
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
||||
- **Timestamps:** every executed decision bumps `updated_at`, so the node exits the stale window on the next run.
|
||||
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
|
||||
|
||||
## Logging
|
||||
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
|
||||
|
||||
@@ -17,8 +17,9 @@ Changelog:
|
||||
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
|
||||
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
|
||||
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
|
||||
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
|
||||
abiba (.24).
|
||||
Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24).
|
||||
(v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she
|
||||
moved to her own container, kagentz CT 105 / .14, and is monitored there.)
|
||||
v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service,
|
||||
.110 llama-server.service, .15 strix-server.service) — .8 was probing a stale
|
||||
llama-server unit that reads inactive, producing false UNREACHABLE legs.
|
||||
@@ -44,6 +45,11 @@ Changelog:
|
||||
separate from `failures`. Every run prints absolute execution provenance
|
||||
(script + cwd) in the header, in the cron ALERT line, and in --json output so
|
||||
a stale-consumer report is distinguishable from a fault at read time.
|
||||
v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed
|
||||
from the AGENTS dict when she moved off this host onto her own container
|
||||
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
|
||||
from her side. This script must not probe mumuni or .24 — the v2 changelog
|
||||
roster line was the last reference still placing her at .24 / CT100.
|
||||
"""
|
||||
|
||||
import subprocess, json, sys, os, time, re, io, contextlib
|
||||
|
||||
Executable
+227
@@ -0,0 +1,227 @@
|
||||
#!/usr/bin/env bash
|
||||
# capture-dsh-token.sh — refresh the dsh-web login token WITHOUT restarting dsh-web.
|
||||
#
|
||||
# Context (CT 112 / tankodhs.sysloggh.net)
|
||||
# ----------------------------------------
|
||||
# The dsh-web UI (systemd unit `dsh-web.service`, 127.0.0.1:3080) prints a random
|
||||
# launch token to the journal on every start:
|
||||
#
|
||||
# dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||||
#
|
||||
# That token is the only way to bootstrap the authority-bound 30-day browser
|
||||
# cookie. It rotates on every dsh-web start, so the Authentik-gated
|
||||
# `location = /dsh-web-login` in /etc/nginx/sites-available/dsh must always
|
||||
# reference the token of the RUNNING process.
|
||||
#
|
||||
# This script:
|
||||
# 1. selects the launch token the RUNNING service actually accepts from the
|
||||
# current systemd invocation — it NEVER stops or starts dsh-web,
|
||||
# 2. records it in /etc/dsh-web/launch-token,
|
||||
# 3. regenerates the nginx include /etc/dsh-web/nginx-login.conf (the
|
||||
# `proxy_pass ...?token=` line consumed by /dsh-web-login),
|
||||
# 4. reloads nginx ONLY when the on-disk include differs from the generated
|
||||
# one or the applied-state stamp does not match the token (the stamp is
|
||||
# written only after a successful reload), rolling the include back on
|
||||
# failure so the next run retries,
|
||||
# 5. removes the legacy unauthenticated :8081 endpoint if it ever reappears.
|
||||
#
|
||||
# Idempotent and safe to run at any time (systemd ExecStartPost or timer).
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
|
||||
|
||||
JOURNAL_UNIT="dsh-web.service"
|
||||
TOKEN_FILE="/etc/dsh-web/launch-token"
|
||||
INCLUDE_FILE="/etc/dsh-web/nginx-login.conf"
|
||||
STAMP_FILE="/etc/dsh-web/nginx-login.conf.applied"
|
||||
PENDING_FILE="/etc/dsh-web/nginx-reload.pending"
|
||||
SITE_ENABLED="/etc/nginx/sites-enabled/dsh"
|
||||
LEGACY_8081="/etc/nginx/sites-enabled/dsh.token"
|
||||
STASH_DIR="/etc/nginx/sites-available"
|
||||
LOCK_FILE="/run/capture-dsh-token.lock"
|
||||
LOGIN_HOST="tankodhs.sysloggh.net"
|
||||
LOGIN_UPSTREAM="http://127.0.0.1:3080"
|
||||
TOKEN_WAIT=120
|
||||
|
||||
log() { printf 'capture-dsh-token: %s\n' "$*" >&2; }
|
||||
die() { printf 'capture-dsh-token: ERROR: %s\n' "$*" >&2; exit 1; }
|
||||
|
||||
[ "$(id -u)" -eq 0 ] || die "must run as root"
|
||||
|
||||
# ── 0. Serialize runs so timer/ExecStartPost/manual runs cannot interleave ──
|
||||
exec 9>"$LOCK_FILE"
|
||||
flock -n 9 || { log "another capture-dsh-token run holds $LOCK_FILE; exiting"; exit 0; }
|
||||
mkdir -p "$(dirname "$PENDING_FILE")"
|
||||
|
||||
# ── 0b. Guarantee the generated include exists before any `nginx -t` ──────
|
||||
# The :80 site includes /etc/dsh-web/nginx-login.conf by literal path, so a
|
||||
# missing include makes every `nginx -t` fail and can wedge recovery. Seed it
|
||||
# from the last known token (or a placeholder); step 4 replaces it.
|
||||
if [ ! -f "$INCLUDE_FILE" ]; then
|
||||
SEED="placeholder"
|
||||
if [ -f "$TOKEN_FILE" ]; then
|
||||
SEED="$(cat "$TOKEN_FILE" 2>/dev/null || true)"
|
||||
[ -n "$SEED" ] || SEED="placeholder"
|
||||
fi
|
||||
printf '%s' "$SEED" | grep -qE '^[A-Za-z0-9._~+/=:@-]+$' || SEED="placeholder"
|
||||
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$SEED" > "$INCLUDE_FILE"
|
||||
chmod 600 "$INCLUDE_FILE"
|
||||
log "created missing $INCLUDE_FILE"
|
||||
fi
|
||||
|
||||
# ── 1. Remove the legacy unauthenticated :8081 endpoint, if present ─────────
|
||||
# It bypassed Authentik entirely (listened on 0.0.0.0:8081 with no auth_request)
|
||||
# and must never come back. Stash it rather than delete so it is auditable.
|
||||
if [ -e "$LEGACY_8081" ] || [ -L "$LEGACY_8081" ]; then
|
||||
TS="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
STASHED="$STASH_DIR/dsh.token.disabled-$TS"
|
||||
mv "$LEGACY_8081" "$STASHED"
|
||||
chmod 600 "$STASHED" 2>/dev/null || true
|
||||
touch "$PENDING_FILE"
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
die "nginx config test failed after disabling $LEGACY_8081 (kept disabled at $STASHED): $NGINX_TEST_OUT; a pending reload is recorded so running nginx is reloaded once the config is fixed. The legacy :8081 endpoint will NOT be restored."
|
||||
fi
|
||||
if ! nginx -s reload; then
|
||||
die "nginx reload failed after disabling $LEGACY_8081 (kept disabled at $STASHED); a pending reload is recorded so running nginx is reloaded on the next run. The legacy :8081 endpoint will NOT be restored."
|
||||
fi
|
||||
rm -f "$PENDING_FILE"
|
||||
log "removed legacy :8081 endpoint -> $STASHED"
|
||||
fi
|
||||
|
||||
# ── 1b. Honor a recorded pending reload regardless of token selection ───────
|
||||
# A failed reload leaves PENDING_FILE set so a stashed legacy :8081 file can
|
||||
# never remain loaded in the running nginx while dsh-web is down or not yet
|
||||
# answering. Reconcile it before the token wait.
|
||||
if [ -e "$PENDING_FILE" ]; then
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
log "WARNING: pending nginx reload recorded but 'nginx -t' fails: $NGINX_TEST_OUT; continuing so the include can be regenerated; will retry next run"
|
||||
elif ! nginx -s reload; then
|
||||
log "WARNING: pending nginx reload recorded but 'nginx -s reload' failed; will retry next run"
|
||||
else
|
||||
rm -f "$PENDING_FILE"
|
||||
log "completed pending nginx reload"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── 2. Select the token the RUNNING service actually accepts ────────────────
|
||||
# Re-sample the service's CURRENT systemd invocation on every pass and read
|
||||
# candidates only from it, so a restart that lands during the wait immediately
|
||||
# switches to the new invocation; there is no whole-journal or cross-invocation
|
||||
# fallback, and an empty/unknown invocation just waits. Each candidate is then
|
||||
# functionally verified against the local dsh-web using the public authority,
|
||||
# exactly as the /dsh-web-login proxy does, and the first that answers 303 is
|
||||
# the live token. Candidates are re-probed newest-first on each pass (connection
|
||||
# failures stay eligible) until one is accepted or the wait elapses.
|
||||
journal_tokens() {
|
||||
journalctl -u "$JOURNAL_UNIT" "_SYSTEMD_INVOCATION_ID=$1" --no-pager -o cat 2>/dev/null \
|
||||
| grep -oE 'dsh web: https?://[^[:space:]]+[?&]token=[^[:space:]]+' \
|
||||
| sed -E 's/.*[?&]token=//' \
|
||||
| grep -E '^[A-Za-z0-9._~+/=:@-]+$' \
|
||||
| tac | awk '!seen[$0]++' || true
|
||||
}
|
||||
|
||||
TOKEN=""
|
||||
DEADLINE=$((SECONDS + TOKEN_WAIT))
|
||||
NO_INVOCATION_WARNED=0
|
||||
while [ -z "$TOKEN" ] && [ "$SECONDS" -lt "$DEADLINE" ]; do
|
||||
INVOCATION="$(systemctl show -p InvocationID --value "$JOURNAL_UNIT" 2>/dev/null || true)"
|
||||
if [ -z "$INVOCATION" ] || [ "$INVOCATION" = "n/a" ]; then
|
||||
if [ "$NO_INVOCATION_WARNED" -eq 0 ]; then
|
||||
log "WARNING: no invocation id for $JOURNAL_UNIT; waiting for a live invocation"
|
||||
NO_INVOCATION_WARNED=1
|
||||
fi
|
||||
sleep 2
|
||||
continue
|
||||
fi
|
||||
for cand in $(journal_tokens "$INVOCATION"); do
|
||||
code="$(curl -s -o /dev/null --max-time 5 -w '%{http_code}' \
|
||||
-H "Host: $LOGIN_HOST" "$LOGIN_UPSTREAM/?token=$cand" || true)"
|
||||
if [ "$code" = "303" ]; then
|
||||
TOKEN="$cand"
|
||||
break
|
||||
fi
|
||||
done
|
||||
[ -n "$TOKEN" ] && break
|
||||
sleep 2
|
||||
done
|
||||
|
||||
if [ -z "$TOKEN" ]; then
|
||||
log "no accepted launch token in the current invocation within ${TOKEN_WAIT}s; leaving the include untouched for the next run"
|
||||
[ -e "$PENDING_FILE" ] && die "pending nginx reload could not be completed; will retry next run"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ── 3. Record the token (atomic, private) ──────────────────────────────────
|
||||
mkdir -p "$(dirname "$TOKEN_FILE")"
|
||||
if ! printf '%s\n' "$TOKEN" | cmp -s - "$TOKEN_FILE" 2>/dev/null; then
|
||||
printf '%s\n' "$TOKEN" > "$TOKEN_FILE.tmp"
|
||||
chmod 600 "$TOKEN_FILE.tmp"
|
||||
mv "$TOKEN_FILE.tmp" "$TOKEN_FILE"
|
||||
log "recorded live launch token in $TOKEN_FILE"
|
||||
fi
|
||||
chmod 600 "$TOKEN_FILE"
|
||||
|
||||
# ── 4. Regenerate the nginx login include (reload only when it changes) ────
|
||||
NEW_INCLUDE="$(mktemp "$INCLUDE_FILE.XXXXXX")"
|
||||
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$TOKEN" > "$NEW_INCLUDE"
|
||||
chmod 600 "$NEW_INCLUDE"
|
||||
|
||||
# The stamp records the token nginx actually loaded. It is written only after a
|
||||
# successful reload, so the early exit is safe only when both the stamp and the
|
||||
# on-disk include agree with the live token; anything else falls through to the
|
||||
# reload path so the include can never silently diverge from what nginx serves.
|
||||
APPLIED=""
|
||||
[ -f "$STAMP_FILE" ] && APPLIED="$(cat "$STAMP_FILE" 2>/dev/null || true)"
|
||||
[ -f "$INCLUDE_FILE" ] && chmod 600 "$INCLUDE_FILE"
|
||||
[ -f "$STAMP_FILE" ] && chmod 600 "$STAMP_FILE"
|
||||
|
||||
if [ "$APPLIED" = "$TOKEN" ] && [ -f "$INCLUDE_FILE" ] && cmp -s "$NEW_INCLUDE" "$INCLUDE_FILE" \
|
||||
&& [ ! -e "$PENDING_FILE" ]; then
|
||||
rm -f "$NEW_INCLUDE"
|
||||
log "token unchanged; nginx not reloaded"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[ -e "$SITE_ENABLED" ] || { rm -f "$NEW_INCLUDE"; die "$SITE_ENABLED missing; refusing to reload"; }
|
||||
|
||||
RESTORE=""
|
||||
if [ -f "$INCLUDE_FILE" ]; then
|
||||
RESTORE="$(mktemp "$INCLUDE_FILE.bak.XXXXXX")"
|
||||
cp -p "$INCLUDE_FILE" "$RESTORE"
|
||||
chmod 600 "$RESTORE"
|
||||
fi
|
||||
|
||||
mv "$NEW_INCLUDE" "$INCLUDE_FILE"
|
||||
chmod 600 "$INCLUDE_FILE"
|
||||
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
if [ -n "$RESTORE" ]; then
|
||||
mv "$RESTORE" "$INCLUDE_FILE"
|
||||
else
|
||||
rm -f "$INCLUDE_FILE"
|
||||
fi
|
||||
die "nginx config test failed: $NGINX_TEST_OUT; previous include restored"
|
||||
fi
|
||||
|
||||
if ! nginx -s reload; then
|
||||
if [ -n "$RESTORE" ]; then
|
||||
mv "$RESTORE" "$INCLUDE_FILE"
|
||||
else
|
||||
rm -f "$INCLUDE_FILE"
|
||||
fi
|
||||
touch "$PENDING_FILE"
|
||||
die "nginx reload failed; previous include restored; a pending reload is recorded so the next run retries"
|
||||
fi
|
||||
|
||||
if [ -n "$RESTORE" ]; then
|
||||
rm -f "$RESTORE"
|
||||
fi
|
||||
|
||||
rm -f "$PENDING_FILE"
|
||||
|
||||
printf '%s\n' "$TOKEN" > "$STAMP_FILE.tmp"
|
||||
chmod 600 "$STAMP_FILE.tmp"
|
||||
mv "$STAMP_FILE.tmp" "$STAMP_FILE"
|
||||
|
||||
log "token changed; nginx reloaded"
|
||||
log "login endpoint: https://$LOGIN_HOST/dsh-web-login (Authentik-gated)"
|
||||
@@ -15,7 +15,7 @@ import smtplib, json, subprocess, os, sys, datetime, re
|
||||
from email.mime.text import MIMEText
|
||||
from email.mime.multipart import MIMEMultipart
|
||||
|
||||
PVE = "https://minipve.sysloggh.net"
|
||||
PVE = "https://192.168.68.12:8006"
|
||||
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
|
||||
# ── Shared credentials —─
|
||||
@@ -38,10 +38,16 @@ TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
|
||||
# ── Helpers ──
|
||||
|
||||
def pve_get(path):
|
||||
cmd = f'curl -sfk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list)."""
|
||||
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||
try:
|
||||
return json.loads(subprocess.check_output(cmd, shell=True))["data"]
|
||||
except: return []
|
||||
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
|
||||
if r.returncode != 0:
|
||||
return None
|
||||
data = json.loads(r.stdout)
|
||||
return data.get("data", [])
|
||||
except:
|
||||
return None
|
||||
|
||||
def ssh(host, cmd):
|
||||
try:
|
||||
@@ -97,6 +103,12 @@ def collect():
|
||||
|
||||
# ── Proxmox Nodes ──
|
||||
nodes = pve_get("/api2/json/nodes")
|
||||
if nodes is None:
|
||||
report["nodes"] = {}
|
||||
report["node_count"] = 0
|
||||
report["nodes_online"] = 0
|
||||
report["pve_probe_status"] = "unreachable"
|
||||
else:
|
||||
report["nodes"] = {n["node"]: {
|
||||
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
|
||||
@@ -108,10 +120,16 @@ def collect():
|
||||
} for n in nodes}
|
||||
report["node_count"] = len(nodes)
|
||||
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
|
||||
report["pve_probe_status"] = "ok"
|
||||
|
||||
# ── VMs/CTs ──
|
||||
resources = pve_get("/api2/json/cluster/resources")
|
||||
if resources is None:
|
||||
vms = []
|
||||
report["resources_probe_status"] = "unreachable"
|
||||
else:
|
||||
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||
report["resources_probe_status"] = "ok"
|
||||
report["total_vms"] = len(vms)
|
||||
report["running_vms"] = sum(1 for v in vms if v.get("status") == "running")
|
||||
stopped = [v for v in vms if v.get("status") != "running"]
|
||||
@@ -183,7 +201,7 @@ def collect():
|
||||
("Authentik", "https://auth.sysloggh.net"),
|
||||
("Zulip", "https://chat.sysloggh.net"),
|
||||
("Pulse", "https://pulse.sysloggh.net"),
|
||||
("Proxmox", "https://minipve.sysloggh.net"),
|
||||
("Proxmox", "https://192.168.68.12:8006"),
|
||||
("SearXNG", "http://192.168.68.7:8888"),
|
||||
("Firecrawl", "http://192.168.68.7:3002/health"),
|
||||
]
|
||||
@@ -247,11 +265,13 @@ def collect():
|
||||
zulip_health = json.loads(health_body) if health_body else {}
|
||||
except:
|
||||
zulip_health = {}
|
||||
report["zulip_ext"]["connected"] = zulip_health.get("connected", False)
|
||||
report["zulip_ext"]["queue_id"] = zulip_health.get("queue_id")
|
||||
report["zulip_ext"]["last_error"] = zulip_health.get("last_error")
|
||||
report["zulip_ext"]["messages_processed"] = zulip_health.get("messages_processed", 0)
|
||||
report["zulip_ext"]["retry_count"] = zulip_health.get("retry_count", 0)
|
||||
# Live state is nested under 'zulip' key
|
||||
zulip_state = zulip_health.get("zulip", {})
|
||||
report["zulip_ext"]["connected"] = zulip_state.get("connected", False)
|
||||
report["zulip_ext"]["queue_id"] = zulip_state.get("queue_id")
|
||||
report["zulip_ext"]["last_error"] = zulip_state.get("last_error")
|
||||
report["zulip_ext"]["messages_processed"] = zulip_state.get("messages_processed", 0)
|
||||
report["zulip_ext"]["skipped"] = zulip_state.get("skipped", 0)
|
||||
|
||||
# Phase 2: PM2 process check
|
||||
pm2_raw = subprocess.check_output(
|
||||
@@ -289,8 +309,8 @@ def collect():
|
||||
# Abiba (pi)
|
||||
report["agents"]["abiba"] = {
|
||||
"platform": "pi", "ct": 100, "ip": "192.168.68.24",
|
||||
"zulip_connected": zulip_health.get("connected", False),
|
||||
"zulip_processed": zulip_health.get("messages_processed", 0),
|
||||
"zulip_connected": zulip_state.get("connected", False),
|
||||
"zulip_processed": zulip_state.get("messages_processed", 0),
|
||||
"pm2_status": pm2.get("status", "unknown"),
|
||||
"pm2_restarts": pm2.get("restarts", "?"),
|
||||
"pm2_uptime": pm2.get("uptime", "?"),
|
||||
@@ -308,26 +328,13 @@ def collect():
|
||||
"updated_at": "",
|
||||
}
|
||||
|
||||
# Mumuni (CT 100, IP 192.168.68.24)
|
||||
mumuni_state = ssh("192.168.68.24", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
||||
mumuni_data = {}
|
||||
try:
|
||||
mumuni_data = json.loads(mumuni_state) if mumuni_state else {}
|
||||
except:
|
||||
mumuni_data = {}
|
||||
mumuni_platforms = mumuni_data.get("platforms", {})
|
||||
report["agents"]["mumuni"] = {
|
||||
"platform": "hermes", "ct": 100, "ip": "192.168.68.24",
|
||||
"gateway_state": mumuni_data.get("gateway_state", "unknown"),
|
||||
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
|
||||
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
|
||||
"email_state": mumuni_platforms.get("email", {}).get("state", "unknown"),
|
||||
"hermes_version": "",
|
||||
}
|
||||
# Get Hermes version
|
||||
ver = ssh("192.168.68.24", "hermes --version 2>/dev/null | head -1")
|
||||
if ver:
|
||||
report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip()
|
||||
# Mumuni is deliberately absent from this digest: captain ruling 2026-09-10.
|
||||
# She moved off this host onto her own container (kagentz CT 105 on minipve,
|
||||
# 192.168.68.14, dedicated `hermes` user) and is monitored from her side. The
|
||||
# former probe ssh'd to 192.168.68.24 for the decommissioned deployment's
|
||||
# ~/.hermes/gateway_state.json, always read "unknown", and published a false
|
||||
# "mumuni:unknown" line in the agent table and the gateway-unknown issue
|
||||
# count of every digest. Do NOT re-add an .24 / gateway_state probe.
|
||||
|
||||
return report
|
||||
|
||||
@@ -419,7 +426,7 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
<div class="alert {'good' if not issues else 'bad' if any('🔴' in i for i in issues) else 'warn'}">
|
||||
<p style="margin:0;font-size:16px"><b>{status}</b></p>
|
||||
<p style="margin:4px 0 0 0;font-size:13px">
|
||||
{r['node_count']} PVE nodes · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
||||
Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']}) · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
||||
{r['docker_vm']['total'] + r['docker_syslog']['total'] + r['docker_netbird']['total']} containers ·
|
||||
{len(r['endpoints'])} endpoints · {len(r.get('agents',{}))} agents
|
||||
</p>
|
||||
@@ -435,8 +442,10 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
|
||||
# ── Quick Stats ──
|
||||
html += '<div class="card"><h2>📊 Quick Stats</h2><div class="grid">'
|
||||
pve_status_label = "unreachable" if r.get('pve_probe_status') == 'unreachable' else f"{r['nodes_online']}/{r['node_count']}"
|
||||
pve_status_color = "red" if r.get('pve_probe_status') == 'unreachable' or r['nodes_online'] != r['node_count'] else "green"
|
||||
stats = [
|
||||
("PVE Nodes", f"{r['nodes_online']}/{r['node_count']}", "green" if r['nodes_online'] == r['node_count'] else "red"),
|
||||
("PVE Nodes", pve_status_label, pve_status_color),
|
||||
("VMs/CTs", f"{r['running_vms']}/{r['total_vms']}", "green" if r['running_vms'] == r['total_vms'] else "red"),
|
||||
("Containers", f"{r['docker_vm']['running']}/{r['docker_vm']['total']}", "green" if r['docker_vm']['running'] == r['docker_vm']['total'] else "yellow"),
|
||||
("LiteLLM Ctrs", f"{r['docker_syslog']['running']}/{r['docker_syslog']['total']}", "green" if r['docker_syslog']['running'] == r['docker_syslog']['total'] else "red"),
|
||||
@@ -539,12 +548,6 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
processed = "DSH"
|
||||
elif name == "mumuni":
|
||||
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
tg = "✅" if agent.get("telegram_state") == "connected" else "❌"
|
||||
ver = agent.get("hermes_version", "")
|
||||
processed = f"TG:{tg} v{ver}"
|
||||
else:
|
||||
zulip_state = "⬜"
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
|
||||
@@ -5,6 +5,7 @@
|
||||
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||
|
||||
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
||||
LOG="/root/pm2-self-heal.log"
|
||||
TELEGRAM_CHAT_ID="5822977936"
|
||||
|
||||
notify_tg() {
|
||||
@@ -16,8 +17,6 @@ notify_tg() {
|
||||
-d "text=${msg}" \
|
||||
-d "parse_mode=HTML" > /dev/null 2>&1 || true
|
||||
}
|
||||
ALERTS="${ALERTS}$msg"
|
||||
}
|
||||
|
||||
# Log-only mode: replaced by prose contract pm2-self-heal.prose.md
|
||||
# Only alerts Telegram on actual failure (status != online)
|
||||
|
||||
+14
-19
@@ -1,7 +1,10 @@
|
||||
#!/bin/bash
|
||||
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
|
||||
# Implements zulip-health.prose.md v2
|
||||
# Implements zulip-health.prose.md v3
|
||||
# Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'.
|
||||
# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B
|
||||
# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes
|
||||
# agent leg is retired — see the note after the Tanko leg.
|
||||
set -euo pipefail
|
||||
|
||||
ZULIP_SITE="https://chat.sysloggh.net"
|
||||
@@ -21,7 +24,8 @@ notify() {
|
||||
|
||||
# Zulip DM to owner
|
||||
local content="${severity} Zulip Monitor: ${msg}"
|
||||
local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
@@ -141,23 +145,14 @@ else
|
||||
esac
|
||||
fi
|
||||
|
||||
# ── Platform B: Hermes (Mumuni) ──
|
||||
MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \
|
||||
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
|
||||
MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c "
|
||||
import sys,json
|
||||
d=json.load(sys.stdin)
|
||||
p=d.get('platforms',{}).get('zulip',{})
|
||||
print(p.get('state','unknown'))
|
||||
" 2>/dev/null)
|
||||
|
||||
if [ "$MUMUNI_ZULIP" != "connected" ]; then
|
||||
notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG"
|
||||
else
|
||||
echo " Mumuni: ✅ Zulip connected" >> "$LOG"
|
||||
fi
|
||||
# ── Removed: the former "Platform B: Hermes" agent leg ──
|
||||
# Captain ruling 2026-09-10: that agent moved off this host onto her own
|
||||
# container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now
|
||||
# monitored on her side — see the out-of-scope note in zulip-health.prose.md.
|
||||
# The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway
|
||||
# state, which reported "unknown" on every run and posted a false 🔴 DM plus an
|
||||
# #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must
|
||||
# never contact her former host.
|
||||
|
||||
# ── Platform C: Agent Zero (kagentz) ──
|
||||
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
"""
|
||||
Regression tests for daily-infra-report.py fixes (PR #64).
|
||||
|
||||
Tests:
|
||||
(a) Asserts the nested zulip read feeds the agent-card fields
|
||||
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
|
||||
"""
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch, MagicMock
|
||||
|
||||
# Add scripts to path
|
||||
sys.path.insert(0, str(Path(__file__).parent.parent / "scripts"))
|
||||
import importlib.util
|
||||
|
||||
def load_script():
|
||||
"""Load the daily-infra-report script as a module."""
|
||||
script_path = Path(__file__).parent.parent / "scripts" / "daily-infra-report.py"
|
||||
spec = importlib.util.spec_from_file_location("daily_infra_report", script_path)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
def test_nested_zulip_read_feeds_agent_card():
|
||||
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
|
||||
# Mock the http_get_body response with nested structure
|
||||
mock_health_response = json.dumps({
|
||||
"status": "ok",
|
||||
"platform": "pi",
|
||||
"agent": "abiba",
|
||||
"zulip": {
|
||||
"connected": True,
|
||||
"queue_id": "test-queue-id",
|
||||
"messages_processed": 42,
|
||||
"skipped": 5,
|
||||
"last_error": None
|
||||
}
|
||||
})
|
||||
|
||||
# Import and patch
|
||||
report_mod = load_script()
|
||||
|
||||
with patch.object(report_mod, 'http_get_body', return_value=mock_health_response):
|
||||
# Simulate the collect() function's Zulip section
|
||||
zulip_health = json.loads(report_mod.http_get_body("http://localhost:9200/health"))
|
||||
zulip_state = zulip_health.get("zulip", {})
|
||||
|
||||
# Assert the nested key is read correctly
|
||||
assert zulip_state.get("connected") == True, "Zulip connected should be True from nested key"
|
||||
assert zulip_state.get("messages_processed") == 42, "messages_processed should be 42 from nested key"
|
||||
assert zulip_state.get("queue_id") == "test-queue-id", "queue_id should be read from nested key"
|
||||
|
||||
# Simulate the agent card field population
|
||||
agent_card = {
|
||||
"zulip_connected": zulip_state.get("connected", False),
|
||||
"zulip_processed": zulip_state.get("messages_processed", 0),
|
||||
}
|
||||
|
||||
assert agent_card["zulip_connected"] == True, "Agent card should show Zulip connected"
|
||||
assert agent_card["zulip_processed"] == 42, "Agent card should show 42 processed messages"
|
||||
|
||||
|
||||
def test_unreachable_pve_get_renders_labelled_unreachable():
|
||||
"""Test that an unreachable PVE API renders 'unreachable' instead of '0/0'."""
|
||||
# Import and patch
|
||||
report_mod = load_script()
|
||||
|
||||
# Test pve_get returns None on error
|
||||
with patch.object(report_mod.subprocess, 'run') as mock_run:
|
||||
mock_run.return_value.returncode = 7 # Connection failure
|
||||
result = report_mod.pve_get("/api2/json/nodes")
|
||||
assert result is None, "pve_get should return None on connection failure"
|
||||
|
||||
# Test the render logic
|
||||
report = {
|
||||
"nodes": {},
|
||||
"node_count": 0,
|
||||
"nodes_online": 0,
|
||||
"pve_probe_status": "unreachable",
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
}
|
||||
|
||||
# The render should show "unreachable" not "0/0"
|
||||
pve_status_label = "unreachable" if report.get('pve_probe_status') == 'unreachable' else f"{report['nodes_online']}/{report['node_count']}"
|
||||
|
||||
assert pve_status_label == "unreachable", "PVE status should show 'unreachable' when probe fails, not '0/0'"
|
||||
|
||||
|
||||
def test_unreachable_resources_renders_labelled_unreachable():
|
||||
"""Test that unreachable resources probe renders 'unreachable' instead of '0/0'."""
|
||||
report_mod = load_script()
|
||||
|
||||
# Test resources probe returns None
|
||||
with patch.object(report_mod.subprocess, 'run') as mock_run:
|
||||
mock_run.return_value.returncode = 7
|
||||
result = report_mod.pve_get("/api2/json/cluster/resources")
|
||||
assert result is None, "pve_get for resources should return None on connection failure"
|
||||
|
||||
# Test the render logic
|
||||
report = {
|
||||
"resources_probe_status": "unreachable",
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
}
|
||||
|
||||
resources_label = "unreachable" if report.get('resources_probe_status') == 'unreachable' else f"{report['running_vms']}/{report['total_vms']}"
|
||||
|
||||
assert resources_label == "unreachable", "Resources status should show 'unreachable' when probe fails, not '0/0'"
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print("Running tests...")
|
||||
try:
|
||||
test_nested_zulip_read_feeds_agent_card()
|
||||
print("✓ test_nested_zulip_read_feeds_agent_card passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_nested_zulip_read_feeds_agent_card failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_unreachable_pve_get_renders_labelled_unreachable()
|
||||
print("✓ test_unreachable_pve_get_renders_labelled_unreachable passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_unreachable_pve_get_renders_labelled_unreachable failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_unreachable_resources_renders_labelled_unreachable()
|
||||
print("✓ test_unreachable_resources_renders_labelled_unreachable passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
print("All tests passed!")
|
||||
@@ -0,0 +1,352 @@
|
||||
"""Regression tests for the 2026-09-10 retirement of the Mumuni monitoring leg.
|
||||
|
||||
WHY THIS FILE EXISTS: captain ruling 2026-09-10 — Mumuni moved off this host
|
||||
onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated
|
||||
`hermes` user) and is monitored from her side. The monitor nevertheless kept
|
||||
ssh'ing to root@192.168.68.24 for `~/.hermes/gateway_state.json` on the
|
||||
decommissioned deployment, read "unknown" on every run, and posted a false 🔴
|
||||
"Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert to the
|
||||
captain. The daily infra digest published a matching `mumuni:unknown` row.
|
||||
|
||||
CONTRACT UNDER TEST:
|
||||
* `scripts/zulip-monitor.sh` carries NO Mumuni probe and NO 192.168.68.24
|
||||
reference; it never ssh'es .24, and even on a failing run it emits no Mumuni
|
||||
notify (stdout alert, Zulip payload, or log line).
|
||||
* The Abiba (pi — the Zulip bridge), Tanko (DSH) and Agent Zero (kagentz) legs
|
||||
still work: deleting the Mumuni leg must not have gutted the rest.
|
||||
* `scripts/daily-infra-report.py` no longer probes .24 for a Hermes gateway
|
||||
state and no longer emits a `mumuni` agent entry.
|
||||
* `scripts/agent-health-check.py`'s AGENTS roster has no mumuni entry. This is
|
||||
a pin, not a behavior change — verify the probe was already gone.
|
||||
* `zulip-health.prose.md` retires the Mumuni-only steps and says explicitly
|
||||
that Mumuni is not monitored from this host.
|
||||
|
||||
HOW: behavioral execution plus one named deliverable-text contract. The sandbox
|
||||
copies the shipped monitor verbatim and rewrites only its LOG constant, then
|
||||
runs it with stub ssh/curl on PATH; the ssh stub records every host it is asked
|
||||
to reach, so "never probes .24" and "no Mumuni notify" are asserted from
|
||||
observed behavior. The daily digest is pinned by importing it and exercising
|
||||
collect() and build_html() directly. The single source-text assertion is the
|
||||
deliverable-text contract the captain acceptance names for the shipped monitor.
|
||||
|
||||
Usage: python3 -m pytest tests/test_mumuni_monitor_removal.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
import os
|
||||
import pathlib
|
||||
import stat
|
||||
import subprocess
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||
DAILY_REPORT = ROOT / "scripts" / "daily-infra-report.py"
|
||||
AHC = ROOT / "scripts" / "agent-health-check.py"
|
||||
HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
|
||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||
|
||||
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
|
||||
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||
|
||||
|
||||
# ── scripts/zulip-monitor.sh: deliverable-text contract ─────────────
|
||||
|
||||
def test_zulip_monitor_deliverable_text_contract():
|
||||
"""Owned deliverable-text contract for scripts/zulip-monitor.sh.
|
||||
|
||||
Captain acceptance requires the shipped monitor to contain no Mumuni probe
|
||||
identifier and no 192.168.68.24 literal. Behavioral proof that the monitor
|
||||
never contacts that host and never emits a Mumuni notify lives in the
|
||||
sandbox tests below; this only pins the named text contract.
|
||||
"""
|
||||
text = ZULIP_MONITOR.read_text()
|
||||
assert "mumuni" not in text.lower()
|
||||
assert MUMUNI_IP not in text
|
||||
|
||||
|
||||
# ── scripts/zulip-monitor.sh: behavioral sandbox ─────────────────────
|
||||
|
||||
SSH_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub ssh: record the target host, then answer by host + remote command.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||
host=""
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
*@192.168.*) host="${a##*@}" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||
cmd="${*: -1}"
|
||||
case "$host" in
|
||||
192.168.68.15)
|
||||
case "$cmd" in
|
||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||
esac ;;
|
||||
192.168.68.14)
|
||||
case "$cmd" in
|
||||
*agent.json*) printf '%s' "$AZ_A2A" ;;
|
||||
*"ps aux"*) printf '%s\n' "$AZ_PS" ;;
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
|
||||
# record every call (including notify) payloads.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
case " $* " in
|
||||
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
path.write_text(body)
|
||||
path.chmod(path.stat().st_mode
|
||||
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a='{"name":"kagentz"}',
|
||||
az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
Everything else — legs, labels, notify logic — is the shipped script.
|
||||
"""
|
||||
sandbox = tmp_path / "sandbox"
|
||||
bindir = sandbox / "bin"
|
||||
record = sandbox / "record"
|
||||
bindir.mkdir(parents=True)
|
||||
record.mkdir()
|
||||
|
||||
_write_exec(bindir / "ssh", SSH_STUB)
|
||||
_write_exec(bindir / "curl", CURL_STUB)
|
||||
|
||||
source = ZULIP_MONITOR.read_text()
|
||||
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||
log_path = sandbox / "zulip-health-monitor.log"
|
||||
script = sandbox / "zulip-monitor.sh"
|
||||
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||
|
||||
env = dict(os.environ)
|
||||
env.update({
|
||||
"PATH": f"{bindir}:{env['PATH']}",
|
||||
"RECORD_DIR": str(record),
|
||||
"TANKO_SVC": tanko_svc,
|
||||
"TANKO_HTTP": tanko_http,
|
||||
"AZ_A2A": az_a2a,
|
||||
"AZ_PS": az_ps,
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
return proc, record, log_path
|
||||
|
||||
|
||||
def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
||||
proc, record, log_path = _run_monitor(tmp_path)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# Every retained leg actually ran and passed.
|
||||
assert "Server: ✅ HTTP 200" in log
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "Tanko: ✅ service=active http=200" in log
|
||||
assert "kagentz: ✅ A2A alive" in log
|
||||
assert "kagentz: ✅ Adapter running" in log
|
||||
assert "Result: ✅ All healthy" in log
|
||||
|
||||
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||
assert proc.stdout == ""
|
||||
assert "Mumuni" not in log
|
||||
assert "🔴" not in log
|
||||
|
||||
# Observed behavior: .24 is never resolved, only Tanko's vantage and the
|
||||
# Agent Zero host are contacted.
|
||||
hosts = record.joinpath("ssh.hosts").read_text().split()
|
||||
assert MUMUNI_IP not in hosts
|
||||
assert set(hosts) == {TANKO_VANTAGE, AGENT_ZERO_HOST}
|
||||
assert not record.joinpath("unexpected-ssh").exists()
|
||||
|
||||
|
||||
def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
||||
# Failure path: exercises notify() end to end so "no Mumuni notify" is
|
||||
# proven on the alert path, not only on the quiet healthy path.
|
||||
proc, record, log_path = _run_monitor(tmp_path, tanko_svc="inactive",
|
||||
tanko_http="000")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
|
||||
alerts = proc.stdout
|
||||
assert "Tanko (DSH dsh-web) service state: inactive" in alerts
|
||||
assert "1 issue(s) found" in alerts
|
||||
|
||||
# No Mumuni text in stdout, the log, or any Zulip DM/stream payload.
|
||||
assert "Mumuni" not in alerts
|
||||
assert "Mumuni" not in log_path.read_text()
|
||||
assert MUMUNI_IP not in alerts + log_path.read_text()
|
||||
payloads = record.joinpath("curl.calls").read_text()
|
||||
assert "Mumuni" not in payloads
|
||||
assert MUMUNI_IP not in payloads
|
||||
|
||||
# The rest of the monitor still ran alongside the failing Tanko leg.
|
||||
log = log_path.read_text()
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "kagentz: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
|
||||
|
||||
# ── scripts/daily-infra-report.py: behavioral digest checks ──────────
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def daily():
|
||||
spec = importlib.util.spec_from_file_location("daily_infra_report", DAILY_REPORT)
|
||||
assert spec and spec.loader
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
DAILY_AGENTS = {
|
||||
"abiba": {
|
||||
"platform": "pi", "ct": 100, "ip": MUMUNI_IP,
|
||||
"zulip_connected": True, "zulip_processed": 5,
|
||||
"pm2_status": "online", "pm2_restarts": "0", "pm2_uptime": "1h",
|
||||
},
|
||||
"tanko": {
|
||||
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
|
||||
"gateway_state": "n/a (DSH)", "zulip_state": "connected",
|
||||
"telegram_state": "unknown", "gateway_pid": None, "updated_at": "",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def _fabricated_report(agents):
|
||||
return {
|
||||
"nodes": {},
|
||||
"node_count": 1,
|
||||
"nodes_online": 1,
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
"stopped_vms": [],
|
||||
"vms_by_node": {n: [] for n in
|
||||
["amdpve", "minipve", "storepve", "acerpve", "ocupve"]},
|
||||
"storage": [],
|
||||
"docker_vm": {"total": 0, "running": 0, "unhealthy": [],
|
||||
"containers": [], "reclaimable": "", "disk_used": "1%"},
|
||||
"docker_syslog": {"total": 0, "running": 0, "containers": []},
|
||||
"docker_netbird": {"total": 0, "running": 0, "containers": []},
|
||||
"endpoints": [],
|
||||
"litellm": {"checks": []},
|
||||
"nfs": [],
|
||||
"zulip_ext": {
|
||||
"connected": True, "queue_id": "queue", "last_error": None,
|
||||
"messages_processed": 0, "retry_count": 0, "pm2": {},
|
||||
"pm2_healthy": True, "bot_skipped_15min": 0, "finalized_1h": 0,
|
||||
"failed_finalize_1h": 0, "finalize_fail_pct": 0,
|
||||
"server_status": "200",
|
||||
},
|
||||
"agents": agents,
|
||||
}
|
||||
|
||||
|
||||
def _agent_status_card(html):
|
||||
start = html.index("🤖 Agent Status")
|
||||
end = html.index("💬 Zulip Extension")
|
||||
return html[start:end]
|
||||
|
||||
|
||||
def test_daily_report_renders_only_abiba_and_tanko_agents(daily):
|
||||
"""build_html() over a Mumuni-free agent set must render no Mumuni row and
|
||||
no Mumuni gateway-unknown issue, while abiba and tanko rows still render."""
|
||||
html = daily.build_html(_fabricated_report(dict(DAILY_AGENTS)))
|
||||
card = _agent_status_card(html)
|
||||
assert "mumuni" not in card.lower()
|
||||
assert "abiba" in card
|
||||
assert "tanko" in card
|
||||
assert "mumuni" not in html.lower()
|
||||
|
||||
|
||||
def test_daily_report_collect_never_probes_mumuni(monkeypatch, daily):
|
||||
"""collect() with ssh stubbed must add no mumuni agent and must never ssh
|
||||
its decommissioned .24 host."""
|
||||
probed = []
|
||||
|
||||
class _NoSubprocess:
|
||||
@staticmethod
|
||||
def check_output(*args, **kwargs):
|
||||
return b""
|
||||
|
||||
def fake_ssh(host, cmd):
|
||||
probed.append(host)
|
||||
return ""
|
||||
|
||||
monkeypatch.setattr(daily, "pve_get", lambda path: [])
|
||||
monkeypatch.setattr(daily, "ssh_jerome", lambda host, cmd: "")
|
||||
monkeypatch.setattr(daily, "ssh", fake_ssh)
|
||||
monkeypatch.setattr(daily, "http_get",
|
||||
lambda url, auth=None, timeout=10: "200")
|
||||
monkeypatch.setattr(daily, "http_get_body",
|
||||
lambda url, auth=None, timeout=10: "")
|
||||
monkeypatch.setattr(daily, "count_in_log", lambda *a, **k: 0)
|
||||
monkeypatch.setattr(daily, "subprocess", _NoSubprocess)
|
||||
|
||||
report = daily.collect()
|
||||
assert "mumuni" not in report["agents"]
|
||||
assert MUMUNI_IP not in probed
|
||||
|
||||
|
||||
# ── scripts/agent-health-check.py: roster pin ───────────────────────
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def ahc():
|
||||
spec = importlib.util.spec_from_file_location("agent_health_check_roster", AHC)
|
||||
assert spec and spec.loader
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
def test_agent_health_roster_has_no_mumuni_entry(ahc):
|
||||
assert "mumuni" not in ahc.AGENTS
|
||||
|
||||
|
||||
# ── zulip-health.prose.md: contract reconciliation ──────────────────
|
||||
|
||||
def test_health_contract_retires_mumuni_only_steps():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
assert MUMUNI_IP not in text
|
||||
for step in ("**B4:", "**B5:", "**B6:"):
|
||||
assert step not in text
|
||||
|
||||
|
||||
def test_health_contract_states_mumuni_is_not_monitored_from_this_host():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
assert "Mumuni is NOT monitored from this host" in text
|
||||
assert "monitored on her side" in text
|
||||
assert "her own container" in text
|
||||
|
||||
|
||||
def test_health_contract_keeps_tanko_agent_zero_and_bridge_steps():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
for marker in ("**B1:", "**B2:", "**B3:", "Step 4: Platform C",
|
||||
"Step 2: Platform A", "Step 1: Zulip Server Liveness"):
|
||||
assert marker in text, marker
|
||||
+197
-43
@@ -1,9 +1,9 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-health
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||
version: 3.0.0
|
||||
version: 3.2.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
report_only_agents:
|
||||
@@ -12,13 +12,23 @@ report_only_agents:
|
||||
|
||||
# Zulip Mesh Health Monitor
|
||||
|
||||
Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero).
|
||||
Runs every 15 minutes in the background. Also triggers on session start.
|
||||
Monitors the Zulip-connected agents under this host's operational control (pi,
|
||||
DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on
|
||||
session start.
|
||||
|
||||
> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She
|
||||
> moved off this host onto her own container — kagentz CT 105 on minipve
|
||||
> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her
|
||||
> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may
|
||||
> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on
|
||||
> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat,
|
||||
> response delivery) are retired: they always read "unknown" against the
|
||||
> decommissioned deployment and produced a false 🔴 alert on every run.
|
||||
|
||||
## Requires
|
||||
|
||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); Mumuni (192.168.68.14, kagentz CT105 on minipve); and Agent Zero Docker host (192.168.68.14)
|
||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||
- **PM2** on localhost for pi process management
|
||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||
@@ -81,13 +91,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
|
||||
## Streaming Support (2026-07-05)
|
||||
|
||||
Zulip agents now support progressive message editing during agent generation.
|
||||
When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is
|
||||
When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is
|
||||
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
||||
|
||||
- Adapter implements `edit_message()` using `_api_patch()` helper
|
||||
- Gateway stream consumer progressively edits the Zulip message
|
||||
- User sees real-time agent thinking instead of waiting for full response
|
||||
- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active
|
||||
- Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here
|
||||
|
||||
### Verification
|
||||
```bash
|
||||
@@ -183,7 +193,10 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
|
||||
| Crash loop >10/h | Alert user |
|
||||
|
||||
|
||||
### Step 3: Platform B — Tanko (DSH on amdpve CT 112) & Mumuni (Hermes)
|
||||
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
|
||||
|
||||
Mumuni is out of scope for this host (see the note above): she runs on her own
|
||||
container and is monitored on her side.
|
||||
|
||||
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
|
||||
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
|
||||
@@ -240,44 +253,185 @@ timeout only. Never expect a bare `200` — the public URL terminates in the
|
||||
token-gated authentik chain. Statuses outside the healthy set are
|
||||
logged/reported as a warning — reported, never healed on.
|
||||
|
||||
**B4: Gateway Process** (Hermes agent Mumuni only — Tanko runs no Hermes gateway)
|
||||
|
||||
```bash
|
||||
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
|
||||
```
|
||||
|
||||
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more
|
||||
than one `gateway run` process is found, the gateway has a collision (typically
|
||||
one `--force` and one `--replace` process). Kill the newer/duplicate process,
|
||||
then restart the remaining gateway per-agent (parameterized 2026-08-09, captain
|
||||
ruling). Check the gateway log for "Gateway running with 2 platform(s)" (not 1)
|
||||
to confirm Zulip reloaded.
|
||||
|
||||
**B5: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
||||
```
|
||||
|
||||
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
|
||||
Silence > 300s → warning. Silence > 600s → critical.
|
||||
|
||||
**B6: Response Delivery** (Hermes agent Mumuni only)
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
||||
```
|
||||
|
||||
> 50% fail rate → critical.
|
||||
|
||||
**Platform B Actions**
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service |
|
||||
| No heartbeat in 10min | Same as above |
|
||||
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
||||
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
||||
| `dsh-web` service not `active` | Restart Tanko via DSH service |
|
||||
| HTTP `:3080` connection refused/timeout (`000`) | Same as above |
|
||||
| HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
|
||||
|
||||
**B4: dsh-web Authentication (Tanko — restart-persistent login)**
|
||||
|
||||
The dsh-web UI is token-gated. On every start the process prints a random
|
||||
launch token to the journal:
|
||||
|
||||
```
|
||||
dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||||
```
|
||||
|
||||
The token only bootstraps an authority-bound, HMAC-signed browser cookie with a
|
||||
30-day lifetime. The signing secret is durable in
|
||||
`/root/.dsh/.credentials.yaml` (key `client-connection/browser-session`), so a
|
||||
cookie minted once keeps working across `dsh-web` restarts; the launch token
|
||||
itself rotates on every restart.
|
||||
|
||||
**Login endpoint (public, Authentik-gated):**
|
||||
`https://tankodhs.sysloggh.net/dsh-web-login`
|
||||
|
||||
It lives inside the Authentik-gated `:80` server block
|
||||
(`/etc/nginx/sites-available/dsh`, symlinked from
|
||||
`/etc/nginx/sites-enabled/dsh`) as `location = /dsh-web-login`, guarded by
|
||||
`auth_request /outpost.goauthentik.io/auth/nginx`. It proxies to dsh-web with
|
||||
`Host: tankodhs.sysloggh.net`, so the minted cookie is bound to the public
|
||||
authority — never to `127.0.0.1:3080`. The token-dependent line is isolated in
|
||||
the generated include `/etc/dsh-web/nginx-login.conf`:
|
||||
|
||||
```
|
||||
proxy_pass http://127.0.0.1:3080/?token=<TOKEN>;
|
||||
```
|
||||
|
||||
**Token refresh (non-disruptive):**
|
||||
`/opt/deepseek-harness/capture-dsh-token.sh` (source:
|
||||
`scripts/capture-dsh-token.sh`) reads candidate launch tokens from the journal
|
||||
**scoped to the service's current systemd invocation**
|
||||
(`systemctl show -p InvocationID` + `_SYSTEMD_INVOCATION_ID=`), re-sampling the
|
||||
invocation on every pass so a restart that lands during the wait switches to the
|
||||
new invocation; a restarted process's stale token is never considered while its
|
||||
new startup banner is still pending and there is no whole-journal or
|
||||
cross-invocation fallback. Each candidate
|
||||
is then functionally verified against dsh-web with `Host: tankodhs.sysloggh.net`,
|
||||
using the first the running process accepts with `303`. It waits up to 120s for
|
||||
a restarted process to accept a token and re-probes every current-invocation
|
||||
candidate on each pass, so a token that briefly returns `000` while the service
|
||||
is still starting is not disqualified. If none is accepted it leaves the include
|
||||
untouched and exits so the timer retries (exiting non-zero when a pending reload
|
||||
is still outstanding). It writes
|
||||
`/etc/dsh-web/launch-token` and regenerates `/etc/dsh-web/nginx-login.conf`,
|
||||
reloading nginx only when the on-disk include differs from the generated one or
|
||||
the applied-state stamp does not match the token (`nginx -t` guards the reload,
|
||||
and the stamp is written only after a successful `nginx -s reload`, so a failed
|
||||
or interrupted reload is retried on the next run). Any failed reload records a
|
||||
pending-reload marker under `/etc/dsh-web/`; the next run attempts the reload
|
||||
before the token wait, independent of token state, and clears the marker only
|
||||
once the reload succeeds, so a disabled legacy `:8081` file can never leave the
|
||||
running nginx unreloaded. The generated include is recreated before any
|
||||
`nginx -t` if it is missing, so a failed run cannot wedge recovery.
|
||||
Runs are serialized with `flock` on `/run/capture-dsh-token.lock`. It **never
|
||||
stops or starts `dsh-web`**.
|
||||
It is triggered by the `dsh-web.service` drop-in
|
||||
`/etc/systemd/system/dsh-web.service.d/20-token-refresh.conf`
|
||||
(`ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service`) and by
|
||||
`dsh-web-token.timer` every 2 minutes for reconciliation.
|
||||
|
||||
<details><summary>Installed systemd wiring (CT 112)</summary>
|
||||
|
||||
```ini
|
||||
# /etc/systemd/system/dsh-web-token.service
|
||||
[Unit]
|
||||
Description=Refresh the dsh-web launch token for the nginx login endpoint
|
||||
After=dsh-web.service
|
||||
[Service]
|
||||
Type=oneshot
|
||||
TimeoutStartSec=180
|
||||
ExecStart=/opt/deepseek-harness/capture-dsh-token.sh
|
||||
|
||||
# /etc/systemd/system/dsh-web-token.timer
|
||||
[Unit]
|
||||
Description=Periodically refresh the dsh-web login token
|
||||
[Timer]
|
||||
OnBootSec=90s
|
||||
OnUnitActiveSec=120s
|
||||
AccuracySec=10s
|
||||
Persistent=true
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
|
||||
# /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf
|
||||
[Service]
|
||||
ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
> **Do NOT reintroduce the `:8081` endpoint.** It listened on `0.0.0.0:8081`
|
||||
> with no `auth_request` and was a full Authentik bypass for anyone on the LAN.
|
||||
> The script now removes `/etc/nginx/sites-enabled/dsh.token` automatically if
|
||||
> it ever reappears.
|
||||
|
||||
**Authentication flow:**
|
||||
1. `GET https://tankodhs.sysloggh.net/dsh-web-login`
|
||||
2. Unauthenticated → Authentik sign-in; once authenticated the request reaches
|
||||
dsh-web with `Host: tankodhs.sysloggh.net`.
|
||||
3. dsh-web accepts the launch token on `GET /`, writes the
|
||||
`dsh-auth-<authority-hash>` cookie (30 days, `HttpOnly`, `SameSite=Strict`)
|
||||
and returns `303` to `/`.
|
||||
4. Every later request through `/` presents that cookie; the token is not needed
|
||||
again until the cookie expires or a new browser is used.
|
||||
|
||||
**Verification** (amdpve vantage):
|
||||
```bash
|
||||
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
|
||||
# Expected: 302
|
||||
|
||||
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
||||
-w '%{http_code}\n' http://192.168.68.122:8081/"
|
||||
# Expected: 000
|
||||
|
||||
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
|
||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the minted dsh-auth-... cookie (authority
|
||||
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
|
||||
|
||||
# 4. Token refresh is non-disruptive and idempotent.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
||||
# Expected: "token unchanged; nginx not reloaded" when nothing changed
|
||||
```
|
||||
|
||||
**Restart durability (acceptance):** after `systemctl restart dsh-web`, (a) the
|
||||
cookie minted before the restart still returns `200` on `/`, and (b) the
|
||||
refreshed `/etc/dsh-web/nginx-login.conf` carries the new token and mints a
|
||||
fresh cookie. Both verified live 2026-09-11.
|
||||
|
||||
```bash
|
||||
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
|
||||
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
|
||||
# until the socket answers (any status but 000) before asserting the cookie.
|
||||
for i in $(seq 1 60); do
|
||||
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
|
||||
[ "$UP" != "000" ] && break
|
||||
sleep 2
|
||||
done
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the pre-restart cookie is still accepted.
|
||||
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
|
||||
# manual run may no-op on the flock, so poll until the include carries a token
|
||||
# the running process accepts (bounded wait) before the mint+reuse check.
|
||||
for i in $(seq 1 60); do
|
||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
||||
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
|
||||
[ "$CODE" = "303" ] && break
|
||||
sleep 2
|
||||
done
|
||||
# Expected: 303 — the include now holds the token the running process accepts.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the refreshed token minted a fresh cookie.
|
||||
```
|
||||
|
||||
|
||||
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
|
||||
|
||||
@@ -333,7 +487,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I
|
||||
|
||||
Check each agent's log for excessive bot-to-bot chatter:
|
||||
- Abiba: `Skipped.*bot msgs` count
|
||||
- Tanko/Mumuni: Repeated DM exchanges between bots
|
||||
- Tanko: Repeated DM exchanges between bots
|
||||
- kagentz: Adapter log for bot DMs being processed
|
||||
|
||||
If any bot processes >50 bot-originated messages in 15min → warning.
|
||||
|
||||
Reference in New Issue
Block a user