Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b001b657d4 | ||
|
|
a1ffeaad34 | ||
|
|
287657a77a | ||
|
|
21f9073e0b | ||
|
|
32fe7c0652 | ||
|
|
25cf2f5eef | ||
|
|
26f2301188 | ||
|
|
a6a459acc0 | ||
|
|
14bed6e916 | ||
|
|
2e70c834cb | ||
|
|
4f59b82404 | ||
|
|
8a4ee4cf5a | ||
|
|
952fca9c92 | ||
|
|
b1462f3e79 | ||
|
|
bfdff13ae7 | ||
|
|
8bf32f6f0f |
@@ -628,18 +628,18 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 3.0.0
|
||||
version: 3.1.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/15 * * * *'
|
||||
description: "Every 15 minutes \u2014 monitors all Zulip-connected agents"
|
||||
description: "Every 15 minutes \u2014 monitors the Zulip-connected agents under this host's control (pi, DSH, Agent Zero)"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 120
|
||||
requires:
|
||||
- Zulip API key for abiba-bot@chat.sysloggh.net
|
||||
- SSH access to all Hermes agents
|
||||
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
|
||||
verification:
|
||||
postconditions:
|
||||
- check: bot registration active
|
||||
|
||||
@@ -3,15 +3,16 @@ kind: responsibility
|
||||
name: infrastructure-update
|
||||
description: >
|
||||
Autonomous system-wide update contract covering all 5 Proxmox nodes,
|
||||
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
|
||||
15+ containers/VMs, and 5 Docker ecosystems (docker-vm .7, CT 116 .116,
|
||||
CT 117, hwpve .11, NetBird VPS 72.61.0.17). Updates apt packages,
|
||||
Docker images, and container stacks in safe waves with health checks
|
||||
and automatic rollback on failure.
|
||||
agent: abiba
|
||||
triggers:
|
||||
- on "infra update" command
|
||||
- weekly (Sunday 03:00 EDT) via cron
|
||||
- weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
|
||||
- on security advisory relay from Mumuni
|
||||
version: 1.2.0
|
||||
version: 1.3.0
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -80,7 +81,13 @@ Before ANY update wave:
|
||||
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
|
||||
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
|
||||
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
|
||||
| CT 117 (zulip, storepve) | Zulip | `docker pull zulip/docker-zulip:latest && docker restart zulip-zulip-1` |
|
||||
| CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) |
|
||||
| CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) |
|
||||
| hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` |
|
||||
| NetBird VPS (72.61.0.17) | NetBird (server, dashboard, proxy, traefik, crowdsec) | `ssh root@72.61.0.17 'cd /root && docker compose pull && docker compose up -d'` |
|
||||
| VM 109 (.7) | Trove test | `cd /opt/trove-test && docker compose pull && docker compose up -d` |
|
||||
| VM 109 (.7) | docker-stats | `cd /opt/docker-stats && docker compose pull && docker compose up -d` |
|
||||
| CT 116 (.116) | Monitoring (Grafana, Prometheus, Alertmanager, PVE exporter) | `cd /opt/monitoring && docker compose pull && docker compose up -d` |
|
||||
|
||||
**Verify after Wave 3:**
|
||||
- All containers healthy: `docker ps` on each host
|
||||
@@ -88,8 +95,12 @@ Before ANY update wave:
|
||||
- MCP integration test: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools (23 RA-H OS + 67 GitHub)
|
||||
- Zulip test: send test message to #agent-hub
|
||||
- Dashboard loading: `curl localhost:3001/` (via CT 116)
|
||||
- Firecrawl test: `curl :3002/`
|
||||
- Firecrawl test: `curl -X POST http://192.168.68.7:3002/v1/search -H 'Content-Type: application/json' -d '{"query":"health","limit":1}'` → `"success":true` (GET `/` returns 200)
|
||||
- Authentik test: `curl http://192.168.68.11:9000/` → 302 redirect to login
|
||||
- NetBird test: `curl -s -o /dev/null -w '%{http_code}' https://netbird.sysloggh.net/` → 200
|
||||
- harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
|
||||
- SearXNG test: `curl :8888`
|
||||
- Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
|
||||
|
||||
## Wave 4: Proxmox Kernel Reboot
|
||||
|
||||
@@ -158,6 +169,8 @@ Before Wave 1, snapshot these files:
|
||||
/opt/search-stack/searxng/docker-compose.yml (VM 109 .7)
|
||||
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
||||
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
||||
/root/compose.yml (hwpve .11 — Authentik server/worker/postgres)
|
||||
/root/docker-compose.yml (NetBird VPS — netbird server/dashboard/proxy, traefik, crowdsec)
|
||||
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
||||
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
|
||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||
@@ -220,7 +233,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
||||
|
||||
- [ ] All 5 PVE nodes updated, no reboot-loop
|
||||
- [ ] All VMs/CTs running post-update
|
||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
|
||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117 + hwpve .11 + NetBird VPS)
|
||||
- [ ] LiteLLM inference passing (syslog-auto test)
|
||||
- [ ] Zulip server + all 3 agents connected
|
||||
- [ ] GPU fleet at full capacity (3/3)
|
||||
|
||||
@@ -113,7 +113,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||
|
||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||
|
||||
## Maintains
|
||||
|
||||
+28
-6
@@ -6,7 +6,7 @@ name: memory-fixer
|
||||
description: >
|
||||
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
|
||||
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
|
||||
version: 2.0.0
|
||||
version: 2.1.0
|
||||
---
|
||||
---
|
||||
|
||||
@@ -69,13 +69,15 @@ FROM nodes
|
||||
WHERE json_extract(metadata, '$.namespace') IS NULL;
|
||||
```
|
||||
|
||||
### 3. Staleness Review Tagging
|
||||
### 3. Staleness Review Tagging (refresh-suggested nodes only)
|
||||
|
||||
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
|
||||
|
||||
**Archive-suggested nodes are NO LONGER tagged — they are archived outright (see Level 1 fix 4).** Tagging with `[REVIEW: refresh]` applies only to living nodes (infrastructure, deployment, system, system-health, business, philosophy, research, learning, investigation, analysis, project, agent, registry, policy).
|
||||
|
||||
**Exclusion Rules:**
|
||||
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
|
||||
- Nodes whose `description` already starts with `[REVIEW:` are NOT re-processed
|
||||
- Nodes whose `description` already starts with `[REVIEW:` or `[ARCHIVED]` are NOT re-processed
|
||||
|
||||
```sql
|
||||
SELECT id, title, json_extract(metadata, '$.type') as node_type,
|
||||
@@ -104,9 +106,28 @@ LIMIT 10;
|
||||
|
||||
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`.
|
||||
|
||||
### 4. Stale-Node Archiving (Level 1 — standing Kwame directive, 2026-09-11)
|
||||
|
||||
**Kwame's standing directive: stale nodes CAN be archived by the fixer. No per-batch escalation, no `[REVIEW: archive]` tagging — archive them.**
|
||||
|
||||
For every node whose suggested action is `archive` (i.e. its type is NOT one of the living types in fix 3), archive it in a **single** `updateNode` call:
|
||||
|
||||
```python
|
||||
updateNode(id, {
|
||||
"description": "[ARCHIVED] " + originalDescriptionWithoutReviewTag,
|
||||
"metadata": {"state": "archived"}
|
||||
})
|
||||
```
|
||||
|
||||
- `state` transitions **DO work through `updateNode`** (`archived`, and back to `active`). The former "state only accepts processed/not_processed, use SSH" claim was wrong — verified 2026-09-11 by archiving 7 nodes (#61, #373, #388, #465, #475, #526, #1476) over the bridge with `updated_at` auto-bumping. **SSH to the bridge host is a fallback, not a requirement**, and it is blocked from kagentz anyway.
|
||||
- Pass `description` and `metadata` in the **same** call, and always keep the `updates` object nested: `{"id": N, "updates": {…}}`.
|
||||
- Archiving is non-destructive: the node stays in the graph, marked `state: archived` + `[ARCHIVED] ` prefix. **Living nodes (refresh-suggested) are NEVER archived** without a specific Kwame decision — they are the cluster/agent/business canon.
|
||||
|
||||
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
|
||||
|
||||
## Level 2 Escalations (Kwame Decision Required)
|
||||
|
||||
1. **Stale nodes** flagged with `[REVIEW: …]` — Archive, refresh, or keep?
|
||||
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
|
||||
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
|
||||
3. **Orphan Nodes >90 days old** — Archive or connect?
|
||||
|
||||
@@ -144,7 +165,7 @@ Reply with:
|
||||
|
||||
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
|
||||
|
||||
> ⚠️ `updateNode` cannot set `state` to non-standard values (restricted to `processed`/`not_processed`) and cannot add metadata keys. For state transitions and `updated_at` bumps, use **direct SSH + SQLite** on the bridge host:
|
||||
> ⚠️ **Corrected 2026-09-11:** `updateNode` DOES accept `state` changes — `{"updates": {"description": …, "metadata": {"state": "archived"}}}` works over the bridge, and `updated_at` bumps automatically. The old "use direct SSH + SQLite for state transitions" instruction was based on a wrong assumption; SSH is a fallback only (and is blocked from kagentz). Use one `updateNode` call for both the tag and the state.
|
||||
> ```bash
|
||||
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
|
||||
> ```
|
||||
@@ -172,8 +193,9 @@ The result must be 0 rows when all decisions are executed. Report what was done.
|
||||
## Checks
|
||||
|
||||
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
|
||||
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
|
||||
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
|
||||
- **Timestamps:** every executed decision bumps `updated_at`, so the node exits the stale window on the next run.
|
||||
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
|
||||
|
||||
## Logging
|
||||
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
|
||||
|
||||
@@ -17,8 +17,9 @@ Changelog:
|
||||
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
|
||||
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
|
||||
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
|
||||
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
|
||||
abiba (.24).
|
||||
Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24).
|
||||
(v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she
|
||||
moved to her own container, kagentz CT 105 / .14, and is monitored there.)
|
||||
v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service,
|
||||
.110 llama-server.service, .15 strix-server.service) — .8 was probing a stale
|
||||
llama-server unit that reads inactive, producing false UNREACHABLE legs.
|
||||
@@ -44,6 +45,11 @@ Changelog:
|
||||
separate from `failures`. Every run prints absolute execution provenance
|
||||
(script + cwd) in the header, in the cron ALERT line, and in --json output so
|
||||
a stale-consumer report is distinguishable from a fault at read time.
|
||||
v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed
|
||||
from the AGENTS dict when she moved off this host onto her own container
|
||||
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
|
||||
from her side. This script must not probe mumuni or .24 — the v2 changelog
|
||||
roster line was the last reference still placing her at .24 / CT100.
|
||||
"""
|
||||
|
||||
import subprocess, json, sys, os, time, re, io, contextlib
|
||||
|
||||
@@ -15,7 +15,7 @@ import smtplib, json, subprocess, os, sys, datetime, re
|
||||
from email.mime.text import MIMEText
|
||||
from email.mime.multipart import MIMEMultipart
|
||||
|
||||
PVE = "https://minipve.sysloggh.net"
|
||||
PVE = "https://192.168.68.12:8006"
|
||||
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
|
||||
# ── Shared credentials —─
|
||||
@@ -38,10 +38,16 @@ TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
|
||||
# ── Helpers ──
|
||||
|
||||
def pve_get(path):
|
||||
cmd = f'curl -sfk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list)."""
|
||||
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
|
||||
try:
|
||||
return json.loads(subprocess.check_output(cmd, shell=True))["data"]
|
||||
except: return []
|
||||
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
|
||||
if r.returncode != 0:
|
||||
return None
|
||||
data = json.loads(r.stdout)
|
||||
return data.get("data", [])
|
||||
except:
|
||||
return None
|
||||
|
||||
def ssh(host, cmd):
|
||||
try:
|
||||
@@ -97,21 +103,33 @@ def collect():
|
||||
|
||||
# ── Proxmox Nodes ──
|
||||
nodes = pve_get("/api2/json/nodes")
|
||||
report["nodes"] = {n["node"]: {
|
||||
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
|
||||
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
|
||||
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
|
||||
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
|
||||
"uptime_h": n.get('uptime',0)//3600,
|
||||
"status": n["status"]
|
||||
} for n in nodes}
|
||||
report["node_count"] = len(nodes)
|
||||
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
|
||||
if nodes is None:
|
||||
report["nodes"] = {}
|
||||
report["node_count"] = 0
|
||||
report["nodes_online"] = 0
|
||||
report["pve_probe_status"] = "unreachable"
|
||||
else:
|
||||
report["nodes"] = {n["node"]: {
|
||||
"cpu_pct": round(n.get('cpu',0)*100, 1),
|
||||
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
|
||||
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
|
||||
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
|
||||
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
|
||||
"uptime_h": n.get('uptime',0)//3600,
|
||||
"status": n["status"]
|
||||
} for n in nodes}
|
||||
report["node_count"] = len(nodes)
|
||||
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
|
||||
report["pve_probe_status"] = "ok"
|
||||
|
||||
# ── VMs/CTs ──
|
||||
resources = pve_get("/api2/json/cluster/resources")
|
||||
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||
if resources is None:
|
||||
vms = []
|
||||
report["resources_probe_status"] = "unreachable"
|
||||
else:
|
||||
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
|
||||
report["resources_probe_status"] = "ok"
|
||||
report["total_vms"] = len(vms)
|
||||
report["running_vms"] = sum(1 for v in vms if v.get("status") == "running")
|
||||
stopped = [v for v in vms if v.get("status") != "running"]
|
||||
@@ -183,7 +201,7 @@ def collect():
|
||||
("Authentik", "https://auth.sysloggh.net"),
|
||||
("Zulip", "https://chat.sysloggh.net"),
|
||||
("Pulse", "https://pulse.sysloggh.net"),
|
||||
("Proxmox", "https://minipve.sysloggh.net"),
|
||||
("Proxmox", "https://192.168.68.12:8006"),
|
||||
("SearXNG", "http://192.168.68.7:8888"),
|
||||
("Firecrawl", "http://192.168.68.7:3002/health"),
|
||||
]
|
||||
@@ -247,11 +265,13 @@ def collect():
|
||||
zulip_health = json.loads(health_body) if health_body else {}
|
||||
except:
|
||||
zulip_health = {}
|
||||
report["zulip_ext"]["connected"] = zulip_health.get("connected", False)
|
||||
report["zulip_ext"]["queue_id"] = zulip_health.get("queue_id")
|
||||
report["zulip_ext"]["last_error"] = zulip_health.get("last_error")
|
||||
report["zulip_ext"]["messages_processed"] = zulip_health.get("messages_processed", 0)
|
||||
report["zulip_ext"]["retry_count"] = zulip_health.get("retry_count", 0)
|
||||
# Live state is nested under 'zulip' key
|
||||
zulip_state = zulip_health.get("zulip", {})
|
||||
report["zulip_ext"]["connected"] = zulip_state.get("connected", False)
|
||||
report["zulip_ext"]["queue_id"] = zulip_state.get("queue_id")
|
||||
report["zulip_ext"]["last_error"] = zulip_state.get("last_error")
|
||||
report["zulip_ext"]["messages_processed"] = zulip_state.get("messages_processed", 0)
|
||||
report["zulip_ext"]["skipped"] = zulip_state.get("skipped", 0)
|
||||
|
||||
# Phase 2: PM2 process check
|
||||
pm2_raw = subprocess.check_output(
|
||||
@@ -289,8 +309,8 @@ def collect():
|
||||
# Abiba (pi)
|
||||
report["agents"]["abiba"] = {
|
||||
"platform": "pi", "ct": 100, "ip": "192.168.68.24",
|
||||
"zulip_connected": zulip_health.get("connected", False),
|
||||
"zulip_processed": zulip_health.get("messages_processed", 0),
|
||||
"zulip_connected": zulip_state.get("connected", False),
|
||||
"zulip_processed": zulip_state.get("messages_processed", 0),
|
||||
"pm2_status": pm2.get("status", "unknown"),
|
||||
"pm2_restarts": pm2.get("restarts", "?"),
|
||||
"pm2_uptime": pm2.get("uptime", "?"),
|
||||
@@ -308,26 +328,13 @@ def collect():
|
||||
"updated_at": "",
|
||||
}
|
||||
|
||||
# Mumuni (CT 100, IP 192.168.68.24)
|
||||
mumuni_state = ssh("192.168.68.24", "cat ~/.hermes/gateway_state.json 2>/dev/null")
|
||||
mumuni_data = {}
|
||||
try:
|
||||
mumuni_data = json.loads(mumuni_state) if mumuni_state else {}
|
||||
except:
|
||||
mumuni_data = {}
|
||||
mumuni_platforms = mumuni_data.get("platforms", {})
|
||||
report["agents"]["mumuni"] = {
|
||||
"platform": "hermes", "ct": 100, "ip": "192.168.68.24",
|
||||
"gateway_state": mumuni_data.get("gateway_state", "unknown"),
|
||||
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
|
||||
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
|
||||
"email_state": mumuni_platforms.get("email", {}).get("state", "unknown"),
|
||||
"hermes_version": "",
|
||||
}
|
||||
# Get Hermes version
|
||||
ver = ssh("192.168.68.24", "hermes --version 2>/dev/null | head -1")
|
||||
if ver:
|
||||
report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip()
|
||||
# Mumuni is deliberately absent from this digest: captain ruling 2026-09-10.
|
||||
# She moved off this host onto her own container (kagentz CT 105 on minipve,
|
||||
# 192.168.68.14, dedicated `hermes` user) and is monitored from her side. The
|
||||
# former probe ssh'd to 192.168.68.24 for the decommissioned deployment's
|
||||
# ~/.hermes/gateway_state.json, always read "unknown", and published a false
|
||||
# "mumuni:unknown" line in the agent table and the gateway-unknown issue
|
||||
# count of every digest. Do NOT re-add an .24 / gateway_state probe.
|
||||
|
||||
return report
|
||||
|
||||
@@ -419,7 +426,7 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
<div class="alert {'good' if not issues else 'bad' if any('🔴' in i for i in issues) else 'warn'}">
|
||||
<p style="margin:0;font-size:16px"><b>{status}</b></p>
|
||||
<p style="margin:4px 0 0 0;font-size:13px">
|
||||
{r['node_count']} PVE nodes · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
||||
Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']}) · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
|
||||
{r['docker_vm']['total'] + r['docker_syslog']['total'] + r['docker_netbird']['total']} containers ·
|
||||
{len(r['endpoints'])} endpoints · {len(r.get('agents',{}))} agents
|
||||
</p>
|
||||
@@ -435,8 +442,10 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
|
||||
# ── Quick Stats ──
|
||||
html += '<div class="card"><h2>📊 Quick Stats</h2><div class="grid">'
|
||||
pve_status_label = "unreachable" if r.get('pve_probe_status') == 'unreachable' else f"{r['nodes_online']}/{r['node_count']}"
|
||||
pve_status_color = "red" if r.get('pve_probe_status') == 'unreachable' or r['nodes_online'] != r['node_count'] else "green"
|
||||
stats = [
|
||||
("PVE Nodes", f"{r['nodes_online']}/{r['node_count']}", "green" if r['nodes_online'] == r['node_count'] else "red"),
|
||||
("PVE Nodes", pve_status_label, pve_status_color),
|
||||
("VMs/CTs", f"{r['running_vms']}/{r['total_vms']}", "green" if r['running_vms'] == r['total_vms'] else "red"),
|
||||
("Containers", f"{r['docker_vm']['running']}/{r['docker_vm']['total']}", "green" if r['docker_vm']['running'] == r['docker_vm']['total'] else "yellow"),
|
||||
("LiteLLM Ctrs", f"{r['docker_syslog']['running']}/{r['docker_syslog']['total']}", "green" if r['docker_syslog']['running'] == r['docker_syslog']['total'] else "red"),
|
||||
@@ -539,12 +548,6 @@ th {{ color: #8b949e; font-weight: normal; }}
|
||||
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
processed = "DSH"
|
||||
elif name == "mumuni":
|
||||
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
tg = "✅" if agent.get("telegram_state") == "connected" else "❌"
|
||||
ver = agent.get("hermes_version", "")
|
||||
processed = f"TG:{tg} v{ver}"
|
||||
else:
|
||||
zulip_state = "⬜"
|
||||
gateway = agent.get("gateway_state", "?")
|
||||
@@ -698,6 +701,6 @@ if __name__ == "__main__":
|
||||
print(f" Zulip Ext: {'✅' if report.get('zulip_ext',{}).get('connected') else '❌'}")
|
||||
print(f" LiteLLM: {sum(1 for c in report.get('litellm',{}).get('checks',[]) if c['status']=='pass')}/{len(report.get('litellm',{}).get('checks',[]))} checks pass")
|
||||
agent_parts = []
|
||||
for k,v in report.get('agents',{}).items():
|
||||
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
||||
print(f" Agents: {', '.join(agent_parts)}")
|
||||
for k,v in report.get('agents',{}).items():
|
||||
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
|
||||
print(f" Agents: {', '.join(agent_parts)}")
|
||||
|
||||
@@ -5,6 +5,7 @@
|
||||
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||
|
||||
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
||||
LOG="/root/pm2-self-heal.log"
|
||||
TELEGRAM_CHAT_ID="5822977936"
|
||||
|
||||
notify_tg() {
|
||||
@@ -16,8 +17,6 @@ notify_tg() {
|
||||
-d "text=${msg}" \
|
||||
-d "parse_mode=HTML" > /dev/null 2>&1 || true
|
||||
}
|
||||
ALERTS="${ALERTS}$msg"
|
||||
}
|
||||
|
||||
# Log-only mode: replaced by prose contract pm2-self-heal.prose.md
|
||||
# Only alerts Telegram on actual failure (status != online)
|
||||
|
||||
+14
-19
@@ -1,7 +1,10 @@
|
||||
#!/bin/bash
|
||||
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
|
||||
# Implements zulip-health.prose.md v2
|
||||
# Implements zulip-health.prose.md v3
|
||||
# Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'.
|
||||
# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B
|
||||
# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes
|
||||
# agent leg is retired — see the note after the Tanko leg.
|
||||
set -euo pipefail
|
||||
|
||||
ZULIP_SITE="https://chat.sysloggh.net"
|
||||
@@ -21,7 +24,8 @@ notify() {
|
||||
|
||||
# Zulip DM to owner
|
||||
local content="${severity} Zulip Monitor: ${msg}"
|
||||
local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
@@ -141,23 +145,14 @@ else
|
||||
esac
|
||||
fi
|
||||
|
||||
# ── Platform B: Hermes (Mumuni) ──
|
||||
MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \
|
||||
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
|
||||
MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c "
|
||||
import sys,json
|
||||
d=json.load(sys.stdin)
|
||||
p=d.get('platforms',{}).get('zulip',{})
|
||||
print(p.get('state','unknown'))
|
||||
" 2>/dev/null)
|
||||
|
||||
if [ "$MUMUNI_ZULIP" != "connected" ]; then
|
||||
notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG"
|
||||
else
|
||||
echo " Mumuni: ✅ Zulip connected" >> "$LOG"
|
||||
fi
|
||||
# ── Removed: the former "Platform B: Hermes" agent leg ──
|
||||
# Captain ruling 2026-09-10: that agent moved off this host onto her own
|
||||
# container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now
|
||||
# monitored on her side — see the out-of-scope note in zulip-health.prose.md.
|
||||
# The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway
|
||||
# state, which reported "unknown" on every run and posted a false 🔴 DM plus an
|
||||
# #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must
|
||||
# never contact her former host.
|
||||
|
||||
# ── Platform C: Agent Zero (kagentz) ──
|
||||
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
"""
|
||||
Regression tests for daily-infra-report.py fixes (PR #64).
|
||||
|
||||
Tests:
|
||||
(a) Asserts the nested zulip read feeds the agent-card fields
|
||||
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
|
||||
"""
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch, MagicMock
|
||||
|
||||
# Add scripts to path
|
||||
sys.path.insert(0, str(Path(__file__).parent.parent / "scripts"))
|
||||
import importlib.util
|
||||
|
||||
def load_script():
|
||||
"""Load the daily-infra-report script as a module."""
|
||||
script_path = Path(__file__).parent.parent / "scripts" / "daily-infra-report.py"
|
||||
spec = importlib.util.spec_from_file_location("daily_infra_report", script_path)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
def test_nested_zulip_read_feeds_agent_card():
|
||||
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
|
||||
# Mock the http_get_body response with nested structure
|
||||
mock_health_response = json.dumps({
|
||||
"status": "ok",
|
||||
"platform": "pi",
|
||||
"agent": "abiba",
|
||||
"zulip": {
|
||||
"connected": True,
|
||||
"queue_id": "test-queue-id",
|
||||
"messages_processed": 42,
|
||||
"skipped": 5,
|
||||
"last_error": None
|
||||
}
|
||||
})
|
||||
|
||||
# Import and patch
|
||||
report_mod = load_script()
|
||||
|
||||
with patch.object(report_mod, 'http_get_body', return_value=mock_health_response):
|
||||
# Simulate the collect() function's Zulip section
|
||||
zulip_health = json.loads(report_mod.http_get_body("http://localhost:9200/health"))
|
||||
zulip_state = zulip_health.get("zulip", {})
|
||||
|
||||
# Assert the nested key is read correctly
|
||||
assert zulip_state.get("connected") == True, "Zulip connected should be True from nested key"
|
||||
assert zulip_state.get("messages_processed") == 42, "messages_processed should be 42 from nested key"
|
||||
assert zulip_state.get("queue_id") == "test-queue-id", "queue_id should be read from nested key"
|
||||
|
||||
# Simulate the agent card field population
|
||||
agent_card = {
|
||||
"zulip_connected": zulip_state.get("connected", False),
|
||||
"zulip_processed": zulip_state.get("messages_processed", 0),
|
||||
}
|
||||
|
||||
assert agent_card["zulip_connected"] == True, "Agent card should show Zulip connected"
|
||||
assert agent_card["zulip_processed"] == 42, "Agent card should show 42 processed messages"
|
||||
|
||||
|
||||
def test_unreachable_pve_get_renders_labelled_unreachable():
|
||||
"""Test that an unreachable PVE API renders 'unreachable' instead of '0/0'."""
|
||||
# Import and patch
|
||||
report_mod = load_script()
|
||||
|
||||
# Test pve_get returns None on error
|
||||
with patch.object(report_mod.subprocess, 'run') as mock_run:
|
||||
mock_run.return_value.returncode = 7 # Connection failure
|
||||
result = report_mod.pve_get("/api2/json/nodes")
|
||||
assert result is None, "pve_get should return None on connection failure"
|
||||
|
||||
# Test the render logic
|
||||
report = {
|
||||
"nodes": {},
|
||||
"node_count": 0,
|
||||
"nodes_online": 0,
|
||||
"pve_probe_status": "unreachable",
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
}
|
||||
|
||||
# The render should show "unreachable" not "0/0"
|
||||
pve_status_label = "unreachable" if report.get('pve_probe_status') == 'unreachable' else f"{report['nodes_online']}/{report['node_count']}"
|
||||
|
||||
assert pve_status_label == "unreachable", "PVE status should show 'unreachable' when probe fails, not '0/0'"
|
||||
|
||||
|
||||
def test_unreachable_resources_renders_labelled_unreachable():
|
||||
"""Test that unreachable resources probe renders 'unreachable' instead of '0/0'."""
|
||||
report_mod = load_script()
|
||||
|
||||
# Test resources probe returns None
|
||||
with patch.object(report_mod.subprocess, 'run') as mock_run:
|
||||
mock_run.return_value.returncode = 7
|
||||
result = report_mod.pve_get("/api2/json/cluster/resources")
|
||||
assert result is None, "pve_get for resources should return None on connection failure"
|
||||
|
||||
# Test the render logic
|
||||
report = {
|
||||
"resources_probe_status": "unreachable",
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
}
|
||||
|
||||
resources_label = "unreachable" if report.get('resources_probe_status') == 'unreachable' else f"{report['running_vms']}/{report['total_vms']}"
|
||||
|
||||
assert resources_label == "unreachable", "Resources status should show 'unreachable' when probe fails, not '0/0'"
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print("Running tests...")
|
||||
try:
|
||||
test_nested_zulip_read_feeds_agent_card()
|
||||
print("✓ test_nested_zulip_read_feeds_agent_card passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_nested_zulip_read_feeds_agent_card failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_unreachable_pve_get_renders_labelled_unreachable()
|
||||
print("✓ test_unreachable_pve_get_renders_labelled_unreachable passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_unreachable_pve_get_renders_labelled_unreachable failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
test_unreachable_resources_renders_labelled_unreachable()
|
||||
print("✓ test_unreachable_resources_renders_labelled_unreachable passed")
|
||||
except AssertionError as e:
|
||||
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
|
||||
sys.exit(1)
|
||||
|
||||
print("All tests passed!")
|
||||
@@ -0,0 +1,352 @@
|
||||
"""Regression tests for the 2026-09-10 retirement of the Mumuni monitoring leg.
|
||||
|
||||
WHY THIS FILE EXISTS: captain ruling 2026-09-10 — Mumuni moved off this host
|
||||
onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated
|
||||
`hermes` user) and is monitored from her side. The monitor nevertheless kept
|
||||
ssh'ing to root@192.168.68.24 for `~/.hermes/gateway_state.json` on the
|
||||
decommissioned deployment, read "unknown" on every run, and posted a false 🔴
|
||||
"Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert to the
|
||||
captain. The daily infra digest published a matching `mumuni:unknown` row.
|
||||
|
||||
CONTRACT UNDER TEST:
|
||||
* `scripts/zulip-monitor.sh` carries NO Mumuni probe and NO 192.168.68.24
|
||||
reference; it never ssh'es .24, and even on a failing run it emits no Mumuni
|
||||
notify (stdout alert, Zulip payload, or log line).
|
||||
* The Abiba (pi — the Zulip bridge), Tanko (DSH) and Agent Zero (kagentz) legs
|
||||
still work: deleting the Mumuni leg must not have gutted the rest.
|
||||
* `scripts/daily-infra-report.py` no longer probes .24 for a Hermes gateway
|
||||
state and no longer emits a `mumuni` agent entry.
|
||||
* `scripts/agent-health-check.py`'s AGENTS roster has no mumuni entry. This is
|
||||
a pin, not a behavior change — verify the probe was already gone.
|
||||
* `zulip-health.prose.md` retires the Mumuni-only steps and says explicitly
|
||||
that Mumuni is not monitored from this host.
|
||||
|
||||
HOW: behavioral execution plus one named deliverable-text contract. The sandbox
|
||||
copies the shipped monitor verbatim and rewrites only its LOG constant, then
|
||||
runs it with stub ssh/curl on PATH; the ssh stub records every host it is asked
|
||||
to reach, so "never probes .24" and "no Mumuni notify" are asserted from
|
||||
observed behavior. The daily digest is pinned by importing it and exercising
|
||||
collect() and build_html() directly. The single source-text assertion is the
|
||||
deliverable-text contract the captain acceptance names for the shipped monitor.
|
||||
|
||||
Usage: python3 -m pytest tests/test_mumuni_monitor_removal.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
import os
|
||||
import pathlib
|
||||
import stat
|
||||
import subprocess
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||
DAILY_REPORT = ROOT / "scripts" / "daily-infra-report.py"
|
||||
AHC = ROOT / "scripts" / "agent-health-check.py"
|
||||
HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
|
||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||
|
||||
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
|
||||
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||
|
||||
|
||||
# ── scripts/zulip-monitor.sh: deliverable-text contract ─────────────
|
||||
|
||||
def test_zulip_monitor_deliverable_text_contract():
|
||||
"""Owned deliverable-text contract for scripts/zulip-monitor.sh.
|
||||
|
||||
Captain acceptance requires the shipped monitor to contain no Mumuni probe
|
||||
identifier and no 192.168.68.24 literal. Behavioral proof that the monitor
|
||||
never contacts that host and never emits a Mumuni notify lives in the
|
||||
sandbox tests below; this only pins the named text contract.
|
||||
"""
|
||||
text = ZULIP_MONITOR.read_text()
|
||||
assert "mumuni" not in text.lower()
|
||||
assert MUMUNI_IP not in text
|
||||
|
||||
|
||||
# ── scripts/zulip-monitor.sh: behavioral sandbox ─────────────────────
|
||||
|
||||
SSH_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub ssh: record the target host, then answer by host + remote command.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||
host=""
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
*@192.168.*) host="${a##*@}" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||
cmd="${*: -1}"
|
||||
case "$host" in
|
||||
192.168.68.15)
|
||||
case "$cmd" in
|
||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||
esac ;;
|
||||
192.168.68.14)
|
||||
case "$cmd" in
|
||||
*agent.json*) printf '%s' "$AZ_A2A" ;;
|
||||
*"ps aux"*) printf '%s\n' "$AZ_PS" ;;
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
|
||||
# record every call (including notify) payloads.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
case " $* " in
|
||||
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
path.write_text(body)
|
||||
path.chmod(path.stat().st_mode
|
||||
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a='{"name":"kagentz"}',
|
||||
az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
Everything else — legs, labels, notify logic — is the shipped script.
|
||||
"""
|
||||
sandbox = tmp_path / "sandbox"
|
||||
bindir = sandbox / "bin"
|
||||
record = sandbox / "record"
|
||||
bindir.mkdir(parents=True)
|
||||
record.mkdir()
|
||||
|
||||
_write_exec(bindir / "ssh", SSH_STUB)
|
||||
_write_exec(bindir / "curl", CURL_STUB)
|
||||
|
||||
source = ZULIP_MONITOR.read_text()
|
||||
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||
log_path = sandbox / "zulip-health-monitor.log"
|
||||
script = sandbox / "zulip-monitor.sh"
|
||||
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||
|
||||
env = dict(os.environ)
|
||||
env.update({
|
||||
"PATH": f"{bindir}:{env['PATH']}",
|
||||
"RECORD_DIR": str(record),
|
||||
"TANKO_SVC": tanko_svc,
|
||||
"TANKO_HTTP": tanko_http,
|
||||
"AZ_A2A": az_a2a,
|
||||
"AZ_PS": az_ps,
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
return proc, record, log_path
|
||||
|
||||
|
||||
def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
||||
proc, record, log_path = _run_monitor(tmp_path)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# Every retained leg actually ran and passed.
|
||||
assert "Server: ✅ HTTP 200" in log
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "Tanko: ✅ service=active http=200" in log
|
||||
assert "kagentz: ✅ A2A alive" in log
|
||||
assert "kagentz: ✅ Adapter running" in log
|
||||
assert "Result: ✅ All healthy" in log
|
||||
|
||||
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||
assert proc.stdout == ""
|
||||
assert "Mumuni" not in log
|
||||
assert "🔴" not in log
|
||||
|
||||
# Observed behavior: .24 is never resolved, only Tanko's vantage and the
|
||||
# Agent Zero host are contacted.
|
||||
hosts = record.joinpath("ssh.hosts").read_text().split()
|
||||
assert MUMUNI_IP not in hosts
|
||||
assert set(hosts) == {TANKO_VANTAGE, AGENT_ZERO_HOST}
|
||||
assert not record.joinpath("unexpected-ssh").exists()
|
||||
|
||||
|
||||
def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
||||
# Failure path: exercises notify() end to end so "no Mumuni notify" is
|
||||
# proven on the alert path, not only on the quiet healthy path.
|
||||
proc, record, log_path = _run_monitor(tmp_path, tanko_svc="inactive",
|
||||
tanko_http="000")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
|
||||
alerts = proc.stdout
|
||||
assert "Tanko (DSH dsh-web) service state: inactive" in alerts
|
||||
assert "1 issue(s) found" in alerts
|
||||
|
||||
# No Mumuni text in stdout, the log, or any Zulip DM/stream payload.
|
||||
assert "Mumuni" not in alerts
|
||||
assert "Mumuni" not in log_path.read_text()
|
||||
assert MUMUNI_IP not in alerts + log_path.read_text()
|
||||
payloads = record.joinpath("curl.calls").read_text()
|
||||
assert "Mumuni" not in payloads
|
||||
assert MUMUNI_IP not in payloads
|
||||
|
||||
# The rest of the monitor still ran alongside the failing Tanko leg.
|
||||
log = log_path.read_text()
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "kagentz: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
|
||||
|
||||
# ── scripts/daily-infra-report.py: behavioral digest checks ──────────
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def daily():
|
||||
spec = importlib.util.spec_from_file_location("daily_infra_report", DAILY_REPORT)
|
||||
assert spec and spec.loader
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
DAILY_AGENTS = {
|
||||
"abiba": {
|
||||
"platform": "pi", "ct": 100, "ip": MUMUNI_IP,
|
||||
"zulip_connected": True, "zulip_processed": 5,
|
||||
"pm2_status": "online", "pm2_restarts": "0", "pm2_uptime": "1h",
|
||||
},
|
||||
"tanko": {
|
||||
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
|
||||
"gateway_state": "n/a (DSH)", "zulip_state": "connected",
|
||||
"telegram_state": "unknown", "gateway_pid": None, "updated_at": "",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def _fabricated_report(agents):
|
||||
return {
|
||||
"nodes": {},
|
||||
"node_count": 1,
|
||||
"nodes_online": 1,
|
||||
"total_vms": 0,
|
||||
"running_vms": 0,
|
||||
"stopped_vms": [],
|
||||
"vms_by_node": {n: [] for n in
|
||||
["amdpve", "minipve", "storepve", "acerpve", "ocupve"]},
|
||||
"storage": [],
|
||||
"docker_vm": {"total": 0, "running": 0, "unhealthy": [],
|
||||
"containers": [], "reclaimable": "", "disk_used": "1%"},
|
||||
"docker_syslog": {"total": 0, "running": 0, "containers": []},
|
||||
"docker_netbird": {"total": 0, "running": 0, "containers": []},
|
||||
"endpoints": [],
|
||||
"litellm": {"checks": []},
|
||||
"nfs": [],
|
||||
"zulip_ext": {
|
||||
"connected": True, "queue_id": "queue", "last_error": None,
|
||||
"messages_processed": 0, "retry_count": 0, "pm2": {},
|
||||
"pm2_healthy": True, "bot_skipped_15min": 0, "finalized_1h": 0,
|
||||
"failed_finalize_1h": 0, "finalize_fail_pct": 0,
|
||||
"server_status": "200",
|
||||
},
|
||||
"agents": agents,
|
||||
}
|
||||
|
||||
|
||||
def _agent_status_card(html):
|
||||
start = html.index("🤖 Agent Status")
|
||||
end = html.index("💬 Zulip Extension")
|
||||
return html[start:end]
|
||||
|
||||
|
||||
def test_daily_report_renders_only_abiba_and_tanko_agents(daily):
|
||||
"""build_html() over a Mumuni-free agent set must render no Mumuni row and
|
||||
no Mumuni gateway-unknown issue, while abiba and tanko rows still render."""
|
||||
html = daily.build_html(_fabricated_report(dict(DAILY_AGENTS)))
|
||||
card = _agent_status_card(html)
|
||||
assert "mumuni" not in card.lower()
|
||||
assert "abiba" in card
|
||||
assert "tanko" in card
|
||||
assert "mumuni" not in html.lower()
|
||||
|
||||
|
||||
def test_daily_report_collect_never_probes_mumuni(monkeypatch, daily):
|
||||
"""collect() with ssh stubbed must add no mumuni agent and must never ssh
|
||||
its decommissioned .24 host."""
|
||||
probed = []
|
||||
|
||||
class _NoSubprocess:
|
||||
@staticmethod
|
||||
def check_output(*args, **kwargs):
|
||||
return b""
|
||||
|
||||
def fake_ssh(host, cmd):
|
||||
probed.append(host)
|
||||
return ""
|
||||
|
||||
monkeypatch.setattr(daily, "pve_get", lambda path: [])
|
||||
monkeypatch.setattr(daily, "ssh_jerome", lambda host, cmd: "")
|
||||
monkeypatch.setattr(daily, "ssh", fake_ssh)
|
||||
monkeypatch.setattr(daily, "http_get",
|
||||
lambda url, auth=None, timeout=10: "200")
|
||||
monkeypatch.setattr(daily, "http_get_body",
|
||||
lambda url, auth=None, timeout=10: "")
|
||||
monkeypatch.setattr(daily, "count_in_log", lambda *a, **k: 0)
|
||||
monkeypatch.setattr(daily, "subprocess", _NoSubprocess)
|
||||
|
||||
report = daily.collect()
|
||||
assert "mumuni" not in report["agents"]
|
||||
assert MUMUNI_IP not in probed
|
||||
|
||||
|
||||
# ── scripts/agent-health-check.py: roster pin ───────────────────────
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def ahc():
|
||||
spec = importlib.util.spec_from_file_location("agent_health_check_roster", AHC)
|
||||
assert spec and spec.loader
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
def test_agent_health_roster_has_no_mumuni_entry(ahc):
|
||||
assert "mumuni" not in ahc.AGENTS
|
||||
|
||||
|
||||
# ── zulip-health.prose.md: contract reconciliation ──────────────────
|
||||
|
||||
def test_health_contract_retires_mumuni_only_steps():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
assert MUMUNI_IP not in text
|
||||
for step in ("**B4:", "**B5:", "**B6:"):
|
||||
assert step not in text
|
||||
|
||||
|
||||
def test_health_contract_states_mumuni_is_not_monitored_from_this_host():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
assert "Mumuni is NOT monitored from this host" in text
|
||||
assert "monitored on her side" in text
|
||||
assert "her own container" in text
|
||||
|
||||
|
||||
def test_health_contract_keeps_tanko_agent_zero_and_bridge_steps():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
for marker in ("**B1:", "**B2:", "**B3:", "Step 4: Platform C",
|
||||
"Step 2: Platform A", "Step 1: Zulip Server Liveness"):
|
||||
assert marker in text, marker
|
||||
+25
-43
@@ -1,9 +1,9 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-health
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||
version: 3.0.0
|
||||
version: 3.1.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
report_only_agents:
|
||||
@@ -12,13 +12,23 @@ report_only_agents:
|
||||
|
||||
# Zulip Mesh Health Monitor
|
||||
|
||||
Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero).
|
||||
Runs every 15 minutes in the background. Also triggers on session start.
|
||||
Monitors the Zulip-connected agents under this host's operational control (pi,
|
||||
DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on
|
||||
session start.
|
||||
|
||||
> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She
|
||||
> moved off this host onto her own container — kagentz CT 105 on minipve
|
||||
> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her
|
||||
> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may
|
||||
> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on
|
||||
> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat,
|
||||
> response delivery) are retired: they always read "unknown" against the
|
||||
> decommissioned deployment and produced a false 🔴 alert on every run.
|
||||
|
||||
## Requires
|
||||
|
||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); Mumuni (192.168.68.14, kagentz CT105 on minipve); and Agent Zero Docker host (192.168.68.14)
|
||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||
- **PM2** on localhost for pi process management
|
||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||
@@ -81,13 +91,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
|
||||
## Streaming Support (2026-07-05)
|
||||
|
||||
Zulip agents now support progressive message editing during agent generation.
|
||||
When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is
|
||||
When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is
|
||||
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
||||
|
||||
- Adapter implements `edit_message()` using `_api_patch()` helper
|
||||
- Gateway stream consumer progressively edits the Zulip message
|
||||
- User sees real-time agent thinking instead of waiting for full response
|
||||
- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active
|
||||
- Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here
|
||||
|
||||
### Verification
|
||||
```bash
|
||||
@@ -183,7 +193,10 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
|
||||
| Crash loop >10/h | Alert user |
|
||||
|
||||
|
||||
### Step 3: Platform B — Tanko (DSH on amdpve CT 112) & Mumuni (Hermes)
|
||||
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
|
||||
|
||||
Mumuni is out of scope for this host (see the note above): she runs on her own
|
||||
container and is monitored on her side.
|
||||
|
||||
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
|
||||
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
|
||||
@@ -240,44 +253,13 @@ timeout only. Never expect a bare `200` — the public URL terminates in the
|
||||
token-gated authentik chain. Statuses outside the healthy set are
|
||||
logged/reported as a warning — reported, never healed on.
|
||||
|
||||
**B4: Gateway Process** (Hermes agent Mumuni only — Tanko runs no Hermes gateway)
|
||||
|
||||
```bash
|
||||
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
|
||||
```
|
||||
|
||||
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more
|
||||
than one `gateway run` process is found, the gateway has a collision (typically
|
||||
one `--force` and one `--replace` process). Kill the newer/duplicate process,
|
||||
then restart the remaining gateway per-agent (parameterized 2026-08-09, captain
|
||||
ruling). Check the gateway log for "Gateway running with 2 platform(s)" (not 1)
|
||||
to confirm Zulip reloaded.
|
||||
|
||||
**B5: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
|
||||
```
|
||||
|
||||
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
|
||||
Silence > 300s → warning. Silence > 600s → critical.
|
||||
|
||||
**B6: Response Delivery** (Hermes agent Mumuni only)
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
|
||||
```
|
||||
|
||||
> 50% fail rate → critical.
|
||||
|
||||
**Platform B Actions**
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service |
|
||||
| No heartbeat in 10min | Same as above |
|
||||
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
||||
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
||||
| `dsh-web` service not `active` | Restart Tanko via DSH service |
|
||||
| HTTP `:3080` connection refused/timeout (`000`) | Same as above |
|
||||
| HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
|
||||
|
||||
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
|
||||
|
||||
@@ -333,7 +315,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I
|
||||
|
||||
Check each agent's log for excessive bot-to-bot chatter:
|
||||
- Abiba: `Skipped.*bot msgs` count
|
||||
- Tanko/Mumuni: Repeated DM exchanges between bots
|
||||
- Tanko: Repeated DM exchanges between bots
|
||||
- kagentz: Adapter log for bot DMs being processed
|
||||
|
||||
If any bot processes >50 bot-originated messages in 15min → warning.
|
||||
|
||||
Reference in New Issue
Block a user