Compare commits

...
Author SHA1 Message Date
root 2f961d7e7a fix: agent-health-check gateway leg — deterministic probe-failed reporting
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output

Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind

Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
2026-09-14 15:21:36 +00:00
abiba-bot 4fe4f3621d Merge pull request 'fix(monitoring): probe precision - any HTTP status means alive, failed probes never become service verdicts' (#92) from fix/probe-precision-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 15:07:35 +00:00
root 86d2987ad8 fix: probe precision — add retry + probe-failed reporting to zulip-health
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Per defect report 1150.msg:
- Add standing probe rules section (2026-09-14)
- Step 1 (Zulip API): retry once at 25s on 000, print target + code
- Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host
- Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed

Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
2026-09-14 14:53:57 +00:00
root efe9381283 fix: probe precision — correct GPU exporter path, Grafana port, add retry + probe-failed reporting
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed

Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
2026-09-14 14:44:37 +00:00
abiba-bot 6a1c4967db Merge pull request 'fix(disk-gc): deterministic reachability verdict and correctly labelled disk figures' (#91) from fix/disk-gc-probe-deterministic-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 14:43:43 +00:00
3 changed files with 216 additions and 126 deletions
+120 -96
View File
@@ -17,7 +17,7 @@ description: >
⚠️ This contract is target-state aspirational — but GPU export + alerting ⚠️ This contract is target-state aspirational — but GPU export + alerting
are now as-built (verified 2026-08-09). are now as-built (verified 2026-08-09).
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
version: 1.0.0 version: 1.0.1
--- ---
## Architecture ## Architecture
@@ -120,132 +120,156 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
`/health` — an unexpected status (`401`/`403` from a bad or missing credential, `/health` — an unexpected status (`401`/`403` from a bad or missing credential,
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive". `5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
service answered — report the code, never "down". A redirect is not a failure.
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe,
and that is a statement about YOUR PROBE, not about the service.
2. **A failed probe is never a service verdict.** Print
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
then report. Apply the same shape as scripts/disk-gc-scan.py.
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
(retried at 25s: also timeout)" is actionable.
### check-health ### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
tool calls; never repeat a prior report unless a live probe fails.**
**PROBE SHAPE (per standing rules above):**
- Every probe prints the target name + URL + HTTP code (or failure kind)
- Retry once on connection failure at longer timeout
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
- Report the actual probe command and its result, not a summary verdict
```bash ```bash
# Provenance — run first; paste the absolute path into the report # Provenance — run first; paste the absolute path into the report
pwd -P pwd -P
# Zulip API health (POST ping) # ============================================================
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
# ============================================================
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via: # NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2") zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
if [ -z "$zulip_key" ]; then if [ -z "$zulip_key" ]; then
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env" echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
else else
ZULIP_USER="abiba-bot@chat.sysloggh.net" ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}" code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
# Expected: 200 (HTTP 000 = unreachable/cache) echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
# Expected: 200 (bare-200 probe; any other status is an ALERT)
fi fi
# PM2 process health # ============================================================
# 2. PM2 PROCESS HEALTH
# ============================================================
pm2 jlist pm2 jlist
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service) # Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
# spoton-service removed 2026-09-14 (not in live set)
# GPU exporters (may be down per DEPLOYMENT STATUS) # ============================================================
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL" # 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL" # ============================================================
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL" # Probe /metrics (the Prometheus scrape target), not bare /
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
if [ "$code" == "000" ]; then
# Retry with longer timeout
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "GPU exporter http://$host:9400/metrics -> $code"
fi
done
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
# Router health (via nginx on port 80) # ============================================================
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health # 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
# Expected: 200 (Router is up and responding) # ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Router http://192.168.68.116/health -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# LiteLLM health (via nginx on port 80) # ============================================================
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health # 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive # ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
fi
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring # ============================================================
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 — # 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host, # ============================================================
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE # Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
# DOWN = connection refused (000) or timeout only.
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
printf '%s:8006 -> %s\n' "$node" \ code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
"$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")" if [ "$code" == "000" ]; then
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "PVE API https://$node:8006/api2/json/version -> $code"
fi
done done
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12) # Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault. # 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
# Prometheus targets # ============================================================
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' # 7. PROMETHEUS TARGETS — bare-200 probe
# Expected: All targets UP (may show some down if exporters not deployed) # ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# Grafana health # ============================================================
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}' # 8. GRAFANA HEALTH — any-HTTP probe
# Expected: {"status":"ok","version":"..."} # ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
# LiteLLM metrics (Prometheus endpoint) # ============================================================
curl -s http://192.168.68.116:4000/metrics | head -20 # 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
# Expected: Prometheus-formatted metrics output # ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
``` ```
**Report format**: Begin every report with the **absolute path the probe executed **Report format**: Begin every report with the **absolute path the probe executed
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault
distinguishable from a real fault at read time. Summarize actual results from at read time. For each probe, print the target name, the full URL, and the HTTP
each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE code (or failure kind with retry details). Apply the standing probe rules: any
API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
DOWN; empty output is a warning. For probes whose expected result is a bare `200` a failure.
(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected
status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200`
expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect)
is a stale expectation, not a fault.
### Phase 1: GPU Exporters ### Phase 1: GPU Exporters
**NVIDIA (.8 and .110)**: **NVIDIA (.8 and .110)**:
1. Download `nvidia_gpu_exporter` binary 1. Download `nvidia_gpu_exporter` binary
2. Create systemd service `nvidia-gpu-exporter.service`
3. Start and enable
**AMD (.15)**:
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
2. Parses `amdgpu_top --json -d 1000` output
3. Exposes key metrics at `:9400/metrics` via Python http.server
4. Create systemd service
5. Start and enable
### Phase 2: Prometheus
1. Create `/opt/monitoring/` directory on CT 116
2. Write `prometheus.yml` with scrape configs for all targets
3. Add to docker-compose (or separate compose file)
4. Start container
### Phase 3: Grafana
1. Create `/opt/monitoring/grafana/` directories
2. Provision Prometheus datasource
3. Provision GPU fleet dashboard JSON
4. Provision LiteLLM dashboard JSON
5. Add to docker-compose
6. Start container
### Phase 4: Verification
1. Verify all 3 GPU exporters return 200 at :9400/metrics
2. Verify Prometheus targets all UP at :9090/targets
3. Verify Grafana accessible at :3001 with dashboards
4. Verify LiteLLM metrics flowing to Prometheus
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
## Verification Commands
```bash
# GPU exporters
curl -s http://192.168.68.8:9400/metrics | grep nvidia
curl -s http://192.168.68.110:9400/metrics | grep nvidia
curl -s http://192.168.68.15:9400/metrics | grep amdgpu
# Prometheus
curl -s http://192.168.68.116:9090/api/v1/targets
# Grafana
curl -s http://192.168.68.116:3001/api/health
# LiteLLM metrics (already live)
curl -s http://192.168.68.116:4000/metrics | head -20
```
+65 -27
View File
@@ -340,6 +340,36 @@ def check_gpu_ports():
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents) # CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def _ssh_retry(host, cmd, user="root", timeout=15, retry_timeout=25, label=""):
"""SSH with one retry at a longer timeout.
Returns (stdout_or_None, probe_failed_bool, fail_kind).
When probe_failed is True, fail_kind is one of: timeout, ssh-failed.
"""
import subprocess as _sp
def _attempt(tmo, conn_tmo):
try:
r = _sp.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", f"ConnectTimeout={conn_tmo}",
f"{user}@{host}", cmd],
capture_output=True, text=True, timeout=tmo)
return r.stdout.strip() if r.returncode == 0 else None
except _sp.TimeoutExpired:
return "__timeout__"
except:
return None
result = _attempt(timeout, 8)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
prefix = f"{label} " if label else ""
print(f" probe-failed: {prefix}ssh {user}@{host} — {kind} (retrying at {retry_timeout}s…)")
result = _attempt(retry_timeout, 15)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
return None, True, kind
return result, False, None
def check_agents(): def check_agents():
for name, agent in AGENTS.items(): for name, agent in AGENTS.items():
host = agent.get("host") host = agent.get("host")
@@ -355,42 +385,49 @@ def check_agents():
is_dsh = agent.get("runtime") == "dsh" is_dsh = agent.get("runtime") == "dsh"
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime" label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
since = "since 2026-08-27" if is_dsh else "since the harness purge" since = "since 2026-08-27" if is_dsh else "since the harness purge"
live = ssh(host, "true", user=user) live, probe_failed, fail_kind = _ssh_retry(host, "true", user=user)
print(f" {'✅' if live is not None else '❌'} {name}: {label} — " if probe_failed:
f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})") print(f" ❌ {name}: {label} — probe-failed: ssh {user}@{host} {fail_kind} "
if live is None: f"(retried at 25s: also {fail_kind}) [CT {ct}]")
_fail(f"unreachable:{name}", name) _fail(f"probe-failed:{name}:{fail_kind}", name)
else:
print(f" ✅ {name}: {label} — no Hermes gateway {since} "
f"(ssh {user}@{host} OK, CT {ct})")
continue continue
if not host or not user: if not host or not user:
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check") print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
continue continue
# Resolve the Hermes gateway PID once, before the report-only branch: # Resolve the Hermes gateway PID with retry. The probe target is
# the summary line below renders `pid`, and it used to be bound only in # explicit: ssh {user}@{host} pgrep -f hermes gateway.
# the report-only path — leaving it unbound on the abiba/koonimo path pid, probe_failed, fail_kind = _ssh_retry(
# raised UnboundLocalError and crashed the whole check. Agents without host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
# a gateway get pid=?. if not pid and not probe_failed:
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user) pid, probe_failed, fail_kind = _ssh_retry(
if not pid: host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user) if not pid and not probe_failed:
if not pid:
pid = "?" pid = "?"
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only if probe_failed:
if report_only: print(f" ❌ {name}: probe-failed: ssh {user}@{host} {fail_kind} "
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)") f"(retried at 25s: also {fail_kind}) [CT {ct}] — gateway status UNDETERMINED")
# Still check gateway status for reporting purposes _fail(f"probe-failed:{name}:{fail_kind}", name)
if pid == "?":
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
_fail(f"gateway-down:{name}", name)
continue continue
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only (captain's 2026-08-17 ruling)
if report_only:
if pid == "?":
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> no process found (reported only, NOT counted) [CT {ct}]")
_fail(f"gateway-down:{name}", name)
else: else:
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)") print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> pid={pid} (running, reported only, NOT repaired) [CT {ct}]")
continue # Skip the rest of the check for Koby continue # Skip the rest of the check for Koby
# Gateway state file # Gateway state file
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user) state, _, _ = _ssh_retry(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
if state: if state:
try: try:
st = json.loads(state) st = json.loads(state)
@@ -408,21 +445,22 @@ def check_agents():
] ]
streaming = "no" streaming = "no"
for p in adapter_paths: for p in adapter_paths:
has_edit = ssh(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user) has_edit, _, _ = _ssh_retry(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
if has_edit and has_edit != "0": if has_edit and has_edit != "0":
streaming = "yes" streaming = "yes"
break break
# Recent errors # Recent errors
recent_errors = ssh(host, recent_errors, _, _ = _ssh_retry(
host,
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null " r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0", r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
user=user) user=user)
recent_errors = (recent_errors or "0").strip().split("\n")[-1] recent_errors = (recent_errors or "0").strip().split("\n")[-1]
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} " print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} " f"{name}: probe: ssh {user}@{host} — gw={gw_state} zulip={zulip} "
f"errors_10m={recent_errors.strip() or '0'} pid={pid}") f"streaming={streaming} errors_10m={recent_errors.strip() or '0'} pid={pid} [CT {ct}]")
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
+32 -4
View File
@@ -127,20 +127,48 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
## Execution ## Execution
### Liveness rule (scoped)
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
### Step 1: Zulip Server Liveness ### Step 1: Zulip Server Liveness
```bash ```bash
curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \ # Probe the Zulip API (authenticated, any HTTP status = ALIVE)
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY' code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code"
fi
``` ```
Expected: `200`. If not → mark `zulip_server_status: "down"`, skip per-platform checks, alert. Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault.
### Step 2: Platform A — pi (Abiba, localhost) ### Step 2: Platform A — pi (Abiba, localhost)
**A1: Health Endpoint** **A1: Health Endpoint**
Fetch `http://localhost:9200/health` as JSON. Check: ```bash
# Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE)
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health)
echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Abiba extension http://127.0.0.1:9200/health -> $code"
fi
```
Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.**
Check the JSON payload:
| Field | Healthy | Critical | | Field | Healthy | Critical |
|-------|---------|----------| |-------|---------|----------|