Compare commits
5
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
2f961d7e7a | ||
|
|
4fe4f3621d | ||
|
|
86d2987ad8 | ||
|
|
efe9381283 | ||
|
|
6a1c4967db |
@@ -17,7 +17,7 @@ description: >
|
|||||||
⚠️ This contract is target-state aspirational — but GPU export + alerting
|
⚠️ This contract is target-state aspirational — but GPU export + alerting
|
||||||
are now as-built (verified 2026-08-09).
|
are now as-built (verified 2026-08-09).
|
||||||
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
||||||
version: 1.0.0
|
version: 1.0.1
|
||||||
---
|
---
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
@@ -120,132 +120,156 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
|||||||
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
|
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
|
||||||
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
|
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
|
||||||
|
|
||||||
|
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||||
|
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||||
|
service answered — report the code, never "down". A redirect is not a failure.
|
||||||
|
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe,
|
||||||
|
and that is a statement about YOUR PROBE, not about the service.
|
||||||
|
2. **A failed probe is never a service verdict.** Print
|
||||||
|
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
||||||
|
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
||||||
|
then report. Apply the same shape as scripts/disk-gc-scan.py.
|
||||||
|
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
||||||
|
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
||||||
|
(retried at 25s: also timeout)" is actionable.
|
||||||
|
|
||||||
### check-health
|
### check-health
|
||||||
|
|
||||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
||||||
|
tool calls; never repeat a prior report unless a live probe fails.**
|
||||||
|
|
||||||
|
**PROBE SHAPE (per standing rules above):**
|
||||||
|
- Every probe prints the target name + URL + HTTP code (or failure kind)
|
||||||
|
- Retry once on connection failure at longer timeout
|
||||||
|
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
||||||
|
- Report the actual probe command and its result, not a summary verdict
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Provenance — run first; paste the absolute path into the report
|
# Provenance — run first; paste the absolute path into the report
|
||||||
pwd -P
|
pwd -P
|
||||||
|
|
||||||
# Zulip API health (POST ping)
|
# ============================================================
|
||||||
|
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
|
||||||
|
# ============================================================
|
||||||
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
|
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
|
||||||
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
|
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
|
||||||
if [ -z "$zulip_key" ]; then
|
if [ -z "$zulip_key" ]; then
|
||||||
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
|
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
|
||||||
else
|
else
|
||||||
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
||||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}"
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
|
||||||
# Expected: 200 (HTTP 000 = unreachable/cache)
|
echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
|
||||||
|
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# PM2 process health
|
# ============================================================
|
||||||
|
# 2. PM2 PROCESS HEALTH
|
||||||
|
# ============================================================
|
||||||
pm2 jlist
|
pm2 jlist
|
||||||
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
|
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
|
||||||
|
# spoton-service removed 2026-09-14 (not in live set)
|
||||||
|
|
||||||
# GPU exporters (may be down per DEPLOYMENT STATUS)
|
# ============================================================
|
||||||
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
|
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
|
||||||
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
|
# ============================================================
|
||||||
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
|
# Probe /metrics (the Prometheus scrape target), not bare /
|
||||||
|
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
|
||||||
|
if [ "$code" == "000" ]; then
|
||||||
|
# Retry with longer timeout
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
|
||||||
|
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "GPU exporter http://$host:9400/metrics -> $code"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
|
||||||
|
|
||||||
# Router health (via nginx on port 80)
|
# ============================================================
|
||||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
|
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
|
||||||
# Expected: 200 (Router is up and responding)
|
# ============================================================
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
|
||||||
|
if [ "$code" == "000" ]; then
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
|
||||||
|
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "Router http://192.168.68.116/health -> $code"
|
||||||
|
fi
|
||||||
|
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||||
|
|
||||||
# LiteLLM health (via nginx on port 80)
|
# ============================================================
|
||||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
|
||||||
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
|
# ============================================================
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
|
||||||
|
if [ "$code" == "000" ]; then
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
|
||||||
|
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
|
||||||
|
fi
|
||||||
|
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
|
||||||
|
|
||||||
# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring
|
# ============================================================
|
||||||
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 —
|
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
|
||||||
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host,
|
# ============================================================
|
||||||
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE
|
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
|
||||||
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
|
|
||||||
# DOWN = connection refused (000) or timeout only.
|
|
||||||
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
|
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
|
||||||
printf '%s:8006 -> %s\n' "$node" \
|
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
|
||||||
"$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
|
if [ "$code" == "000" ]; then
|
||||||
|
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
|
||||||
|
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "PVE API https://$node:8006/api2/json/version -> $code"
|
||||||
|
fi
|
||||||
done
|
done
|
||||||
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
|
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
|
||||||
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault.
|
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
|
||||||
|
|
||||||
# Prometheus targets
|
# ============================================================
|
||||||
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
|
# 7. PROMETHEUS TARGETS — bare-200 probe
|
||||||
# Expected: All targets UP (may show some down if exporters not deployed)
|
# ============================================================
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
|
||||||
|
if [ "$code" == "000" ]; then
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
|
||||||
|
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
|
||||||
|
fi
|
||||||
|
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||||
|
|
||||||
# Grafana health
|
# ============================================================
|
||||||
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
|
# 8. GRAFANA HEALTH — any-HTTP probe
|
||||||
# Expected: {"status":"ok","version":"..."}
|
# ============================================================
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
|
||||||
|
if [ "$code" == "000" ]; then
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
|
||||||
|
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
|
||||||
|
fi
|
||||||
|
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
|
||||||
|
|
||||||
# LiteLLM metrics (Prometheus endpoint)
|
# ============================================================
|
||||||
curl -s http://192.168.68.116:4000/metrics | head -20
|
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
|
||||||
# Expected: Prometheus-formatted metrics output
|
# ============================================================
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
|
||||||
|
if [ "$code" == "000" ]; then
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
|
||||||
|
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
|
||||||
|
fi
|
||||||
|
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
|
||||||
```
|
```
|
||||||
|
|
||||||
**Report format**: Begin every report with the **absolute path the probe executed
|
**Report format**: Begin every report with the **absolute path the probe executed
|
||||||
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
|
from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault
|
||||||
distinguishable from a real fault at read time. Summarize actual results from
|
at read time. For each probe, print the target name, the full URL, and the HTTP
|
||||||
each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE
|
code (or failure kind with retry details). Apply the standing probe rules: any
|
||||||
API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is
|
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
||||||
DOWN; empty output is a warning. For probes whose expected result is a bare `200`
|
a failure.
|
||||||
(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected
|
|
||||||
status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200`
|
|
||||||
expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect)
|
|
||||||
is a stale expectation, not a fault.
|
|
||||||
|
|
||||||
|
|
||||||
### Phase 1: GPU Exporters
|
### Phase 1: GPU Exporters
|
||||||
|
|
||||||
**NVIDIA (.8 and .110)**:
|
**NVIDIA (.8 and .110)**:
|
||||||
1. Download `nvidia_gpu_exporter` binary
|
1. Download `nvidia_gpu_exporter` binary
|
||||||
2. Create systemd service `nvidia-gpu-exporter.service`
|
|
||||||
3. Start and enable
|
|
||||||
|
|
||||||
**AMD (.15)**:
|
|
||||||
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
|
|
||||||
2. Parses `amdgpu_top --json -d 1000` output
|
|
||||||
3. Exposes key metrics at `:9400/metrics` via Python http.server
|
|
||||||
4. Create systemd service
|
|
||||||
5. Start and enable
|
|
||||||
|
|
||||||
### Phase 2: Prometheus
|
|
||||||
|
|
||||||
1. Create `/opt/monitoring/` directory on CT 116
|
|
||||||
2. Write `prometheus.yml` with scrape configs for all targets
|
|
||||||
3. Add to docker-compose (or separate compose file)
|
|
||||||
4. Start container
|
|
||||||
|
|
||||||
### Phase 3: Grafana
|
|
||||||
|
|
||||||
1. Create `/opt/monitoring/grafana/` directories
|
|
||||||
2. Provision Prometheus datasource
|
|
||||||
3. Provision GPU fleet dashboard JSON
|
|
||||||
4. Provision LiteLLM dashboard JSON
|
|
||||||
5. Add to docker-compose
|
|
||||||
6. Start container
|
|
||||||
|
|
||||||
### Phase 4: Verification
|
|
||||||
|
|
||||||
1. Verify all 3 GPU exporters return 200 at :9400/metrics
|
|
||||||
2. Verify Prometheus targets all UP at :9090/targets
|
|
||||||
3. Verify Grafana accessible at :3001 with dashboards
|
|
||||||
4. Verify LiteLLM metrics flowing to Prometheus
|
|
||||||
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
|
|
||||||
|
|
||||||
## Verification Commands
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# GPU exporters
|
|
||||||
curl -s http://192.168.68.8:9400/metrics | grep nvidia
|
|
||||||
curl -s http://192.168.68.110:9400/metrics | grep nvidia
|
|
||||||
curl -s http://192.168.68.15:9400/metrics | grep amdgpu
|
|
||||||
|
|
||||||
# Prometheus
|
|
||||||
curl -s http://192.168.68.116:9090/api/v1/targets
|
|
||||||
|
|
||||||
# Grafana
|
|
||||||
curl -s http://192.168.68.116:3001/api/health
|
|
||||||
|
|
||||||
# LiteLLM metrics (already live)
|
|
||||||
curl -s http://192.168.68.116:4000/metrics | head -20
|
|
||||||
```
|
|
||||||
|
|||||||
@@ -340,6 +340,36 @@ def check_gpu_ports():
|
|||||||
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
|
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
|
||||||
# ═══════════════════════════════════════════════════════════════════
|
# ═══════════════════════════════════════════════════════════════════
|
||||||
|
|
||||||
|
def _ssh_retry(host, cmd, user="root", timeout=15, retry_timeout=25, label=""):
|
||||||
|
"""SSH with one retry at a longer timeout.
|
||||||
|
|
||||||
|
Returns (stdout_or_None, probe_failed_bool, fail_kind).
|
||||||
|
When probe_failed is True, fail_kind is one of: timeout, ssh-failed.
|
||||||
|
"""
|
||||||
|
import subprocess as _sp
|
||||||
|
def _attempt(tmo, conn_tmo):
|
||||||
|
try:
|
||||||
|
r = _sp.run(
|
||||||
|
["ssh", "-o", "StrictHostKeyChecking=no", "-o", f"ConnectTimeout={conn_tmo}",
|
||||||
|
f"{user}@{host}", cmd],
|
||||||
|
capture_output=True, text=True, timeout=tmo)
|
||||||
|
return r.stdout.strip() if r.returncode == 0 else None
|
||||||
|
except _sp.TimeoutExpired:
|
||||||
|
return "__timeout__"
|
||||||
|
except:
|
||||||
|
return None
|
||||||
|
result = _attempt(timeout, 8)
|
||||||
|
if result is None or result == "__timeout__":
|
||||||
|
kind = "timeout" if result == "__timeout__" else "ssh-failed"
|
||||||
|
prefix = f"{label} " if label else ""
|
||||||
|
print(f" probe-failed: {prefix}ssh {user}@{host} — {kind} (retrying at {retry_timeout}s…)")
|
||||||
|
result = _attempt(retry_timeout, 15)
|
||||||
|
if result is None or result == "__timeout__":
|
||||||
|
kind = "timeout" if result == "__timeout__" else "ssh-failed"
|
||||||
|
return None, True, kind
|
||||||
|
return result, False, None
|
||||||
|
|
||||||
|
|
||||||
def check_agents():
|
def check_agents():
|
||||||
for name, agent in AGENTS.items():
|
for name, agent in AGENTS.items():
|
||||||
host = agent.get("host")
|
host = agent.get("host")
|
||||||
@@ -355,42 +385,49 @@ def check_agents():
|
|||||||
is_dsh = agent.get("runtime") == "dsh"
|
is_dsh = agent.get("runtime") == "dsh"
|
||||||
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
|
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
|
||||||
since = "since 2026-08-27" if is_dsh else "since the harness purge"
|
since = "since 2026-08-27" if is_dsh else "since the harness purge"
|
||||||
live = ssh(host, "true", user=user)
|
live, probe_failed, fail_kind = _ssh_retry(host, "true", user=user)
|
||||||
print(f" {'✅' if live is not None else '❌'} {name}: {label} — "
|
if probe_failed:
|
||||||
f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
|
print(f" ❌ {name}: {label} — probe-failed: ssh {user}@{host} {fail_kind} "
|
||||||
if live is None:
|
f"(retried at 25s: also {fail_kind}) [CT {ct}]")
|
||||||
_fail(f"unreachable:{name}", name)
|
_fail(f"probe-failed:{name}:{fail_kind}", name)
|
||||||
|
else:
|
||||||
|
print(f" ✅ {name}: {label} — no Hermes gateway {since} "
|
||||||
|
f"(ssh {user}@{host} OK, CT {ct})")
|
||||||
continue
|
continue
|
||||||
|
|
||||||
if not host or not user:
|
if not host or not user:
|
||||||
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
||||||
continue
|
continue
|
||||||
|
|
||||||
# Resolve the Hermes gateway PID once, before the report-only branch:
|
# Resolve the Hermes gateway PID with retry. The probe target is
|
||||||
# the summary line below renders `pid`, and it used to be bound only in
|
# explicit: ssh {user}@{host} pgrep -f hermes gateway.
|
||||||
# the report-only path — leaving it unbound on the abiba/koonimo path
|
pid, probe_failed, fail_kind = _ssh_retry(
|
||||||
# raised UnboundLocalError and crashed the whole check. Agents without
|
host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||||
# a gateway get pid=?.
|
if not pid and not probe_failed:
|
||||||
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
pid, probe_failed, fail_kind = _ssh_retry(
|
||||||
if not pid:
|
host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||||
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
if not pid and not probe_failed:
|
||||||
if not pid:
|
|
||||||
pid = "?"
|
pid = "?"
|
||||||
|
|
||||||
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
|
if probe_failed:
|
||||||
|
print(f" ❌ {name}: probe-failed: ssh {user}@{host} {fail_kind} "
|
||||||
|
f"(retried at 25s: also {fail_kind}) [CT {ct}] — gateway status UNDETERMINED")
|
||||||
|
_fail(f"probe-failed:{name}:{fail_kind}", name)
|
||||||
|
continue
|
||||||
|
|
||||||
|
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only (captain's 2026-08-17 ruling)
|
||||||
if report_only:
|
if report_only:
|
||||||
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
|
|
||||||
# Still check gateway status for reporting purposes
|
|
||||||
if pid == "?":
|
if pid == "?":
|
||||||
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
|
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
|
||||||
|
f"-> no process found (reported only, NOT counted) [CT {ct}]")
|
||||||
_fail(f"gateway-down:{name}", name)
|
_fail(f"gateway-down:{name}", name)
|
||||||
continue
|
|
||||||
else:
|
else:
|
||||||
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
|
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
|
||||||
continue # Skip the rest of the check for Koby
|
f"-> pid={pid} (running, reported only, NOT repaired) [CT {ct}]")
|
||||||
|
continue # Skip the rest of the check for Koby
|
||||||
|
|
||||||
# Gateway state file
|
# Gateway state file
|
||||||
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
state, _, _ = _ssh_retry(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
||||||
if state:
|
if state:
|
||||||
try:
|
try:
|
||||||
st = json.loads(state)
|
st = json.loads(state)
|
||||||
@@ -408,21 +445,22 @@ def check_agents():
|
|||||||
]
|
]
|
||||||
streaming = "no"
|
streaming = "no"
|
||||||
for p in adapter_paths:
|
for p in adapter_paths:
|
||||||
has_edit = ssh(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
|
has_edit, _, _ = _ssh_retry(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
|
||||||
if has_edit and has_edit != "0":
|
if has_edit and has_edit != "0":
|
||||||
streaming = "yes"
|
streaming = "yes"
|
||||||
break
|
break
|
||||||
|
|
||||||
# Recent errors
|
# Recent errors
|
||||||
recent_errors = ssh(host,
|
recent_errors, _, _ = _ssh_retry(
|
||||||
|
host,
|
||||||
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
|
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
|
||||||
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
|
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
|
||||||
user=user)
|
user=user)
|
||||||
recent_errors = (recent_errors or "0").strip().split("\n")[-1]
|
recent_errors = (recent_errors or "0").strip().split("\n")[-1]
|
||||||
|
|
||||||
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
|
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
|
||||||
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
|
f"{name}: probe: ssh {user}@{host} — gw={gw_state} zulip={zulip} "
|
||||||
f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
|
f"streaming={streaming} errors_10m={recent_errors.strip() or '0'} pid={pid} [CT {ct}]")
|
||||||
|
|
||||||
|
|
||||||
# ═══════════════════════════════════════════════════════════════════
|
# ═══════════════════════════════════════════════════════════════════
|
||||||
|
|||||||
+32
-4
@@ -127,20 +127,48 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
|
|||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
|
### Liveness rule (scoped)
|
||||||
|
|
||||||
|
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||||
|
|
||||||
|
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||||
|
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||||
|
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
||||||
|
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
||||||
|
|
||||||
### Step 1: Zulip Server Liveness
|
### Step 1: Zulip Server Liveness
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \
|
# Probe the Zulip API (authenticated, any HTTP status = ALIVE)
|
||||||
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY'
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
|
||||||
|
if [ "$code" == "000" ]; then
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
|
||||||
|
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code"
|
||||||
|
fi
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: `200`. If not → mark `zulip_server_status: "down"`, skip per-platform checks, alert.
|
Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault.
|
||||||
|
|
||||||
### Step 2: Platform A — pi (Abiba, localhost)
|
### Step 2: Platform A — pi (Abiba, localhost)
|
||||||
|
|
||||||
**A1: Health Endpoint**
|
**A1: Health Endpoint**
|
||||||
|
|
||||||
Fetch `http://localhost:9200/health` as JSON. Check:
|
```bash
|
||||||
|
# Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE)
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health)
|
||||||
|
if [ "$code" == "000" ]; then
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health)
|
||||||
|
echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||||
|
else
|
||||||
|
echo "Abiba extension http://127.0.0.1:9200/health -> $code"
|
||||||
|
fi
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.**
|
||||||
|
|
||||||
|
Check the JSON payload:
|
||||||
|
|
||||||
| Field | Healthy | Critical |
|
| Field | Healthy | Critical |
|
||||||
|-------|---------|----------|
|
|-------|---------|----------|
|
||||||
|
|||||||
Reference in New Issue
Block a user