Compare commits
12
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
2e0b737f2d | ||
|
|
bbe9ee533f | ||
|
|
af3d364242 | ||
|
|
2716e55c16 | ||
|
|
0e4eda0abb | ||
|
|
801de0a25c | ||
|
|
36ae464f59 | ||
|
|
a3e97ce72b | ||
|
|
cc5fe0991c | ||
|
|
dc78604360 | ||
|
|
7bbf148778 | ||
|
|
3a25c7cce5 |
+22
-1
@@ -97,7 +97,28 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
|
||||
`curl http://localhost:9100/gpu-data | jq` — Full fleet status
|
||||
|
||||
### check-health
|
||||
`curl http://localhost:9100/health` — Monitor self-check
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# GPU Monitor health
|
||||
curl http://localhost:9100/health | jq
|
||||
# Expected: 200 with {"status": "healthy", "cache_age_seconds": <n>}
|
||||
|
||||
# Router health (via nginx on port 80)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
|
||||
# Expected: 200 (Router is up and responding)
|
||||
|
||||
# LiteLLM health (via nginx on port 80)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
||||
# Expected: 200 (LiteLLM is up and responding)
|
||||
|
||||
# Dashboard
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/
|
||||
# Expected: 200 (Dashboard is up and responding)
|
||||
```
|
||||
|
||||
**Report format**: Summarize actual results from each probe. If any probe returns non-200, flag as alert.
|
||||
|
||||
### view-dashboard
|
||||
Open `http://localhost:9100/` in browser — Live HTML dashboard
|
||||
|
||||
@@ -108,7 +108,9 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
||||
|
||||
```bash
|
||||
# Zulip API health (POST ping)
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY'
|
||||
source /etc/litellm-monitor.env
|
||||
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
|
||||
# Expected: 200 (HTTP 000 = unreachable/cache)
|
||||
|
||||
# PM2 process health
|
||||
@@ -120,6 +122,18 @@ curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
|
||||
# Router health (via nginx on port 80)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
|
||||
# Expected: 200 (Router is up and responding)
|
||||
|
||||
# LiteLLM health (via nginx on port 80)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
||||
# Expected: 200 (LiteLLM is up and responding)
|
||||
|
||||
# PVE API (401 expected for unauthenticated probe — API is up over https)
|
||||
curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json
|
||||
# Expected: 401 (unauthorized — API is up; 000 = unreachable, 500 = API down)
|
||||
|
||||
# Prometheus targets
|
||||
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
|
||||
# Expected: All targets UP (may show some down if exporters not deployed)
|
||||
@@ -128,7 +142,7 @@ curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
|
||||
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
|
||||
# Expected: {"status":"ok","version":"..."}
|
||||
|
||||
# LiteLLM metrics
|
||||
# LiteLLM metrics (Prometheus endpoint)
|
||||
curl -s http://192.168.68.116:4001/metrics | head -20
|
||||
# Expected: Prometheus-formatted metrics output
|
||||
```
|
||||
|
||||
@@ -0,0 +1,117 @@
|
||||
---
|
||||
kind: pattern
|
||||
name: litellm-client-timeouts
|
||||
description: >
|
||||
Standard client timeout and retry policy for ALL agents calling LiteLLM
|
||||
(CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident:
|
||||
a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures
|
||||
(82% from abiba-pi, 13 from mumuni) against a backend that was actually
|
||||
succeeding at 20-70s per call once clients stopped giving up. Grounded in
|
||||
measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes
|
||||
to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h;
|
||||
nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely
|
||||
client-side. Blast radius if wrong: agents fall back to DeepSeek silently
|
||||
(key/timeout failures present as model degradation, not errors) or abandon
|
||||
healthy-but-slow reasoning calls, fragmenting long tasks.
|
||||
---
|
||||
|
||||
## Maintains
|
||||
|
||||
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
|
||||
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
|
||||
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
|
||||
|
||||
## The measured numbers these values come from
|
||||
|
||||
| Model | avg latency | avg TTFT | p-profile (24h) |
|
||||
|---|---|---|---|
|
||||
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
||||
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
|
||||
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
||||
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
|
||||
|
||||
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
||||
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
||||
tolerated 20-70s — 392 successful slow calls in the three hours after recovery.
|
||||
LiteLLM's internal queue time is ~0s; the latency is model inference, not
|
||||
proxy queuing.
|
||||
|
||||
## Parameters
|
||||
|
||||
### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
|
||||
|
||||
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy
|
||||
28.8s average + 120-300s tail. Every 408 in the incident was a client
|
||||
abandoning a request the backend would have answered.
|
||||
- If the transport exposes a timeout setting for the main model, set it to
|
||||
**300s or more**. If it does not (current Hermes custom-provider path has no
|
||||
timeout knob), that is acceptable ONLY because nginx holds the request for
|
||||
600s — but any wrapper, script, or direct API call you write MUST set its own
|
||||
timeout >= 300s for syslog-auto/qwen-class calls.
|
||||
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast
|
||||
here is what caused the incident.
|
||||
|
||||
### 2. Auxiliary tasks — keep template timeouts, one correction
|
||||
|
||||
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
|
||||
these are fine.
|
||||
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
||||
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
||||
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
|
||||
syslog-auto; delegation defaults that assume fast responses will 408 the
|
||||
same way.
|
||||
|
||||
### 3. Retry policy — backoff, not repetition
|
||||
|
||||
- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff
|
||||
(**15s, 45s**) before giving up.
|
||||
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour
|
||||
from one key) was a batch job retrying without backoff while the backend was
|
||||
down — it multiplied load during recovery.
|
||||
- On 401/403: do NOT retry — that is a key/permission problem (see
|
||||
litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
|
||||
- On 429: honor the retry-after header if present, else back off 60s.
|
||||
|
||||
### 4. Health probes — identify yourself and time out sanely
|
||||
|
||||
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12
|
||||
such orphans appeared in the incident window and cost investigation time).
|
||||
Send a real model name and use a real (probe-designated) key.
|
||||
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
|
||||
report "backend slow (>30s)" rather than hanging.
|
||||
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
|
||||
standard; sub-hourly synthetic traffic distorts latency baselines.
|
||||
|
||||
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
|
||||
|
||||
- The incident window showed bulk clients amplifying a backend stall 5:1.
|
||||
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
|
||||
- Any batch loop over N requests MUST sleep >= 5s between requests and honor
|
||||
the retry policy in section 3.
|
||||
|
||||
## Returns
|
||||
|
||||
- A single standard any agent or script can cite: timeouts >= 300s on the
|
||||
syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x,
|
||||
15s/45s), identified probes at 30s/hourly, batch jobs chunked and
|
||||
day-scheduled.
|
||||
- Failure signature recognition: bulk 408s from multiple keys in one window =
|
||||
backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated);
|
||||
single-key 408s = that client's timeout is too short.
|
||||
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s,
|
||||
auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent,
|
||||
stable aliases), litellm-api-keys.prose.md (key/permission failures).
|
||||
|
||||
## Intentionally NOT changed
|
||||
|
||||
- No server-side LiteLLM timeout/cooldown changes proposed — the incident
|
||||
self-recovered and the server is healthy (0.56s live probe); changing
|
||||
server behavior without process-level root cause (CT116 requires root;
|
||||
not reachable from kagentz) would be guessing.
|
||||
- No change to the template's vision/web_extract/compression timeouts —
|
||||
measured data says they are correct.
|
||||
- No per-agent key permission changes — those are litellm-api-keys.prose.md
|
||||
territory (and the open gpu-vision/gemma 403 items are already filed with
|
||||
the key owners).
|
||||
- No model routing changes — syslog-auto's weighted pool behaved correctly
|
||||
throughout the incident.
|
||||
@@ -113,7 +113,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||
|
||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v3 (2026-09-08) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -100,6 +100,32 @@ Edit `build-dashboards.py`, run it, `docker restart harness-grafana`
|
||||
### check-targets
|
||||
`curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'`
|
||||
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Prometheus health (bound to 0.0.0.0:9090 on .116)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy
|
||||
# Expected: 200 (Prometheus is up and healthy)
|
||||
|
||||
# Grafana health (bound to 0.0.0.0:3001 on .116)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health
|
||||
# Expected: 200 (Grafana is up and healthy)
|
||||
|
||||
# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics"
|
||||
# Expected: 200 (docker-stats-exporter is up and responding)
|
||||
|
||||
# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics"
|
||||
# Expected: 200 (pve-exporter is up and responding)
|
||||
```
|
||||
|
||||
**Report format**: Summarize actual results from each probe. If any probe returns non-200, flag as alert.
|
||||
|
||||
**Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN.
|
||||
|
||||
### restart-exporter
|
||||
`cd /opt/monitoring && docker compose restart pve-exporter docker-stats`
|
||||
|
||||
|
||||
@@ -19,6 +19,13 @@ Changelog:
|
||||
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
|
||||
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
|
||||
abiba (.24).
|
||||
v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service,
|
||||
.110 llama-server.service, .15 strix-server.service) — .8 was probing a stale
|
||||
llama-server unit that reads inactive, producing false UNREACHABLE legs.
|
||||
systemctl is-active no longer swallows non-zero exit as SSH failure.
|
||||
Fixed UnboundLocalError on the abiba/koonimo gateway leg (pid unbound in the
|
||||
summary f-string). Abiba's LiteLLM key now comes from /root/.pi/agent/env.sh
|
||||
(#735 agent separation; creds moved out of shared /root/.bashrc).
|
||||
"""
|
||||
|
||||
import subprocess, json, sys, os, time
|
||||
@@ -40,15 +47,25 @@ PVE_NODES = {
|
||||
# Agent definitions: ct, host, user, pve_node, vault_key_name
|
||||
AGENTS = {
|
||||
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
|
||||
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
|
||||
# abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its
|
||||
# local env file (key_env below), not from the shared vault or .bashrc.
|
||||
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve",
|
||||
"vault_key": None,
|
||||
"key_env": {"file": "/root/.pi/agent/env.sh", "var": "LITELLM_API_KEY"}},
|
||||
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
|
||||
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
|
||||
}
|
||||
|
||||
# Systemd units verified live 2026-09-08 (systemctl list-units on each host):
|
||||
# .8 rtx3090 (gpu-dense) -> llama-chat-api.service (active; the old
|
||||
# llama-server.service unit file is stale/inactive — probing it read as
|
||||
# UNREACHABLE for a healthy process)
|
||||
# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active)
|
||||
# .15 strixhalo (amdpve) -> strix-server.service (active)
|
||||
GPU_HOSTS = {
|
||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
|
||||
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"},
|
||||
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server"},
|
||||
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"},
|
||||
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"},
|
||||
"gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"},
|
||||
}
|
||||
|
||||
FAIL = []
|
||||
@@ -161,10 +178,41 @@ def _get_agent_key(agent_name, vault_key_name):
|
||||
|
||||
return None
|
||||
|
||||
# Inject keys from vault for each agent
|
||||
|
||||
def _read_env_export(path, var):
|
||||
"""Parse `export VAR=value` (or `VAR=value`) out of a local env file.
|
||||
|
||||
#735 agent separation (2026-09-06): agent creds moved out of the shared
|
||||
/root/.bashrc into per-agent env files under /root/.pi/agent/ (bashrc's
|
||||
source line keeps abiba shells resolving them, but the file of record is
|
||||
env.sh). Do NOT fall back to /root/.bashrc here: desktop (.200) SSH
|
||||
sessions override LITELLM_API_KEY with mumuni's key, so sourcing bashrc
|
||||
would validate the wrong identity.
|
||||
"""
|
||||
try:
|
||||
with open(os.path.expanduser(path)) as _f:
|
||||
for line in _f:
|
||||
line = line.strip()
|
||||
if not (line.startswith("export " + var + "=") or line.startswith(var + "=")):
|
||||
continue
|
||||
value = line.split("=", 1)[1].strip().strip('"').strip("'")
|
||||
if value:
|
||||
return value
|
||||
except (OSError, UnicodeDecodeError):
|
||||
pass
|
||||
return None
|
||||
|
||||
|
||||
# Inject keys for each agent:
|
||||
# - vault-backed agents (tanko/koby/koonimo): {NAME}_LITELLM_API_KEY from
|
||||
# Infisical (project 322fceab-39da-4854-a55a-568e76c0f13f, env prod).
|
||||
# - abiba (pi agent, no vault key): LITELLM_API_KEY from its local env file
|
||||
# /root/.pi/agent/env.sh (moved there from /root/.bashrc in #735).
|
||||
for agent_name in AGENTS:
|
||||
info = AGENTS[agent_name]
|
||||
key = _get_agent_key(agent_name, info.get("vault_key"))
|
||||
if not key and info.get("key_env"):
|
||||
key = _read_env_export(info["key_env"]["file"], info["key_env"]["var"])
|
||||
AGENTS[agent_name]["key"] = key
|
||||
|
||||
|
||||
@@ -176,7 +224,7 @@ def check_keys():
|
||||
for name, agent in AGENTS.items():
|
||||
key = agent.get("key")
|
||||
if not key:
|
||||
print(f" ❌ {name}: NO KEY FOUND (vault empty or unreachable)")
|
||||
print(f" ❌ {name}: NO KEY FOUND (vault/env empty or unreachable)")
|
||||
FAIL.append(f"key:{name}:no-key")
|
||||
continue
|
||||
data = http_json(f"{LITELLM}/v1/models",
|
||||
@@ -190,7 +238,7 @@ def check_keys():
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
# CHECK 2: GPU Port Conflict Detection (unchanged)
|
||||
# CHECK 2: GPU Port Conflict Detection (unit names verified live 2026-09-08)
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
|
||||
def check_gpu_ports():
|
||||
@@ -199,7 +247,11 @@ def check_gpu_ports():
|
||||
port = gpu["port"]
|
||||
svc = gpu["service"]
|
||||
|
||||
svc_status = ssh(host, f"systemctl is-active {svc}")
|
||||
# `systemctl is-active` exits non-zero when the unit is inactive or
|
||||
# missing, which the ssh() helper would swallow as an SSH failure and
|
||||
# report as UNREACHABLE. `|| true` keeps the real state word so we can
|
||||
# tell "unit inactive" from "host unreachable".
|
||||
svc_status = ssh(host, f"systemctl is-active {svc} || true")
|
||||
port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1")
|
||||
|
||||
if not svc_status:
|
||||
@@ -254,14 +306,22 @@ def check_agents():
|
||||
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
||||
continue
|
||||
|
||||
# Resolve the Hermes gateway PID once, before the report-only branch:
|
||||
# the summary line below renders `pid`, and it used to be bound only in
|
||||
# the report-only path — leaving it unbound on the abiba/koonimo path
|
||||
# raised UnboundLocalError and crashed the whole check. Agents without
|
||||
# a gateway get pid=?.
|
||||
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||
if not pid:
|
||||
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||
if not pid:
|
||||
pid = "?"
|
||||
|
||||
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
|
||||
if report_only:
|
||||
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
|
||||
# Still check gateway status for reporting purposes
|
||||
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||
if not pid:
|
||||
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||
if not pid:
|
||||
if pid == "?":
|
||||
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
|
||||
FAIL.append(f"gateway-down:{name}")
|
||||
continue
|
||||
|
||||
@@ -238,10 +238,11 @@ ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.her
|
||||
**C1: A2A Server Health**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json"
|
||||
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
|
||||
ssh root@192.168.68.14 "curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:50080/a2a/"
|
||||
```
|
||||
|
||||
Expected: `{"name":"kagentz",...}`. Connection refused → A2A server down.
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (000) → A2A server down.
|
||||
|
||||
**C2: Adapter Process**
|
||||
|
||||
@@ -262,12 +263,14 @@ Check: `processed=N` incrementing, `silence < 600s`, `reconnects` ≈ 0.
|
||||
**C4: A2A Response Verification**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:8001/a2a \
|
||||
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
|
||||
ssh root@192.168.68.14 "curl -s -X POST http://127.0.0.1:50080/a2a \
|
||||
-H 'Content-Type: application/json' \
|
||||
-H 'Authorization: Bearer $LITELLM_KEY' \
|
||||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||
```
|
||||
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`.
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
|
||||
|
||||
**Platform C Actions**
|
||||
|
||||
|
||||
Reference in New Issue
Block a user