Compare commits

..
Author SHA1 Message Date
root 2eb901debf fix: correct zulip-health Tanko probe to use CT 115 (amdpve) for SSH access
Fixed SSH key gap issue by probing from CT 115 (amdpve) instead of CT 117 (storepve).
Changed from 'ps aux' to 'systemctl is-active dsh-web' + localhost curl probe.

Fixes false DEGRADED status from SSH authentication failure.
2026-09-08 09:39:40 +00:00
root 801de0a25c fix: correct proxmox-monitor probe endpoints to use localhost for exporters
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only), not 0.0.0.0.
They must be probed from .116 via SSH. Prometheus and Grafana remain 0.0.0.0 (LAN-reachable).

Fixes false alarm from remote probe showing connection-refused (by design for localhost binds).
2026-09-08 08:41:21 +00:00
root 36ae464f59 fix: probe alignment across multiple monitoring contracts
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
- infra-monitoring: Added Router+LiteLLM+PVE API probes with expected HTTP codes (200/401/500)
- gpu-monitor: Added explicit check-health section with Router+LiteLLM probe commands
- proxmox-monitor: Added check-health section for Prometheus/Grafana/exporters (all bound to 0.0.0.0)
- zulip-health: Fixed A2A port from :8001 to :50080 and added auth-gated expectation (401)

Closes: infra-monitoring-probe-alignment, gpu-monitor-probe-alignment, proxmox-monitor-probe-alignment, zulip-health-probe-alignment
2026-09-08 07:20:41 +00:00
root a3e97ce72b fix: add Router+LiteLLM+PVE API probes to infrastructure-monitoring check-health
- Added Router health probe (http://192.168.68.116/health) — expected 200
- Added LiteLLM health probe (http://192.168.68.116/litellm/health) — expected 200
- Added PVE API probe (https://192.168.68.116:8006/api2/json) — 401 expected for unauthenticated
- Clarified that 401 means API is up, 000 means unreachable, 500 means API down

Closes: infra-monitoring-probe-alignment
2026-09-08 07:20:41 +00:00
root cc5fe0991c Move Zulip bot creds from inline to env/vault reference
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The abiba-bot@chat.sysloggh.net:KEY is now sourced from
/etc/litellm-monitor.env as ZULIP_USER and ZULIP_BOT_KEY.
2026-09-08 06:18:21 +00:00
mumuni-bot dc78604360 Merge pull request 'feat: litellm-client-timeouts — standard client timeout/retry policy for all agents' (#61) from fix/litellm-client-timeouts into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
2026-09-08 06:14:03 +00:00
mumuni-bot 7bbf148778 feat: litellm-client-timeouts contract — standard client timeout/retry policy
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Root cause of the 2026-09-06 incident: 157 client-abandoned 408s against a
backend that was succeeding at 20-70s/call once clients stopped giving up.
Measured grounding: syslog-auto 28.8s avg / 25.4s TTFT over 888 calls;
nginx already allows 600s (Rule 5); the gap was entirely client-side.

Values: >=300s primary path, >=120s gpu-dense delegation, 2 retries with
15s/45s backoff, identified probes at 30s/hourly, batch jobs chunked +
day-scheduled. Verified live: Prometheus metrics, 0.56s syslog-auto probe,
gpu-fleet /health/unified topology.
2026-09-08 06:07:43 +00:00
mumuni-bot 3a25c7cce5 Merge pull request 'fix: restore agent-health-check AGENTS roster + land missed Tanko-DSH rows' (#60) from fix/tanko-dsh-reland into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-07 07:34:34 +00:00
5 changed files with 200 additions and 12 deletions
+22 -1
View File
@@ -97,7 +97,28 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
`curl http://localhost:9100/gpu-data | jq` — Full fleet status
### check-health
`curl http://localhost:9100/health` — Monitor self-check
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
```bash
# GPU Monitor health
curl http://localhost:9100/health | jq
# Expected: 200 with {"status": "healthy", "cache_age_seconds": <n>}
# Router health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (Router is up and responding)
# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 200 (LiteLLM is up and responding)
# Dashboard
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/
# Expected: 200 (Dashboard is up and responding)
```
**Report format**: Summarize actual results from each probe. If any probe returns non-200, flag as alert.
### view-dashboard
Open `http://localhost:9100/` in browser — Live HTML dashboard
+16 -2
View File
@@ -108,7 +108,9 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
```bash
# Zulip API health (POST ping)
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY'
source /etc/litellm-monitor.env
ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
# Expected: 200 (HTTP 000 = unreachable/cache)
# PM2 process health
@@ -120,6 +122,18 @@ curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
# Router health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (Router is up and responding)
# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 200 (LiteLLM is up and responding)
# PVE API (401 expected for unauthenticated probe — API is up over https)
curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json
# Expected: 401 (unauthorized — API is up; 000 = unreachable, 500 = API down)
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)
@@ -128,7 +142,7 @@ curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# LiteLLM metrics
# LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4001/metrics | head -20
# Expected: Prometheus-formatted metrics output
```
+117
View File
@@ -0,0 +1,117 @@
---
kind: pattern
name: litellm-client-timeouts
description: >
Standard client timeout and retry policy for ALL agents calling LiteLLM
(CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident:
a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures
(82% from abiba-pi, 13 from mumuni) against a backend that was actually
succeeding at 20-70s per call once clients stopped giving up. Grounded in
measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes
to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h;
nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely
client-side. Blast radius if wrong: agents fall back to DeepSeek silently
(key/timeout failures present as model degradation, not errors) or abandon
healthy-but-slow reasoning calls, fragmenting long tasks.
---
## Maintains
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
## The measured numbers these values come from
| Model | avg latency | avg TTFT | p-profile (24h) |
|---|---|---|---|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
| strix-moe | 7.5s | — | Strix Halo, healthy |
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
backend), full recovery 07:00-08:00 with ZERO client failures once requests
tolerated 20-70s — 392 successful slow calls in the three hours after recovery.
LiteLLM's internal queue time is ~0s; the latency is model inference, not
proxy queuing.
## Parameters
### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy
28.8s average + 120-300s tail. Every 408 in the incident was a client
abandoning a request the backend would have answered.
- If the transport exposes a timeout setting for the main model, set it to
**300s or more**. If it does not (current Hermes custom-provider path has no
timeout knob), that is acceptable ONLY because nginx holds the request for
600s — but any wrapper, script, or direct API call you write MUST set its own
timeout >= 300s for syslog-auto/qwen-class calls.
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast
here is what caused the incident.
### 2. Auxiliary tasks — keep template timeouts, one correction
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
these are fine.
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
syslog-auto; delegation defaults that assume fast responses will 408 the
same way.
### 3. Retry policy — backoff, not repetition
- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff
(**15s, 45s**) before giving up.
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour
from one key) was a batch job retrying without backoff while the backend was
down — it multiplied load during recovery.
- On 401/403: do NOT retry — that is a key/permission problem (see
litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
- On 429: honor the retry-after header if present, else back off 60s.
### 4. Health probes — identify yourself and time out sanely
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12
such orphans appeared in the incident window and cost investigation time).
Send a real model name and use a real (probe-designated) key.
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
report "backend slow (>30s)" rather than hanging.
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
standard; sub-hourly synthetic traffic distorts latency baselines.
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
- The incident window showed bulk clients amplifying a backend stall 5:1.
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
- Any batch loop over N requests MUST sleep >= 5s between requests and honor
the retry policy in section 3.
## Returns
- A single standard any agent or script can cite: timeouts >= 300s on the
syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x,
15s/45s), identified probes at 30s/hourly, batch jobs chunked and
day-scheduled.
- Failure signature recognition: bulk 408s from multiple keys in one window =
backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated);
single-key 408s = that client's timeout is too short.
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s,
auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent,
stable aliases), litellm-api-keys.prose.md (key/permission failures).
## Intentionally NOT changed
- No server-side LiteLLM timeout/cooldown changes proposed — the incident
self-recovered and the server is healthy (0.56s live probe); changing
server behavior without process-level root cause (CT116 requires root;
not reachable from kagentz) would be guessing.
- No change to the template's vision/web_extract/compression timeouts —
measured data says they are correct.
- No per-agent key permission changes — those are litellm-api-keys.prose.md
territory (and the open gpu-vision/gemma 403 items are already filed with
the key owners).
- No model routing changes — syslog-auto's weighted pool behaved correctly
throughout the incident.
+26
View File
@@ -100,6 +100,32 @@ Edit `build-dashboards.py`, run it, `docker restart harness-grafana`
### check-targets
`curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'`
### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
```bash
# Prometheus health (bound to 0.0.0.0:9090 on .116)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy
# Expected: 200 (Prometheus is up and healthy)
# Grafana health (bound to 0.0.0.0:3001 on .116)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health
# Expected: 200 (Grafana is up and healthy)
# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics"
# Expected: 200 (docker-stats-exporter is up and responding)
# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics"
# Expected: 200 (pve-exporter is up and responding)
```
**Report format**: Summarize actual results from each probe. If any probe returns non-200, flag as alert.
**Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN.
### restart-exporter
`cd /opt/monitoring && docker compose restart pve-exporter docker-stats`
+19 -9
View File
@@ -197,15 +197,22 @@ Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error`
**B2: Agent Process**
```bash
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
# Probe from CT 115 (amdpve) to avoid SSH key issues on CT 117
ssh root@192.168.68.15 "pct exec 112 'systemctl is-active dsh-web'"
```
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling):
Gateway should be active. **Dual-gateway detection**: check for duplicate processes.
| Agent | Restart command | Notes |
|-------|-----------------|-------|
**B3: Gateway Health**
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
```bash
# Probe localhost on CT 112 for DSH web service
ssh root@192.168.68.15 "pct exec 112 'curl -s -o /dev/null -w \"%{http_code}\" http://127.0.0.1:3080/'"
```
Expected: any HTTP code (200, 302, 401, etc.) = gateway is alive. Connection refused = gateway down.
**Note**: DSH web service listens on 127.0.0.1:3080 (loopback-only) by design. Public URL returns 302 through proxy-auth chain, not bare 200.
**B3: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
@@ -238,10 +245,11 @@ ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.her
**C1: A2A Server Health**
```bash
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json"
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
ssh root@192.168.68.14 "curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:50080/a2a/"
```
Expected: `{"name":"kagentz",...}`. Connection refused → A2A server down.
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (000) → A2A server down.
**C2: Adapter Process**
@@ -262,12 +270,14 @@ Check: `processed=N` incrementing, `silence < 600s`, `reconnects` ≈ 0.
**C4: A2A Response Verification**
```bash
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:8001/a2a \
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
ssh root@192.168.68.14 "curl -s -X POST http://127.0.0.1:50080/a2a \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer $LITELLM_KEY' \
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
```
Expected: task ID with "working" status. Poll for completion with `tasks/get`.
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
**Platform C Actions**