Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
64790ebd19 | ||
|
|
6fa0a255df | ||
|
|
c72436b406 | ||
|
|
a17379676d | ||
|
|
63990b84f7 | ||
|
|
7a5ddb46a9 | ||
|
|
568fec2efa | ||
|
|
641a52c6da | ||
|
|
58033f39c5 | ||
|
|
6c59988e7f | ||
|
|
9a2ee6faec | ||
|
|
a12abbeb14 | ||
|
|
0b9aebca37 |
+55
-1
@@ -38,6 +38,7 @@ index:
|
||||
by_category:
|
||||
compliance:
|
||||
- hermes-key-enforcement
|
||||
- litellm-api-keys
|
||||
- hermes-config-template
|
||||
- hermes-agent-baseline
|
||||
monitoring:
|
||||
@@ -101,6 +102,7 @@ index:
|
||||
proxmox:
|
||||
- proxmox-monitor
|
||||
litellm:
|
||||
- litellm-api-keys
|
||||
- litellm-health
|
||||
- litellm-self-heal
|
||||
memory:
|
||||
@@ -628,7 +630,7 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 3.3.0
|
||||
version: 3.4.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/15 * * * *'
|
||||
@@ -1869,6 +1871,58 @@ contracts:
|
||||
drift_alerts: []
|
||||
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||
- name: litellm-api-keys
|
||||
file: litellm-api-keys.prose.md
|
||||
kind: function
|
||||
category: compliance
|
||||
sensitivity: critical
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: on_demand
|
||||
cadence: null
|
||||
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 120
|
||||
requires: []
|
||||
protocol:
|
||||
- Load contract from prose-contracts/main
|
||||
- Retrieve master key from Infisical (project=infrastructure env=production)
|
||||
- Read live key-scoped model roster from CT 116 /v1/models
|
||||
- Create/rotate/verify the requested agent key with an EXPLICIT models list
|
||||
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
|
||||
verification:
|
||||
postconditions:
|
||||
- check: standard agent key is local-only
|
||||
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
|
||||
expect: 0 cloud models
|
||||
- check: key exists with correct alias
|
||||
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
|
||||
expect: 200 with matching alias
|
||||
artifact: key creation/rotation report
|
||||
receipt:
|
||||
format: json
|
||||
storage: ~/.hermes/runs/litellm-api-keys/
|
||||
graph_node: true
|
||||
escalation:
|
||||
info:
|
||||
action: log_to_receipt
|
||||
notify: []
|
||||
warning:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
critical:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
- ops
|
||||
|
||||
koby_report_only: true
|
||||
koby_host: "CT 111 (tdunna)"
|
||||
koby_ip: ".129"
|
||||
|
||||
@@ -18,6 +18,25 @@ author: Abiba (pi agent)
|
||||
|
||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
|
||||
|
||||
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
|
||||
|
||||
## Model Access Tiers (2026-09-20)
|
||||
|
||||
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
|
||||
|
||||
| Tier | Models | Who gets it |
|
||||
|------|--------|-------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
|
||||
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
|
||||
|
||||
**Rules:**
|
||||
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
|
||||
2. Agent keys MUST carry an **explicit local-only** `models` list.
|
||||
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
|
||||
4. The master key always bypasses scoping — it is admin-only, never for inference.
|
||||
|
||||
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
|
||||
|
||||
## Scope
|
||||
|
||||
Applies to all Hermes agent configs across all hosts. Covers these config sections:
|
||||
|
||||
@@ -759,4 +759,4 @@ which was kill+nohup outside systemd) are banned by policy.
|
||||
| Script | Why Disabled |
|
||||
|--------|-------------|
|
||||
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
|
||||
| `zulip-monitor.sh` (Abiba) | Replaced by agent-health-check.py + PM2 auto-restart |
|
||||
| `zulip-monitor.sh` (Abiba) | Standalone cron replaced by agent-health-check.py + PM2 auto-restart; the script itself remains active as the execution step of the `zulip-health` responsibility contract (see `zulip-health.prose.md`) |
|
||||
|
||||
@@ -71,6 +71,10 @@ description: >
|
||||
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
||||
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
||||
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
||||
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
|
||||
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
|
||||
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
|
||||
In LiteLLM Community both silently grant access to EVERY model, including cloud.
|
||||
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
||||
- Return the new key
|
||||
5. **If action == "rotate"**:
|
||||
@@ -87,6 +91,66 @@ description: >
|
||||
- Confirm key alias matches agent_name in LiteLLM key list
|
||||
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
||||
|
||||
## Cloud Provider Consolidation (2026-09-20)
|
||||
|
||||
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
|
||||
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
|
||||
|
||||
### Provider map (per-account namespacing)
|
||||
|
||||
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
|
||||
stay separate:
|
||||
|
||||
| Prefix | Upstream | Auth | Vault secret |
|
||||
|--------|----------|------|--------------|
|
||||
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
|
||||
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
|
||||
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
|
||||
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
|
||||
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
|
||||
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
|
||||
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
|
||||
|
||||
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
|
||||
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
|
||||
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
|
||||
|
||||
### Access tiers (MUST be enforced per key)
|
||||
|
||||
| Tier | Model names | Granted to |
|
||||
|------|-------------|------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
|
||||
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
|
||||
|
||||
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
|
||||
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
|
||||
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
|
||||
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
|
||||
|
||||
### Creating the cloud-enabled key
|
||||
|
||||
```bash
|
||||
# ALWAYS read the live roster first (key-scoped):
|
||||
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
|
||||
| jq -r '.data[].id'
|
||||
|
||||
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
|
||||
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
|
||||
```
|
||||
|
||||
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
|
||||
any `<prefix>/` cloud model.
|
||||
|
||||
### Adding a new cloud provider
|
||||
|
||||
1. Add the upstream key to Infisical `infrastructure/production/root`.
|
||||
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
|
||||
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
|
||||
3. Restart the `harness-litellm` container.
|
||||
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
|
||||
captain approval.
|
||||
5. Update this table and the access-tier section.
|
||||
|
||||
## Production Vault Access Process (canonical, 2026-07-17)
|
||||
|
||||
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
||||
|
||||
@@ -55,9 +55,12 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
||||
try:
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
if rc != 0:
|
||||
# Determine failure kind from curl exit code
|
||||
# Check if this is a timeout from run_command (rc=1, stderr="TIMEOUT")
|
||||
if rc == 1 and stderr == "TIMEOUT":
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
# Otherwise, determine failure kind from curl exit code
|
||||
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
||||
if rc == 28:
|
||||
elif rc == 28:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
elif rc == 7:
|
||||
return (000, "connection refused")
|
||||
@@ -73,6 +76,20 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
||||
except subprocess.TimeoutExpired:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
|
||||
def check_host_health(host_ip):
|
||||
"""Check if the GPU host's llama-chat-api health endpoint is reachable
|
||||
|
||||
Returns: (healthy: bool, detail: str)
|
||||
"""
|
||||
code, _ = probe_http("http://" + host_ip + ":8080/health", timeout=10)
|
||||
if code == 200:
|
||||
return True, "host healthy (200)"
|
||||
elif code == 000:
|
||||
return False, "host unreachable (timeout or refused)"
|
||||
else:
|
||||
return False, "host unhealthy (HTTP " + str(code) + ")"
|
||||
|
||||
|
||||
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
||||
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
||||
cmd = "curl -s -m " + str(timeout)
|
||||
@@ -115,43 +132,53 @@ def check_model_probes():
|
||||
|
||||
results = []
|
||||
|
||||
# Host health mapping: model -> host IP
|
||||
model_hosts = {
|
||||
"gpu-dense": "192.168.68.8", # RTX 3090
|
||||
"gpu-vision": "192.168.68.110", # RTX 5070
|
||||
"strix-moe": "192.168.68.15" # Strix Halo
|
||||
}
|
||||
|
||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||
# Single-host aliases: 30s initial timeout, retry once at 45s on failure
|
||||
# gpu-dense (RTX 3090) may need long warmup/prefill or concurrent generation hold
|
||||
host_ip = model_hosts[model]
|
||||
# Single-host aliases: 30s initial timeout, retry once at 90s on failure
|
||||
# Worst-case prefill ~76s, so 90s retry ensures we cover it
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
|
||||
first_kind = None # Track first attempt's failure kind
|
||||
first_kind = None
|
||||
if code == 000 and failure_kind:
|
||||
# Retry once with longer timeout (45s) before declaring failure
|
||||
first_kind = failure_kind
|
||||
time.sleep(1)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=45)
|
||||
timeout=90)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Both attempts failed - report both kinds
|
||||
if first_kind:
|
||||
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
|
||||
# Both attempts failed - check host health to distinguish busy from down
|
||||
host_healthy, host_detail = check_host_health(host_ip)
|
||||
if host_healthy:
|
||||
results.append((model, False, "busy (completion timed out after retry; " + host_detail + ")"))
|
||||
else:
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
|
||||
# Host unreachable - report both kinds
|
||||
if first_kind:
|
||||
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
|
||||
else:
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
|
||||
elif code == 200:
|
||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
elif code in (401, 403):
|
||||
# Credential fault - capture body and key alias
|
||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||
timeout=10)
|
||||
# Resolve key alias
|
||||
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
|
||||
alias = "monitor-20260813"
|
||||
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||
else:
|
||||
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
@@ -261,6 +288,7 @@ def main():
|
||||
print("")
|
||||
|
||||
all_pass = True
|
||||
degraded = [] # Track degraded (busy) checks
|
||||
|
||||
# Run all checks
|
||||
checks = [
|
||||
@@ -280,9 +308,15 @@ def main():
|
||||
# Model probes
|
||||
model_results = check_model_probes()
|
||||
for name, passed, detail in model_results:
|
||||
status = "✅" if passed else "❌"
|
||||
# Check if this is a busy (degraded) verdict
|
||||
if not passed and detail.startswith("busy "):
|
||||
status = "⚠️"
|
||||
degraded.append(name)
|
||||
else:
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
if not passed:
|
||||
# Only set all_pass=False for real failures (not busy)
|
||||
if not passed and not detail.startswith("busy "):
|
||||
all_pass = False
|
||||
|
||||
# Admin key list
|
||||
@@ -308,10 +342,16 @@ def main():
|
||||
|
||||
print("")
|
||||
if all_pass:
|
||||
print("✅ All checks passed")
|
||||
if degraded:
|
||||
print("✅ All checks passed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||
else:
|
||||
print("✅ All checks passed")
|
||||
return 0
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
if degraded:
|
||||
print("❌ Some checks failed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
return 1
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -176,6 +176,11 @@ fi
|
||||
# never contact her former host.
|
||||
|
||||
# ── Platform C: Agent Zero (kagentz) ──
|
||||
# C1: A2A liveness (no credential needed) — probes the container's internal :80/a2a/
|
||||
# C2: A2A response verification (needs LITELLM_KEY) — probes POST /a2a with auth
|
||||
# C3: Public access path (no credential needed) — probes https://kagentz.sysloggh.net/
|
||||
|
||||
# C1: A2A liveness (container-internal probe)
|
||||
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
|
||||
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||
@@ -184,27 +189,53 @@ AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||
if [ "$AZ_A2A_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz A2A server DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||
echo " kagentz C1: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||
else
|
||||
case "$AZ_A2A_CODE" in
|
||||
200|401)
|
||||
echo " kagentz: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||
echo " kagentz C1: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||
echo " kagentz C1: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# C3: Public access path (the captain's point of view)
|
||||
# Probes the public URL that NetBird proxies to the container. 200/302/401 = alive,
|
||||
# 502 = proxy's "upstream refused" page (incident), connection failed = incident.
|
||||
# Never restarts anything — the contract forbids restarting the platform.
|
||||
KAGENTZ_PUBLIC_CODE=$(curl -s -o /dev/null --connect-timeout 10 --max-time 15 \
|
||||
-w '%{http_code}' https://kagentz.sysloggh.net/ 2>/dev/null) || KAGENTZ_PUBLIC_CODE="000"
|
||||
KAGENTZ_PUBLIC_CODE=$(printf '%s' "$KAGENTZ_PUBLIC_CODE" | tr -d '[:space:]')
|
||||
[ -n "$KAGENTZ_PUBLIC_CODE" ] || KAGENTZ_PUBLIC_CODE="000"
|
||||
|
||||
if [ "$KAGENTZ_PUBLIC_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz public URL DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: ❌ public URL down (HTTP 000)" >> "$LOG"
|
||||
elif [ "$KAGENTZ_PUBLIC_CODE" = "502" ]; then
|
||||
notify "🔴" "kagentz public URL 502 (upstream refused)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: ❌ public URL 502 (upstream refused)" >> "$LOG"
|
||||
else
|
||||
case "$KAGENTZ_PUBLIC_CODE" in
|
||||
200|302|401)
|
||||
echo " kagentz C3: ✅ public URL alive (HTTP $KAGENTZ_PUBLIC_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz public URL answered HTTP $KAGENTZ_PUBLIC_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: 🟡 public URL unexpected http=$KAGENTZ_PUBLIC_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# ── Summary ──
|
||||
# The run verdict is non-optimistic: when ISSUES > 0, the run is an INCIDENT.
|
||||
# The lane must quote this Result line verbatim in its status report.
|
||||
if [ "$ISSUES" -eq 0 ]; then
|
||||
if [ "$ZULIP_CRED_OK" -eq 0 ]; then
|
||||
echo " Result: ✅ All healthy" >> "$LOG"
|
||||
else
|
||||
echo " Result: ✅ All healthy" >> "$LOG"
|
||||
fi
|
||||
echo " Result: ✅ 0 issues (all healthy)" >> "$LOG"
|
||||
else
|
||||
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
||||
echo " Result: 🔴 INCIDENT — $ISSUES issue(s) found" >> "$LOG"
|
||||
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
||||
fi
|
||||
|
||||
|
||||
@@ -98,8 +98,8 @@ exit 0
|
||||
"""
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
|
||||
# record every call (including notify) payloads.
|
||||
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||
# kagentz C3 public URL, and record every call (including notify) payloads.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
@@ -109,6 +109,8 @@ case "$*" in
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
*kagentz.sysloggh.net*)
|
||||
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
@@ -121,7 +123,8 @@ def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a_code="401", az_a2a_exit=0):
|
||||
az_a2a_code="401", az_a2a_exit=0,
|
||||
kagentz_public_code="302"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
@@ -154,6 +157,7 @@ def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
@@ -169,8 +173,9 @@ def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
||||
assert "Server: ✅ HTTP 200" in log
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "Tanko: ✅ service=active http=200" in log
|
||||
assert "kagentz: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "Result: ✅ All healthy" in log
|
||||
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||
|
||||
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||
assert proc.stdout == ""
|
||||
@@ -207,8 +212,8 @@ def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
||||
# The rest of the monitor still ran alongside the failing Tanko leg.
|
||||
log = log_path.read_text()
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "kagentz: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "kagentz C1: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
|
||||
|
||||
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||
@@ -216,9 +221,9 @@ def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||
assert "kagentz: ✅ A2A alive" not in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "kagentz C1: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||
assert "kagentz C1: ✅ A2A alive" not in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
assert "kagentz A2A server answered HTTP 500" in proc.stdout
|
||||
|
||||
|
||||
@@ -230,10 +235,10 @@ def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz: ❌ A2A down (HTTP 000)" in log
|
||||
assert "kagentz: ✅ A2A alive" not in log
|
||||
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||
assert "kagentz C1: ✅ A2A alive" not in log
|
||||
assert "unexpected" not in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,209 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Behavioural tests for zulip-monitor.sh kagentz C1/C3 legs and the Result verdict.
|
||||
|
||||
WHY THIS FILE EXISTS: the 2026-09-19 kagentz A2A outage was correctly detected
|
||||
by the monitor (C1 returned 000, the run logged an issue) but the lane's own
|
||||
summarization in ops.status wrote "OK" with the note "A2A server DOWN — expected
|
||||
(no credentials configured)". The optimistic verdict came from the lane, not the
|
||||
script. The fix adds a C3 public-access-path leg and makes the Result line say
|
||||
"INCIDENT" when issues are found, so the lane can quote it verbatim.
|
||||
|
||||
CONTRACT UNDER TEST:
|
||||
* C1 (A2A liveness, no credential): 000 → INCIDENT.
|
||||
* C3 (public access path, no credential): 502 → INCIDENT, 000 → INCIDENT,
|
||||
200/302/401 → alive.
|
||||
* Result verdict: when ISSUES > 0, the log's Result line says "INCIDENT",
|
||||
not just "issues found".
|
||||
* Healthy control: C1 401 + C3 302 → 0 issues, "all healthy".
|
||||
|
||||
HOW: behavioural execution using the sandbox pattern already in this repo
|
||||
(tests/test_mumuni_monitor_removal.py). The sandbox copies the shipped monitor
|
||||
verbatim, rewrites only its LOG constant, and runs it with stub ssh/curl on
|
||||
PATH. Each test asserts from the run's own log/verdict, not from file text.
|
||||
|
||||
Usage: python3 -m pytest tests/test_zulip_kagentz_legs.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import pathlib
|
||||
import stat
|
||||
import subprocess
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||
|
||||
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||
|
||||
|
||||
# ── Stub ssh: answers Tanko and Agent Zero probes by env vars ──────────
|
||||
|
||||
SSH_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub ssh: record the target host, then answer by host + remote command.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||
host=""
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
*@192.168.*) host="${a##*@}" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||
cmd="${*: -1}"
|
||||
case "$host" in
|
||||
192.168.68.15)
|
||||
case "$cmd" in
|
||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||
esac ;;
|
||||
192.168.68.14)
|
||||
case "$cmd" in
|
||||
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "${AZ_A2A_EXIT:-0}" ;;
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
# ── Stub curl: serves Zulip server, Abiba health, and C3 public URL ───
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||
# C3 public URL probe (https://kagentz.sysloggh.net/). Record every call.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
case " $* " in
|
||||
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
*kagentz.sysloggh.net*)
|
||||
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
path.write_text(body)
|
||||
path.chmod(path.stat().st_mode
|
||||
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a_code="401", az_a2a_exit=0,
|
||||
kagentz_public_code="302"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
Everything else — legs, labels, notify logic — is the shipped script.
|
||||
"""
|
||||
sandbox = tmp_path / "sandbox"
|
||||
bindir = sandbox / "bin"
|
||||
record = sandbox / "record"
|
||||
bindir.mkdir(parents=True)
|
||||
record.mkdir()
|
||||
|
||||
_write_exec(bindir / "ssh", SSH_STUB)
|
||||
_write_exec(bindir / "curl", CURL_STUB)
|
||||
|
||||
source = ZULIP_MONITOR.read_text()
|
||||
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||
log_path = sandbox / "zulip-health-monitor.log"
|
||||
script = sandbox / "zulip-monitor.sh"
|
||||
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||
|
||||
env = dict(os.environ)
|
||||
env.update({
|
||||
"PATH": f"{bindir}:{env['PATH']}",
|
||||
"RECORD_DIR": str(record),
|
||||
"TANKO_SVC": tanko_svc,
|
||||
"TANKO_HTTP": tanko_http,
|
||||
"AZ_A2A_CODE": az_a2a_code,
|
||||
"AZ_A2A_EXIT": str(az_a2a_exit),
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
return proc, record, log_path
|
||||
|
||||
|
||||
# ── Required behavioural cases ─────────────────────────────────────────
|
||||
|
||||
def test_c3_502_is_incident(tmp_path):
|
||||
"""C3 public leg returns 502 → the run's verdict is an INCIDENT,
|
||||
the C3 line names the 502, and the run is not summarised as healthy."""
|
||||
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="502")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C3 line names the 502.
|
||||
assert "kagentz C3: ❌ public URL 502 (upstream refused)" in log
|
||||
# The verdict is an INCIDENT, not healthy.
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
# The run is not summarised as healthy.
|
||||
assert "✅ 0 issues" not in log
|
||||
|
||||
|
||||
def test_c3_000_is_incident(tmp_path):
|
||||
"""C3 public leg returns 000 → INCIDENT."""
|
||||
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="000")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C3 line reports the connection failure.
|
||||
assert "kagentz C3: ❌ public URL down (HTTP 000)" in log
|
||||
# The verdict is an INCIDENT.
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
assert "✅ 0 issues" not in log
|
||||
|
||||
|
||||
def test_healthy_control_c1_401_c3_302(tmp_path):
|
||||
"""Healthy control: C1 401 plus C3 302 → 0 issues and a healthy verdict,
|
||||
proving the new leg cannot cry wolf."""
|
||||
proc, record, log_path = _run_monitor(tmp_path,
|
||||
az_a2a_code="401",
|
||||
kagentz_public_code="302")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# Both legs report alive.
|
||||
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||
# Zero issues, healthy verdict.
|
||||
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||
# No INCIDENT.
|
||||
assert "INCIDENT" not in log
|
||||
# No notify fired for kagentz.
|
||||
assert "kagentz public URL" not in proc.stdout
|
||||
assert "kagentz A2A server" not in proc.stdout
|
||||
|
||||
|
||||
def test_c1_000_is_incident(tmp_path):
|
||||
"""C1 returns 000 → INCIDENT, keeping the leg that actually caught this
|
||||
outage covered behaviourally."""
|
||||
proc, record, log_path = _run_monitor(tmp_path,
|
||||
az_a2a_code="000",
|
||||
az_a2a_exit=7,
|
||||
kagentz_public_code="302")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C1 line reports the A2A down.
|
||||
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||
# The verdict is an INCIDENT (even though C3 is healthy).
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
# The notify fired for the A2A down.
|
||||
assert "kagentz A2A server DOWN" in proc.stdout
|
||||
+42
-19
@@ -1,9 +1,9 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-health
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||
version: 3.3.0
|
||||
version: 3.4.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
report_only_agents:
|
||||
@@ -30,7 +30,7 @@ session start.
|
||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||
- **PM2** on localhost for pi process management
|
||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||
- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200`
|
||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||
- **Relay access** via RA-H OS MCP for alert delivery
|
||||
|
||||
@@ -129,10 +129,10 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
|
||||
|
||||
### Liveness rule (scoped)
|
||||
|
||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN.
|
||||
|
||||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above).
|
||||
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
||||
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
||||
|
||||
@@ -469,9 +469,10 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
|
||||
> monitor issued a restart for something that could not start, posting a false
|
||||
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
|
||||
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
|
||||
> liveness only, and a probe must never restart a platform.
|
||||
> liveness/response and public-path access only, and a probe must never restart
|
||||
> a platform.
|
||||
|
||||
**C1: A2A Server Health**
|
||||
**C1: A2A Server Health (no credential needed)**
|
||||
|
||||
```bash
|
||||
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
|
||||
@@ -479,9 +480,9 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
|
||||
```
|
||||
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down (INCIDENT). Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**C2: A2A Response Verification**
|
||||
**C2: A2A Response Verification (requires LITELLM_KEY)**
|
||||
|
||||
```bash
|
||||
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
|
||||
@@ -491,15 +492,30 @@ ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:
|
||||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||
```
|
||||
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set (this is a credential issue, NOT a server-down incident).
|
||||
|
||||
**C3: Public Access Path (no credential needed)**
|
||||
|
||||
```bash
|
||||
# Probes the public URL that NetBird proxies to the agent-zero container.
|
||||
# This is the captain's point of view: if the captain can't reach it, it's down.
|
||||
# 200/302/401 = alive, 502 = proxy's "upstream refused" page (INCIDENT),
|
||||
# connection failed (000) = INCIDENT. Never restarts anything.
|
||||
curl -s -o /dev/null --connect-timeout 10 --max-time 15 -w '%{http_code}' https://kagentz.sysloggh.net/
|
||||
```
|
||||
|
||||
Expected: `200` (Agent Zero login page), `302` (redirect), or `401` (auth-gated) = alive. `502` = NetBird proxy's "upstream refused" page (the container's port 80 is not listening) = INCIDENT. `000` (connection failed) = INCIDENT. Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**Platform C Actions**
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
|
||||
| C1 A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||
| C1 A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||
| C2 A2A returns 401 | Check `LITELLM_KEY` is set (credential issue, NOT server-down) |
|
||||
| C3 public URL returns `502` (upstream refused) | Alert only — the container's port 80 is not listening; investigate the agent-zero container |
|
||||
| C3 public URL returns `000` (connection failed) | Alert only — the public path is down; investigate the NetBird proxy or the container |
|
||||
| C3 public URL returns a status other than `200`/`302`/`401` | Log/report as a warning — reported, never healed on |
|
||||
|
||||
### Step 5: Global Checks
|
||||
|
||||
@@ -513,12 +529,19 @@ If any bot processes >50 bot-originated messages in 15min → warning.
|
||||
|
||||
### Step 6: Compile and Report
|
||||
|
||||
1. Compile all platform checks and severity
|
||||
2. Determine `overall_severity` from worst per-agent severity
|
||||
3. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||
4. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||
5. If any agent critical or >2 degraded: send relay message to user
|
||||
6. Update `last_check` timestamp in `### Maintains` snapshot
|
||||
1. Run `scripts/zulip-monitor.sh` and take its final `Result:` line as the
|
||||
authoritative run verdict. The verdict line is either
|
||||
`Result: ✅ 0 issues (all healthy)` or
|
||||
`Result: 🔴 INCIDENT — N issue(s) found`.
|
||||
2. Quote that `Result:` line verbatim in the status report. When it says
|
||||
`INCIDENT`, the run MUST be reported as an incident — never summarised as
|
||||
OK/healthy and never annotated as "expected".
|
||||
3. Compile all platform checks and severity
|
||||
4. Determine `overall_severity` from worst per-agent severity
|
||||
5. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||
6. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||
7. If any agent critical or >2 degraded: send relay message to user
|
||||
8. Update `last_check` timestamp in `### Maintains` snapshot
|
||||
|
||||
### Restart Debounce
|
||||
|
||||
|
||||
Reference in New Issue
Block a user