Compare commits

..
Author SHA1 Message Date
root dba9cc53d8 fix(litellm): Fix timeout kind reporting + add busy/degraded detection
1. TIMEOUT KIND FIX: run_command returns (1, '', 'TIMEOUT') when its own
   timeout fires. probe_http now checks for this before falling through to
   'curl exit <rc>', so a 30s timeout reports 'timeout after 30s' not
   'curl exit 1'.

2. BUSY/DEGRADED DETECTION: After both model probes fail, check the
   model's host health endpoint (e.g. 192.168.68.8:8080/health for
   gpu-dense). If the host answers 200, report 'busy (completion timed
   out after retry; host healthy 200)' — do NOT fail the run on that
   alone. If the host does not answer, that's a real FAIL.

3. RETRY TIMEOUT INCREASED: Single-host retry timeout raised from 45s to
   90s. Worst-case prefill on a single-slot .8 host is ~76s (observed
   83K-token prompt at 1078 tok/s), so 90s covers it.

New line shapes:
- Busy: 'gpu-dense: busy (completion timed out after retry; host healthy 200)'
- Real failure: 'probe-failed: gpu-dense timeout after 30s then timeout after 90s (2 attempts)'
2026-09-19 11:32:15 +00:00
21 changed files with 80 additions and 1200 deletions
-12
View File
@@ -71,18 +71,6 @@ jobs:
git fetch origin "${{ gitea.ref }}" --depth=50
git checkout "${{ gitea.sha }}"
- name: Committed-credential scan (secret guard)
run: |
# Fails the build on a credential-shaped string in the tree. Patterns
# live in scripts/secret-patterns.tsv; the only tolerated literal
# examples are in scripts/secret-allowlist.tsv, each with a reason.
# Do not turn this into a warning: a warning in a stream nobody reads
# is how six live credentials sat in this repo for weeks.
bash scripts/secret-scan.sh
- name: Secret guard self-test
run: bash tests/test_secret_scan.sh
- name: Structure + regression + consistency lint
run: bash scripts/prose-lint.sh
-6
View File
@@ -51,12 +51,6 @@ Two incidents taught us this:
- `/grafana/` nginx route — was reverted Jul 2, must not reappear
- `CT 122` or `CT 123` as CT ID labels — don't exist in the cluster
- These rules are hardcoded in `scripts/prose-lint.sh`
- **Committed-credential guard:** `scripts/secret-scan.sh` FAILS the build on
credential-shaped strings (patterns in `scripts/secret-patterns.tsv`, prose
included). Tolerated literals are listed one-per-example with a reason in
`scripts/secret-allowlist.tsv`; never allowlist a live credential. It runs in
the CI lint job, in `scripts/prose-lint.sh`, and via
`bash scripts/secret-scan.sh --staged` before committing.
### Stage 3 — AI Review
- Diff is sent to `syslog-auto` model via LiteLLM
+6 -16
View File
@@ -114,22 +114,12 @@ def audit(path):
)
# --- Rule 5: Main Config Base URL ---
# Canonical internal base (hermes-key-enforcement.prose.md:19/38/57/84/91/96)
# and public base (serves /v1 only, per 2026-09-19 probe from CT 116).
# The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only.
# FAIL anything else (do not widen to accept any path ending in /v1).
# Internal /v1 is non-canonical but working (authenticated via nginx), so WARN not FAIL.
canonical_internal = "http://192.168.68.116/litellm/v1"
public_host = "https://litellm.sysloggh.net/v1"
non_canonical_internal = "http://192.168.68.116/v1"
allowed_bases = (canonical_internal, public_host)
actual_base = model.get("base_url")
if actual_base in allowed_bases:
check(True, "Rule 5", f"model.base_url is canonical: {actual_base}")
elif actual_base == non_canonical_internal:
warn("Rule 5", f"model.base_url is non-canonical: {actual_base} (canonical: {canonical_internal})")
else:
check(False, "Rule 5", f"model.base_url must be one of {allowed_bases} (got {actual_base!r})")
expected_base = "http://192.168.68.116/v1"
check(
model.get("base_url") == expected_base,
"Rule 5",
f"model.base_url must be {expected_base} (got {model.get('base_url')!r}) — /v1 not /litellm/v1",
)
# --- Rule 6: max_tokens Is Required ---
check(
+1 -55
View File
@@ -38,7 +38,6 @@ index:
by_category:
compliance:
- hermes-key-enforcement
- litellm-api-keys
- hermes-config-template
- hermes-agent-baseline
monitoring:
@@ -102,7 +101,6 @@ index:
proxmox:
- proxmox-monitor
litellm:
- litellm-api-keys
- litellm-health
- litellm-self-heal
memory:
@@ -630,7 +628,7 @@ contracts:
sensitivity: high
status: active
owner: abiba
version: 3.4.0
version: 3.3.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
@@ -1871,58 +1869,6 @@ contracts:
drift_alerts: []
# Koby Report-Only Registry (2026-08-17 — Captain)
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
- name: litellm-api-keys
file: litellm-api-keys.prose.md
kind: function
category: compliance
sensitivity: critical
status: active
owner: abiba
version: 1.1.0
trigger:
type: on_demand
cadence: null
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
cron_job_id: null
execution:
agent: abiba
timeout: 120
requires: []
protocol:
- Load contract from prose-contracts/main
- Retrieve master key from Infisical (project=infrastructure env=production)
- Read live key-scoped model roster from CT 116 /v1/models
- Create/rotate/verify the requested agent key with an EXPLICIT models list
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
verification:
postconditions:
- check: standard agent key is local-only
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
expect: 0 cloud models
- check: key exists with correct alias
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
expect: 200 with matching alias
artifact: key creation/rotation report
receipt:
format: json
storage: ~/.hermes/runs/litellm-api-keys/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
- ops
koby_report_only: true
koby_host: "CT 111 (tdunna)"
koby_ip: ".129"
+4 -23
View File
@@ -16,26 +16,7 @@ author: Abiba (pi agent)
## Rule (One Sentence)
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
## Model Access Tiers (2026-09-20)
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
| Tier | Models | Who gets it |
|------|--------|-------------|
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
**Rules:**
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
2. Agent keys MUST carry an **explicit local-only** `models` list.
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
4. The master key always bypasses scoping — it is admin-only, never for inference.
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with authenticated path `http://192.168.68.116/litellm/v1/responses` — hardcoded keys AND unauthenticated `/v1` direct access are both forbidden.**
## Scope
@@ -57,8 +38,8 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` non-canonical but working).
The public host `https://litellm.sysloggh.net` serves `/v1` ONLY (404 on `/litellm/v1`).
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid).
Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**.
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
@@ -132,7 +113,7 @@ model:
model:
provider: harness
base_url: http://192.168.68.116/v1 # ← NON-CANONICAL but WORKING (authenticated via nginx, WARN not FAIL)
base_url: http://192.168.68.116/v1 # ← RULE VIOLATION: unauthenticated path
api_key_env: LITELLM_API_KEY
```
+1 -1
View File
@@ -759,4 +759,4 @@ which was kill+nohup outside systemd) are banned by policy.
| Script | Why Disabled |
|--------|-------------|
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
| `zulip-monitor.sh` (Abiba) | Standalone cron replaced by agent-health-check.py + PM2 auto-restart; the script itself remains active as the execution step of the `zulip-health` responsibility contract (see `zulip-health.prose.md`) |
| `zulip-monitor.sh` (Abiba) | Replaced by agent-health-check.py + PM2 auto-restart |
-64
View File
@@ -71,10 +71,6 @@ description: >
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
In LiteLLM Community both silently grant access to EVERY model, including cloud.
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
- Return the new key
5. **If action == "rotate"**:
@@ -91,66 +87,6 @@ description: >
- Confirm key alias matches agent_name in LiteLLM key list
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
## Cloud Provider Consolidation (2026-09-20)
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
### Provider map (per-account namespacing)
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
stay separate:
| Prefix | Upstream | Auth | Vault secret |
|--------|----------|------|--------------|
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
### Access tiers (MUST be enforced per key)
| Tier | Model names | Granted to |
|------|-------------|------------|
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
### Creating the cloud-enabled key
```bash
# ALWAYS read the live roster first (key-scoped):
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
| jq -r '.data[].id'
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
```
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
any `<prefix>/` cloud model.
### Adding a new cloud provider
1. Add the upstream key to Infisical `infrastructure/production/root`.
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
3. Restart the `harness-litellm` container.
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
captain approval.
5. Update this table and the access-tier section.
## Production Vault Access Process (canonical, 2026-07-17)
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
+10 -20
View File
@@ -2,10 +2,10 @@
kind: responsibility
name: pm2-self-heal
description: >
Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
and auto-restarts any that are stopped or errored. Logs every action to
Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary).
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
PM2 process health check for abiba-telegram, abiba-zulip, gitea-runner, and zulip-watchdog.
gpu-monitor is systemd-managed (gpu-monitor.service), NOT PM2.
gpu-watchdog is decommissioned and folded into gpu-monitor.service.
gitea-runner is KEPT. abiba-zulip is KEPT (online for days).
---
## Maintains
@@ -16,8 +16,6 @@ description: >
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed.
## Continuity
@@ -42,7 +40,7 @@ description: >
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
spoton incident). Alerts include the restart count.
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton
four processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog). The spoton
reference above is historical context for the crash-loop guard, not a live process.
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
@@ -53,25 +51,17 @@ description: >
## Execution
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
2. **Check abiba-telegram** (safe to auto-restart):
2. **Check abiba-telegram**:
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
- If restarts > 5 → alert owner
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
- If status is "online" → pass, log restarts count
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
- If restarts > 5 in last hour → alert owner with full diagnostics
4. **Check gitea-runner**:
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 5 → alert owner
5. **Check zulip-watchdog**:
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 5 → alert owner
6. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
8. **Wait 5 min** → repeat from step 1
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
6. **Wait 5 min** → repeat from step 1
## Example Output (when healthy)
+7 -21
View File
@@ -22,12 +22,10 @@ AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TO
ZULIP_SITE = "https://chat.sysloggh.net"
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
# Note: /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential).
# No Zulip API key is required for this call. If a future leg genuinely needs abiba-bot's key,
# it must prove it with a 200 from /api/v1/users/me as abiba-bot and label itself degraded when it cannot.
# Never fall back to the vault's shared ZULIP_API_KEY.
ZULIP_AUTH = None
DEGRADED_LEGS = []
ZULIP_API_KEY = os.environ.get("ZULIP_API_KEY", "")
if not ZULIP_API_KEY:
raise SystemExit("ZULIP_API_KEY not set — refusing to run with no credential")
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_API_KEY}"
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
LITELLM_BACKEND = "192.168.68.116"
@@ -160,7 +158,7 @@ def collect():
# ── Storage ──
storages = pve_get("/api2/json/nodes/storepve/storage")
report["storage"] = []
for s in (storages or []):
for s in storages:
total = s.get("total",0) or 1
used = s.get("used",0)
pct = used/total*100
@@ -675,9 +673,8 @@ def send_email(html_content, subject_prefix=""):
try:
EMAIL_PASSWORD = os.environ.get("EMAIL_PASSWORD") or os.environ.get("SMTP_PASSWORD") or os.environ.get("MAIL_PASSWORD")
if not EMAIL_PASSWORD:
print(" ⚠️ Degraded leg: credential-missing: EMAIL_PASSWORD (or SMTP_PASSWORD/MAIL_PASSWORD)", file=sys.stderr)
DEGRADED_LEGS.append("credential-missing: EMAIL_PASSWORD")
return True, "✅ Email leg degraded (no credential) — report still produced"
print("EMAIL_PASSWORD not set — refusing to send email", file=sys.stderr)
sys.exit(1)
GMAIL_EMAIL = "jtabiri@gmail.com"
server = smtplib.SMTP("smtp.gmail.com", 587)
@@ -716,17 +713,6 @@ if __name__ == "__main__":
print(f" {msg}")
# Show summary
if DEGRADED_LEGS:
print(f"\n⚠️ Degraded legs ({len(DEGRADED_LEGS)}):")
for leg in DEGRADED_LEGS:
print(f" - {leg}")
else:
print("\n✅ All legs fully credentialed")
# A failed send must exit non-zero; a degraded leg (no credential) must stay exit 0
if not ok:
sys.exit(1)
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
print(f"\n📋 Summary:")
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
+5 -18
View File
@@ -163,7 +163,7 @@ def check_model_probes():
# Both attempts failed - check host health to distinguish busy from down
host_healthy, host_detail = check_host_health(host_ip)
if host_healthy:
results.append((model, False, "busy (completion timed out after retry; " + host_detail + ")"))
results.append((model, False, "busy (completion timed out after retry; host healthy " + host_detail + ")"))
else:
# Host unreachable - report both kinds
if first_kind:
@@ -288,7 +288,6 @@ def main():
print("")
all_pass = True
degraded = [] # Track degraded (busy) checks
# Run all checks
checks = [
@@ -308,15 +307,9 @@ def main():
# Model probes
model_results = check_model_probes()
for name, passed, detail in model_results:
# Check if this is a busy (degraded) verdict
if not passed and detail.startswith("busy "):
status = "⚠️"
degraded.append(name)
else:
status = "✅" if passed else "❌"
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
# Only set all_pass=False for real failures (not busy)
if not passed and not detail.startswith("busy "):
if not passed:
all_pass = False
# Admin key list
@@ -342,16 +335,10 @@ def main():
print("")
if all_pass:
if degraded:
print("✅ All checks passed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
else:
print("✅ All checks passed")
print("✅ All checks passed")
return 0
else:
if degraded:
print("❌ Some checks failed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
else:
print("❌ Some checks failed")
print("❌ Some checks failed")
return 1
if __name__ == "__main__":
+2 -68
View File
@@ -1,7 +1,7 @@
#!/bin/bash
# pm2-self-heal — hourly PM2 process check
# Part of the pm2-self-heal prose contract
# Alerts via Telegram (primary) and Zulip DM (secondary, abiba-zulip restored 2026-08-03)
# Alerts via Telegram (abiba-zulip decommissioned 2026-07-04)
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
@@ -46,75 +46,9 @@ if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
fi
fi
# Check abiba-zulip (live Zulip bridge, heartbeating)
ZULIP_LINE=$(echo "$STATUS" | grep "abiba-zulip")
ZULIP_STATUS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $10}' | xargs)
ZULIP_RESTARTS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $9}' | xargs)
if [ "$ZULIP_STATUS" != "online" ]; then
pm2 restart abiba-zulip > /dev/null 2>&1
sleep 3
ZULIP_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-zulip")
ZULIP_STATUS2=$(echo "$ZULIP_LINE2" | awk -F'│' '{print $10}' | xargs)
if [ "$ZULIP_STATUS2" = "online" ]; then
msg="⚠️ abiba-zulip was **$ZULIP_STATUS** → restarted to online"
ALERTS="${ALERTS}${msg}\n"
else
msg="🚨 abiba-zulip **failed restart** (was $ZULIP_STATUS, still $ZULIP_STATUS2)"
ALERTS="${ALERTS}${msg}\n"
fi
elif [ "$ZULIP_RESTARTS" -gt 5 ]; then
msg="⚠️ abiba-zulip has **$ZULIP_RESTARTS** restarts (high count)"
ALERTS="${ALERTS}${msg}\n"
fi
# Check gitea-runner
GITEA_LINE=$(echo "$STATUS" | grep "gitea-runner")
GITEA_STATUS=$(echo "$GITEA_LINE" | awk -F'│' '{print $10}' | xargs)
GITEA_RESTARTS=$(echo "$GITEA_LINE" | awk -F'│' '{print $9}' | xargs)
if [ "$GITEA_STATUS" != "online" ]; then
pm2 restart gitea-runner > /dev/null 2>&1
sleep 3
GITEA_LINE2=$(pm2 status --no-color 2>/dev/null | grep "gitea-runner")
GITEA_STATUS2=$(echo "$GITEA_LINE2" | awk -F'│' '{print $10}' | xargs)
if [ "$GITEA_STATUS2" = "online" ]; then
msg="⚠️ gitea-runner was **$GITEA_STATUS** → restarted to online"
ALERTS="${ALERTS}${msg}\n"
else
msg="🚨 gitea-runner **failed restart** (was $GITEA_STATUS, still $GITEA_STATUS2)"
ALERTS="${ALERTS}${msg}\n"
fi
elif [ "$GITEA_RESTARTS" -gt 5 ]; then
msg="⚠️ gitea-runner has **$GITEA_RESTARTS** restarts (high count)"
ALERTS="${ALERTS}${msg}\n"
fi
# Check zulip-watchdog
WATCHDOG_LINE=$(echo "$STATUS" | grep "zulip-watchdog")
WATCHDOG_STATUS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $10}' | xargs)
WATCHDOG_RESTARTS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $9}' | xargs)
if [ "$WATCHDOG_STATUS" != "online" ]; then
pm2 restart zulip-watchdog > /dev/null 2>&1
sleep 3
WATCHDOG_LINE2=$(pm2 status --no-color 2>/dev/null | grep "zulip-watchdog")
WATCHDOG_STATUS2=$(echo "$WATCHDOG_LINE2" | awk -F'│' '{print $10}' | xargs)
if [ "$WATCHDOG_STATUS2" = "online" ]; then
msg="⚠️ zulip-watchdog was **$WATCHDOG_STATUS** → restarted to online"
ALERTS="${ALERTS}${msg}\n"
else
msg="🚨 zulip-watchdog **failed restart** (was $WATCHDOG_STATUS, still $WATCHDOG_STATUS2)"
ALERTS="${ALERTS}${msg}\n"
fi
elif [ "$WATCHDOG_RESTARTS" -gt 5 ]; then
msg="⚠️ zulip-watchdog has **$WATCHDOG_RESTARTS** restarts (high count)"
ALERTS="${ALERTS}${msg}\n"
fi
# Log check
{
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS zulip=$ZULIP_STATUS gitea=$GITEA_STATUS watchdog=$WATCHDOG_STATUS alerts=${ALERTS:+yes}"
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS alerts=${ALERTS:+yes}"
[ -n "$ALERTS" ] && echo "$ALERTS"
} >> "$LOG"
+1 -16
View File
@@ -135,22 +135,7 @@ fi
echo " Cross-contract: $WARNINGS total warnings across all checks"
# ── 4. Committed-credential scan ──
# The 2026-09-17 purge removed six live credentials that had sat in .md prose
# and scripts for weeks. This step makes that class of commit FAIL the gate
# instead of printing a warning. Patterns: scripts/secret-patterns.tsv.
# Only deliberate synthetic examples may be listed in scripts/secret-allowlist.tsv,
# each with a reason. Run `bash scripts/secret-scan.sh --staged` before committing.
echo ""
echo "── 4. Secret scan (committed credentials) ──"
if bash scripts/secret-scan.sh; then
echo " ✅ No committed credentials"
else
echo " ❌ COMMITTED CREDENTIAL DETECTED"
FAILED=1
fi
# ── 5. Summary ──
# ── 4. Summary ──
echo ""
echo "═══════════════════════════════════"
if [ $FAILED -eq 1 ]; then
-55
View File
@@ -1,55 +0,0 @@
# secret-allowlist.tsv — exceptions for scripts/secret-scan.sh, every entry with a reason.
#
# Format: <rule-id|*><TAB><path-glob><TAB><literal-substring><TAB><reason>
# Blank lines and lines whose first field starts with '#' are ignored.
# A finding is suppressed only when ALL THREE of rule, path and literal match:
# * the rule id equals the finding's rule id, or is '*'
# * the finding's repo-relative path matches <path-glob> (bash glob)
# * the finding's line contains <literal-substring> verbatim
# An entry whose reason is empty is a hard error (exit 2) — no silent exceptions.
#
# RULE: never allowlist a live credential, and never broaden an entry (rule '*',
# a wide path glob, or a short generic literal) just to silence a finding.
# If the finding is real, remove the credential from the file.
#
# Entries are one per deliberate synthetic example, so the file reads as an
# audit trail of reviewed exceptions rather than a list of things to ignore.
# Rule '*' is used only where the same literal is matched by more than one rule.
#
# ── The 2026-09-17 purge placeholders ─────────────────────────────────────
# PR #112 replaced six live credentials with `«vault: <project>/<env> <SECRET>»`
# references. Those references are safe by construction (they name where the
# secret is read from), but they are listed here explicitly rather than being
# filtered by a general "vault" rule, so a new occurrence still needs a
# deliberate, reasoned entry.
secret-assign litellm-api-keys.prose.md MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY» 2026-09-17 purge: replaced the live Mumuni LiteLLM key with its vault reference; no literal credential.
secret-assign litellm-api-keys.prose.md MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY» 2026-09-17 purge: replaced the live Mumuni Zulip key with its vault reference; no literal credential.
* infrastructure-control.prose.md PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
cred-prose infrastructure-control.prose.md Admin credentials: 2026-09-17 purge: the Stirling admin user/password are two `«vault: ...»` references; no literal credential.
* scripts/daily-infra-report.py PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
secret-assign stirling-pdf-agent-access.prose.md «vault: infrastructure/production STIRLING_API_KEY» 2026-09-17 purge: Stirling PDF API key is read from the vault; the curl example only names the vault path.
bearer-token agent-zero-fix-summary.md «vault: agents/production OPENROUTER_API_KEY» 2026-09-17 purge: OpenRouter key is read from the vault; the example curl only names the vault path.
# ── Deliberate synthetic examples in contracts (not from the purge) ───────
# These exist to teach the rule they illustrate. They are listed here so the
# guard is never taught to skip the words "synthetic"/"example" — a fabricated
# example is always an explicit exception, never a pattern-level exemption.
* hermes-key-enforcement.prose.md sk-synthetic-external-example Rule 15 illustration of a hardcoded external key that is tolerated; fabricated, never a live key.
* hermes-key-enforcement.prose.md sk-synthetic-example-12345 Rule 15 illustration of a forbidden hardcoded key; fabricated, never a live key.
openai-key hermes-key-enforcement.prose.md sk-synthetic-litellm- Fabricated key name inside a `grep 'LITELLM_API_KEY=...'` example; not a live key.
secret-assign hermes-key-enforcement.prose.md sk-NEW_KEY Placeholder standing for the rotated key in an `infisical secrets set` command; not a literal key.
openrouter-key agent-zero-openrouter-key.prose.md sk-or-v1-synthetic Synthetic key prefix in the contract's example response; the real key is read from the vault.
openai-key litellm-api-keys.prose.md sk-synthetic-tanko-example Fabricated key name in migration history prose; not a live key.
openai-key litellm-self-heal.prose.md sk-syslog-local-master-key Deprecated local LiteLLM master key name documented as no-live-usage; kept for history, not a usable credential.
# ── Redacted evidence, not a credential ──────────────────────────────────
secret-assign docs/probe-drift-round2-evidence.md =sk-... Probe evidence records redacted key trailers (`sk-...x6uw`); the usable part of the key is not present.
# ── tests/test_secret_scan.sh fixtures ───────────────────────────────────
# The self-test plants these fabricated values into a TEMP tree, whose path no
# entry here covers, so each still fails the guard when planted (see the test's
# "... fails the guard" cases). They are listed only so the repo-wide scan of
# the test file itself stays quiet.
* tests/test_secret_scan.sh sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef Self-test fixture: fabricated OpenRouter-shaped key written to a temp tree; the guard must fail on it there.
bearer-token tests/test_secret_scan.sh Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb Self-test fixture: fabricated Bearer token written to a temp tree; the guard must fail on it there.
proxmox-token tests/test_secret_scan.sh PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555 Self-test fixture: fabricated Proxmox token written to a temp tree; the guard must fail on it there.
private-key tests/test_secret_scan.sh -----BEGIN OPENSSH PRIVATE KEY----- Self-test fixture: fabricated PEM banner written to a temp tree; the guard must fail on it there.
cred-prose tests/test_secret_scan.sh Admin credentials: Self-test fixture: fabricated prose credential line written to a temp tree; the guard must fail on it there.
secret-assign tests/test_secret_scan.sh DB_PASSWORD=correct-horse-battery-staple Self-test fixture: fabricated password assignment written to a temp tree; the guard must fail on it there.
Can't render this file because it contains an unexpected character in line 23 and column 25.
-22
View File
@@ -1,22 +0,0 @@
# secret-patterns.tsv — checked-in pattern list for scripts/secret-scan.sh
#
# Format: <rule-id><TAB><POSIX ERE><TAB><description><TAB><check>
# Blank lines and lines whose first field starts with '#' are ignored.
# <check> is optional; the only value today is "value", which tells the scanner
# to run the matched value through its inert-value classifier (see
# value_is_inert in secret-scan.sh) so bare identifiers, env refs and dotted
# code access are not reported as credentials. Omit the column to report every
# regex hit.
# Matching is case-insensitive, so `API_KEY` and `api_key` both count.
#
# Add a rule here, never inline in secret-scan.sh: this file is the single
# auditable list of what the guard considers credential-shaped.
openai-key \bsk-[A-Za-z0-9_-]{16,} OpenAI/LiteLLM-style "sk-" secret key (also hyphenated sk-proj- keys)
openrouter-key \bsk-or-v1-[A-Za-z0-9_-]{8,} OpenRouter API key
stripe-live-key \bsk_live_[A-Za-z0-9]{8,} Stripe live secret key
proxmox-token PVEAPIToken=[^[:space:]"']+ Proxmox API token literal
bearer-token bearer[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) literal Bearer token (http header or prose)
auth-header authorization:[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) Authorization header carrying a raw literal value
private-key -----BEGIN [A-Z ]*PRIVATE KEY----- PEM private key block
cred-prose credentials?[[:space:]]*[:=][[:space:]]*[^[:space:]] prose credential line carrying a value
secret-assign (api[_-]?key|apikey|passwd|password|secret|token)s?["']?[[:space:]]*[:=][[:space:]]*["']?(«.{3,}»|[A-Za-z0-9_./+=-]{8,}) credential assignment carrying a literal value value
Can't render this file because it contains an unexpected character in line 5 and column 48.
-269
View File
@@ -1,269 +0,0 @@
#!/usr/bin/env bash
# secret-scan.sh — commit-time secret guard. FAILS (exit 1) on a credential-shaped
# string, so a build cannot go green with a credential committed to it.
#
# Usage:
# scripts/secret-scan.sh # scan the whole git-tracked tree (default)
# scripts/secret-scan.sh --tree
# scripts/secret-scan.sh --path DIR # scan an arbitrary directory (git not required)
# scripts/secret-scan.sh --staged # scan added lines in the index (pre-commit)
# scripts/secret-scan.sh --diff REF # scan added lines since REF (e.g. origin/master)
# --quiet only print the verdict and findings, no per-mode banner
#
# Exit codes: 0 clean, 1 credential found, 2 usage/config error.
#
# Patterns live in scripts/secret-patterns.tsv
# Exceptions live in scripts/secret-allowlist.tsv (every entry carries a reason;
# a missing reason is a hard error, so the guard fails closed).
#
# Dependencies are deliberately bash + coreutils + grep + sed/awk + git. The
# Gitea Actions runner executes job steps INSIDE the runner container, which
# has no node and no python by default: keep this script free of both.
set -uo pipefail
SELF_DIR=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
ROOT=$(cd -- "$SELF_DIR/.." && pwd)
PATTERNS_FILE="$SELF_DIR/secret-patterns.tsv"
ALLOWLIST_FILE="$SELF_DIR/secret-allowlist.tsv"
# The guard's own definition files are not scannable content: the pattern list
# necessarily contains the pattern text, and the allowlist necessarily contains
# the allowed literals. Narrow, exact-path exclusion — not a wildcard.
SELF_FILES=(
"scripts/secret-scan.sh"
"scripts/secret-patterns.tsv"
"scripts/secret-allowlist.tsv"
)
MODE="tree"
PATH_DIR=""
DIFF_REF=""
QUIET=0
usage() {
sed -n '2,20p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'
exit 2
}
while [ $# -gt 0 ]; do
case "$1" in
--tree) MODE="tree" ;;
--path) MODE="path"; PATH_DIR="${2:-}"; shift ;;
--staged) MODE="staged" ;;
--diff) MODE="diff"; DIFF_REF="${2:-}"; shift ;;
--quiet) QUIET=1 ;;
-h|--help) usage ;;
*) echo "secret-scan: unknown argument '$1'" >&2; usage ;;
esac
shift
done
[ -f "$PATTERNS_FILE" ] || { echo "secret-scan: missing $PATTERNS_FILE" >&2; exit 2; }
[ -f "$ALLOWLIST_FILE" ] || { echo "secret-scan: missing $ALLOWLIST_FILE" >&2; exit 2; }
if [ "$MODE" = "path" ] && [ -z "$PATH_DIR" ]; then
echo "secret-scan: --path needs a directory" >&2; exit 2
fi
if [ "$MODE" = "diff" ] && [ -z "$DIFF_REF" ]; then
echo "secret-scan: --diff needs a base ref" >&2; exit 2
fi
# ── Load patterns ──────────────────────────────────────────────────────────
RULE_IDS=()
RULE_RES=()
RULE_DESCS=()
RULE_CHECKS=()
COMBINED=""
while IFS=$'\t' read -r id re desc check; do
case "$id" in ''|'#'*) continue ;; esac
[ -n "$re" ] || continue
RULE_IDS+=("$id"); RULE_RES+=("$re"); RULE_DESCS+=("$desc"); RULE_CHECKS+=("${check:-}")
if [ -z "$COMBINED" ]; then COMBINED="($re)"; else COMBINED="$COMBINED|($re)"; fi
done < "$PATTERNS_FILE"
if [ "${#RULE_IDS[@]}" -eq 0 ]; then
echo "secret-scan: no patterns loaded from $PATTERNS_FILE" >&2; exit 2
fi
# ── Load allowlist (fails closed on a missing reason) ──────────────────────
AL_RULES=()
AL_GLOBS=()
AL_LITS=()
AL_REASONS=()
AL_LINENO=0
while IFS=$'\t' read -r rule glob lit reason; do
AL_LINENO=$((AL_LINENO + 1))
case "$rule" in ''|'#'*) continue ;; esac
if [ -z "$glob" ] || [ -z "$lit" ] || [ -z "$reason" ]; then
echo "secret-scan: ❌ $ALLOWLIST_FILE:$AL_LINENO — allowlist entry needs <rule> <path-glob> <literal> <reason>; reason-based exceptions only, refusing to run" >&2
exit 2
fi
AL_RULES+=("$rule"); AL_GLOBS+=("$glob"); AL_LITS+=("$lit"); AL_REASONS+=("$reason")
done < "$ALLOWLIST_FILE"
# nocasematch is toggled only around the regex test; path globs must stay
# case-sensitive, so it is never left on.
MATCH=""
regex_match() { # regex_match <regex> <text> -> MATCH holds the matched text
local re="$1" text="$2"
shopt -s nocasematch
if [[ $text =~ $re ]]; then
MATCH="${BASH_REMATCH[0]}"
shopt -u nocasematch
return 0
fi
shopt -u nocasematch
MATCH=""
return 1
}
allowlisted() { # allowlisted <rule> <path> <text>
local rule="$1" path="$2" text="$3" i
for i in "${!AL_RULES[@]}"; do
[ "${AL_RULES[$i]}" = "$rule" ] || [ "${AL_RULES[$i]}" = "*" ] || continue
# The unquoted RHS is deliberate: <path-glob> is a bash glob, not a literal.
# shellcheck disable=SC2053
[[ $path == ${AL_GLOBS[$i]} ]] || continue
[[ $text == *"${AL_LITS[$i]}"* ]] || continue
return 0
done
return 1
}
mask_value() { # mask_value <text> <match> — never echo a credential to logs.
# Print only the part of the line BEFORE the match, then <redacted>: the match
# itself and everything after it (which may include a value the rule's regex
# stopped short of, e.g. `credentials:` followed by a backticked password) is
# never written to stdout.
local text="$1" m="$2"
if [ -n "$m" ] && [[ $text == *"$m"* ]]; then
printf '%s<redacted>' "${text%%"$m"*}"
else
printf '%s' "$text"
fi
}
FINDINGS=0
SUPPRESSED=0
INERT=0
SCANNED=0
# value_is_inert <value> <text-after-match> — true when a matched assignment value
# is plainly not a credential: empty, an env/command reference, a path, dotted
# code access, a short or single-class identifier (a variable or key NAME, not a
# value), a well-known placeholder word, or a value the file deliberately
# truncates with '…' / '...' (a redacted prefix is not a usable credential).
# Deliberately does NOT know the words "synthetic" or "example": a fabricated
# example must be an explicit allowlist entry.
value_is_inert() {
local v="$1" rest="$2"
case "$rest" in '…'*|'...'*) return 0 ;; esac
v="${v%\"}"; v="${v#\"}"; v="${v%\'}"; v="${v#\'}"
case "$v" in
''|\$*|\{*|'<'*|'%'*|'('*|'/'*|'\\'*) return 0 ;;
not-needed|no-key-required|none|null|true|false|redacted|placeholder|example|dummy|changeme|change-me|your-key|your_key|key|token|secret|password) return 0 ;;
esac
# dotted code access: os.environ.get / process.env.ZULIP_API_KEY / cfg.a
if [[ $v =~ ^[a-z_][a-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)+$ ]]; then return 0; fi
# bare identifier (no punctuation beyond _): a NAME, not a value. A real
# secret in this shape is long and mixes letters with digits.
if [[ $v =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
[ "${#v}" -lt 20 ] && return 0
[[ $v =~ [0-9] ]] || return 0
return 1
fi
return 1
}
report_finding() { # report_finding <path> <line> <text>
local path="$1" line="$2" text="$3" i val
for i in "${!RULE_IDS[@]}"; do
regex_match "${RULE_RES[$i]}" "$text" || continue
SCANNED=$((SCANNED + 1))
if [ "${RULE_CHECKS[$i]}" = "value" ]; then
val="${MATCH#*[:=]}"
val="${val# }"
if value_is_inert "$val" "${text#*"$MATCH"}"; then
INERT=$((INERT + 1))
continue
fi
fi
if allowlisted "${RULE_IDS[$i]}" "$path" "$text"; then
SUPPRESSED=$((SUPPRESSED + 1))
continue
fi
FINDINGS=$((FINDINGS + 1))
printf ' ❌ %s:%s [%s] %s\n' "$path" "$line" "${RULE_IDS[$i]}" "${RULE_DESCS[$i]}"
printf ' | %s\n' "$(mask_value "$text" "$MATCH")"
done
}
self_excluded() { # self_excluded <repo-relative-path>
local p="$1" s
for s in "${SELF_FILES[@]}"; do
[ "$p" = "$s" ] && return 0
done
return 1
}
# ── Collect candidate lines and scan them ─────────────────────────────────
if [ "$MODE" = "tree" ] || [ "$MODE" = "path" ]; then
if [ "$MODE" = "tree" ]; then
BASE="$ROOT"
git -C "$BASE" rev-parse --git-dir >/dev/null 2>&1 || { echo "secret-scan: --tree needs a git checkout (use --path DIR)" >&2; exit 2; }
mapfile -d '' candidate < <(git -C "$BASE" ls-files -z 2>/dev/null)
if [ "${#candidate[@]}" -eq 0 ]; then
echo "secret-scan: ❌ no tracked files — refusing to report clean" >&2; exit 2
fi
else
BASE=$(cd -- "$PATH_DIR" 2>/dev/null && pwd) || { echo "secret-scan: --path '$PATH_DIR' is not a directory" >&2; exit 2; }
mapfile -t candidate < <(cd -- "$BASE" && find . -type f -not -path './.git/*' | sed 's|^\./||')
if [ "${#candidate[@]}" -eq 0 ]; then
echo "secret-scan: ❌ no files under $BASE — refusing to report clean" >&2; exit 2
fi
fi
[ "$QUIET" -eq 1 ] || echo "── secret scan ($MODE): ${#candidate[@]} files under $BASE ──"
for rel in "${candidate[@]}"; do
[ -f "$BASE/$rel" ] || continue
self_excluded "$rel" && continue
while IFS= read -r hit; do
[ -n "$hit" ] || continue
report_finding "$rel" "${hit%%:*}" "${hit#*:}"
done < <(grep -nEIi -e "$COMBINED" "$BASE/$rel" 2>/dev/null || true)
done
else
# --staged / --diff: only ADDED lines, with the post-change line number.
if [ "$MODE" = "staged" ]; then
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines in the index ──"
DIFF_TEXT=$(git -C "$ROOT" diff --cached --unified=0 --no-color -- . 2>/dev/null)
else
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines since $DIFF_REF ──"
DIFF_TEXT=$(git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"...HEAD 2>/dev/null \
|| git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"..HEAD 2>/dev/null)
fi
if [ -z "$DIFF_TEXT" ]; then
[ "$QUIET" -eq 1 ] || echo " (no added lines)"
fi
while IFS=$'\t' read -r rel line text; do
[ -n "$rel" ] || continue
self_excluded "$rel" && continue
report_finding "$rel" "$line" "$text"
done < <(printf '%s\n' "$DIFF_TEXT" | awk '
/^\+\+\+ / { f=$2; sub(/^b\//,"",f); next }
/^@@ / { if (match($0, /\+[0-9]+/)) ln=substr($0, RSTART+1, RLENGTH-1)+0; next }
(/^\+/ && !/^\+\+\+/) { print f "\t" ln "\t" substr($0,2); ln++; next }
')
fi
# ── Verdict ────────────────────────────────────────────────────────────────
if [ "$FINDINGS" -gt 0 ]; then
echo ""
echo "❌ SECRET SCAN FAILED — $FINDINGS credential-shaped string(s) in ${MODE} content."
echo " Fix: remove the credential and read it from the vault/env."
echo " Only a deliberate synthetic example may be added to scripts/secret-allowlist.tsv,"
echo " one entry per file/rule/literal, with a reason. Never allowlist a live credential."
exit 1
fi
echo "✅ secret scan clean (${MODE}; ${SUPPRESSED} allowlisted exception(s), ${INERT} inert value(s) ignored)"
exit 0
+9 -40
View File
@@ -176,11 +176,6 @@ fi
# never contact her former host.
# ── Platform C: Agent Zero (kagentz) ──
# C1: A2A liveness (no credential needed) — probes the container's internal :80/a2a/
# C2: A2A response verification (needs LITELLM_KEY) — probes POST /a2a with auth
# C3: Public access path (no credential needed) — probes https://kagentz.sysloggh.net/
# C1: A2A liveness (container-internal probe)
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
@@ -189,53 +184,27 @@ AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
if [ "$AZ_A2A_CODE" = "000" ]; then
notify "🔴" "kagentz A2A server DOWN (connection failed)"
ISSUES=$((ISSUES + 1))
echo " kagentz C1: ❌ A2A down (HTTP 000)" >> "$LOG"
echo " kagentz: ❌ A2A down (HTTP 000)" >> "$LOG"
else
case "$AZ_A2A_CODE" in
200|401)
echo " kagentz C1: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
echo " kagentz: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
*)
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
ISSUES=$((ISSUES + 1))
echo " kagentz C1: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
esac
fi
# C3: Public access path (the captain's point of view)
# Probes the public URL that NetBird proxies to the container. 200/302/401 = alive,
# 502 = proxy's "upstream refused" page (incident), connection failed = incident.
# Never restarts anything — the contract forbids restarting the platform.
KAGENTZ_PUBLIC_CODE=$(curl -s -o /dev/null --connect-timeout 10 --max-time 15 \
-w '%{http_code}' https://kagentz.sysloggh.net/ 2>/dev/null) || KAGENTZ_PUBLIC_CODE="000"
KAGENTZ_PUBLIC_CODE=$(printf '%s' "$KAGENTZ_PUBLIC_CODE" | tr -d '[:space:]')
[ -n "$KAGENTZ_PUBLIC_CODE" ] || KAGENTZ_PUBLIC_CODE="000"
if [ "$KAGENTZ_PUBLIC_CODE" = "000" ]; then
notify "🔴" "kagentz public URL DOWN (connection failed)"
ISSUES=$((ISSUES + 1))
echo " kagentz C3: ❌ public URL down (HTTP 000)" >> "$LOG"
elif [ "$KAGENTZ_PUBLIC_CODE" = "502" ]; then
notify "🔴" "kagentz public URL 502 (upstream refused)"
ISSUES=$((ISSUES + 1))
echo " kagentz C3: ❌ public URL 502 (upstream refused)" >> "$LOG"
else
case "$KAGENTZ_PUBLIC_CODE" in
200|302|401)
echo " kagentz C3: ✅ public URL alive (HTTP $KAGENTZ_PUBLIC_CODE)" >> "$LOG" ;;
*)
notify "🟡" "kagentz public URL answered HTTP $KAGENTZ_PUBLIC_CODE — running, unexpected status"
ISSUES=$((ISSUES + 1))
echo " kagentz C3: 🟡 public URL unexpected http=$KAGENTZ_PUBLIC_CODE (running, warning)" >> "$LOG" ;;
echo " kagentz: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
esac
fi
# ── Summary ──
# The run verdict is non-optimistic: when ISSUES > 0, the run is an INCIDENT.
# The lane must quote this Result line verbatim in its status report.
if [ "$ISSUES" -eq 0 ]; then
echo " Result: ✅ 0 issues (all healthy)" >> "$LOG"
if [ "$ZULIP_CRED_OK" -eq 0 ]; then
echo " Result: ✅ All healthy" >> "$LOG"
else
echo " Result: ✅ All healthy" >> "$LOG"
fi
else
echo " Result: 🔴 INCIDENT — $ISSUES issue(s) found" >> "$LOG"
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
fi
+2 -72
View File
@@ -24,7 +24,7 @@ BASE = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
base_url: http://192.168.68.116/v1
max_tokens: 4096
default: syslog-auto
provider: harness
@@ -52,7 +52,7 @@ delegation:
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
base_url: http://192.168.68.116/v1
"""
@@ -189,73 +189,3 @@ def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
assert code == 1, out
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
def test_canonical_internal_path_passes(tmp_path):
"""Rule 5 must accept the canonical internal base from hermes-key-enforcement.prose.md."""
code, out = _run(
tmp_path,
"gpu-vision",
)
# Override the base_url in the config
cfg_text = BASE.format(alias="gpu-vision").replace(
" base_url: http://192.168.68.116/v1",
" base_url: http://192.168.68.116/litellm/v1",
)
code, out = _run_config(tmp_path, "canonical-internal.yaml", cfg_text)
assert code == 0, out
assert "RESULT: PASS" in out
# Verify the correct message is shown
assert "model.base_url is canonical" in out
def test_wrong_base_url_fails(tmp_path):
"""Rule 5 must reject paths outside the allowed list."""
cfg_text = BASE.format(alias="gpu-vision").replace(
" base_url: http://192.168.68.116/litellm/v1",
" base_url: http://192.168.68.116/litellm/v1/responses",
)
code, out = _run_config(tmp_path, "wrong-base.yaml", cfg_text)
assert code == 1, out
assert "RESULT: FAIL" in out
assert "model.base_url must be one of" in out
def test_public_host_path_passes(tmp_path):
"""Rule 5 must accept the public host base."""
cfg_text = BASE.format(alias="gpu-vision").replace(
" base_url: http://192.168.68.116/v1",
" base_url: https://litellm.sysloggh.net/v1",
)
code, out = _run_config(tmp_path, "public-host.yaml", cfg_text)
assert code == 0, out
assert "RESULT: PASS" in out
def test_old_rule5_check_would_fail_canonical(tmp_path):
"""
Proof that the OLD Rule 5 check would fail the canonical internal path.
This proves the bug existed before the fix.
"""
# OLD check expected /v1, so the canonical /litellm/v1 would have failed
canonical_cfg = BASE.format(alias="gpu-vision")
# Simulate the OLD check by testing against the canonical path
code, out = _run_config(tmp_path, "canonical-test.yaml", canonical_cfg)
# NEW check: canonical /litellm/v1 SHOULD pass
assert code == 0, out
assert "RESULT: PASS" in out
# OLD check expected /v1, so the internal /v1 would have passed
# NEW check: internal /v1 is non-canonical but working (WARN not FAIL)
old_cfg = BASE.format(alias="gpu-vision").replace(
" base_url: http://192.168.68.116/litellm/v1",
" base_url: http://192.168.68.116/v1",
)
code, out = _run_config(tmp_path, "old-check-test.yaml", old_cfg)
assert code == 0, out
assert "RESULT: PASS" in out
# NEW check should pass
new_cfg = BASE.format(alias="gpu-vision")
code, out = _run_config(tmp_path, "new-check-test.yaml", new_cfg)
assert code == 0, out
assert "RESULT: PASS" in out
+13 -18
View File
@@ -98,8 +98,8 @@ exit 0
"""
CURL_STUB = r"""#!/usr/bin/env bash
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
# kagentz C3 public URL, and record every call (including notify) payloads.
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
# record every call (including notify) payloads.
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
case "$*" in
*:9200/health*)
@@ -109,8 +109,6 @@ case "$*" in
esac ;;
*server_settings*)
printf '%s' "$SERVER_HTTP" ;;
*kagentz.sysloggh.net*)
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
esac
exit 0
"""
@@ -123,8 +121,7 @@ def _write_exec(path: pathlib.Path, body: str) -> None:
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
az_a2a_code="401", az_a2a_exit=0,
kagentz_public_code="302"):
az_a2a_code="401", az_a2a_exit=0):
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
Only the LOG constant is rewritten (to keep the run inside the worktree).
@@ -157,7 +154,6 @@ def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
"PI_HTTP": "200",
"PI_BODY": CONNECTED_FIXTURE.read_text(),
"SERVER_HTTP": "200",
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
})
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
capture_output=True, text=True)
@@ -173,9 +169,8 @@ def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
assert "Server: ✅ HTTP 200" in log
assert "Abiba: ✅ Connected" in log
assert "Tanko: ✅ service=active http=200" in log
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
assert "Result: ✅ 0 issues (all healthy)" in log
assert "kagentz: ✅ A2A alive (HTTP 401)" in log
assert "Result: ✅ All healthy" in log
# A healthy run emits no notify at all — and certainly no Mumuni one.
assert proc.stdout == ""
@@ -212,8 +207,8 @@ def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
# The rest of the monitor still ran alongside the failing Tanko leg.
log = log_path.read_text()
assert "Abiba: ✅ Connected" in log
assert "kagentz C1: ✅ A2A alive" in log
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
assert "kagentz: ✅ A2A alive" in log
assert "Result: 🔴 1 issue(s) found" in log
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
@@ -221,9 +216,9 @@ def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
assert "kagentz C1: 🟡 A2A unexpected http=500 (running, warning)" in log
assert "kagentz C1: ✅ A2A alive" not in log
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
assert "kagentz: 🟡 A2A unexpected http=500 (running, warning)" in log
assert "kagentz: ✅ A2A alive" not in log
assert "Result: 🔴 1 issue(s) found" in log
assert "kagentz A2A server answered HTTP 500" in proc.stdout
@@ -235,10 +230,10 @@ def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
assert "kagentz C1: ✅ A2A alive" not in log
assert "kagentz: ❌ A2A down (HTTP 000)" in log
assert "kagentz: ✅ A2A alive" not in log
assert "unexpected" not in log
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
assert "Result: 🔴 1 issue(s) found" in log
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
-153
View File
@@ -1,153 +0,0 @@
#!/usr/bin/env bash
# test_secret_scan.sh — self-test for the commit-time secret guard.
#
# Run: bash tests/test_secret_scan.sh
# Exit: 0 all cases passed, 1 a case failed.
#
# WHY THIS FILE EXISTS: a scanner that is never observed to fail is not a guard.
# Every fixture below is fabricated and pattern-shaped; the test writes it to a
# temp tree (a path no allowlist entry covers) and asserts the guard FAILS. The
# same fixtures are deliberately listed in scripts/secret-allowlist.tsv, so the
# repo-wide tree scan stays quiet while a planted copy still bites — that is the
# difference between an explicit, reasoned exception and a guard trained to
# ignore a word.
#
# Only bash + coreutils + grep. No python/node: the Gitea runner executes job
# steps inside the runner container, which has neither.
set -uo pipefail
HERE=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
ROOT=$(cd -- "$HERE/.." && pwd)
SCAN="$ROOT/scripts/secret-scan.sh"
PASS=0
FAIL=0
LAST_OUT=""
ok() { PASS=$((PASS + 1)); echo " ✅ $1"; }
bad() { FAIL=$((FAIL + 1)); echo " ❌ $1"; }
expect_exit() { # expect_exit <want-code> <label> <cmd...>
local want="$1" label="$2"; shift 2
local rc
LAST_OUT=$("$@" 2>&1); rc=$?
if [ "$rc" -eq "$want" ]; then ok "$label (exit $rc)"; else
bad "$label (wanted exit $want, got $rc)"
printf '%s\n' "$LAST_OUT" | sed 's/^/ /' | head -8
fi
}
expect_contains() { # expect_contains <label> <needle>
if printf '%s' "$LAST_OUT" | grep -qF -- "$2"; then ok "$1"; else
bad "$1 (output did not mention: $2)"
fi
}
TMPROOT=$(mktemp -d)
trap 'rm -rf "$TMPROOT"' EXIT
echo "── secret-scan self-test ──"
# ── 1. Guard syntax ───────────────────────────────────────────────────────
expect_exit 0 "scanner parses with bash -n" bash -n "$SCAN"
# ── 2. Guard FAILS on planted, pattern-matching fixtures ──────────────────
mkdir -p "$TMPROOT/planted"
cat > "$TMPROOT/planted/ops.env" <<'EOF'
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
EOF
expect_exit 1 "planted sk-or-v1 key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted sk-or-v1 key names the openrouter-key rule" "[openrouter-key]"
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/curl.sh" <<'EOF'
curl -s -H "Authorization: Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb" http://example.invalid/
EOF
expect_exit 1 "planted literal Bearer token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted Bearer token names the bearer-token rule" "[bearer-token]"
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/pve.sh" <<'EOF'
AUTH="Authorization: PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555"
EOF
expect_exit 1 "planted Proxmox token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted Proxmox token names the proxmox-token rule" "[proxmox-token]"
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/deploy-key.pem" <<'EOF'
-----BEGIN OPENSSH PRIVATE KEY-----
b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW
-----END OPENSSH PRIVATE KEY-----
EOF
expect_exit 1 "planted PEM private key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted PEM key names the private-key rule" "[private-key]"
# Prose is scanned exactly like code — the original exposures were in .md files.
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/handover.md" <<'EOF'
- Admin credentials: `admin` / `correct-horse-battery-staple`
EOF
expect_exit 1 "planted prose credential line fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted prose line names the cred-prose rule" "[cred-prose]"
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/config.env" <<'EOF'
DB_PASSWORD=correct-horse-battery-staple
EOF
expect_exit 1 "planted password assignment fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
expect_contains "planted password assignment names the secret-assign rule" "[secret-assign]"
# ── 3. Guard stays QUIET on inert values and on the real tree ─────────────
mkdir -p "$TMPROOT/inert"
cat > "$TMPROOT/inert/config.yaml" <<'EOF'
api_key: not-needed
bearer_token=monitor_key
api_key: $LITELLM_API_KEY
EOF
expect_exit 0 "env refs, sentinels and variable names are not credentials" bash "$SCAN" --path "$TMPROOT/inert" --quiet
expect_exit 0 "current repo tree passes the guard" bash "$SCAN" --tree
expect_contains "tree run reports the allowlisted exceptions it applied" "allowlisted exception(s)"
# ── 4. Allowlist entries are path-explicit, not word-based ────────────────
# This exact line is allowlisted in infrastructure-control.prose.md; the same
# text at an unlisted path must still fail, proving the exception is per-file
# and reviewed, not a blanket "ignore the word vault".
rm -f "$TMPROOT/planted/"*
cat > "$TMPROOT/planted/unlisted.md" <<'EOF'
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
EOF
expect_exit 1 "allowlisted text at an unlisted path still fails" bash "$SCAN" --path "$TMPROOT/planted" --quiet
# ── 5. Commit-time mode: the guard blocks a STAGED credential ─────────────
# A throwaway git repo with its own copy of the scanner, so this exercises the
# real pre-commit path (--staged) without touching this repo's index.
mkdir -p "$TMPROOT/repo/scripts"
cp "$SCAN" "$TMPROOT/repo/scripts/secret-scan.sh"
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/repo/scripts/secret-patterns.tsv"
cp "$ROOT/scripts/secret-allowlist.tsv" "$TMPROOT/repo/scripts/secret-allowlist.tsv"
git -C "$TMPROOT/repo" init -q
git -C "$TMPROOT/repo" -c user.email=t@example.invalid -c user.name=test commit -q --allow-empty -m base
cat > "$TMPROOT/repo/planted.env" <<'EOF'
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
EOF
git -C "$TMPROOT/repo" add planted.env
expect_exit 1 "staged credential fails at commit time (--staged)" bash "$TMPROOT/repo/scripts/secret-scan.sh" --staged --quiet
expect_contains "staged credential names the openrouter-key rule" "[openrouter-key]"
# ── 6. Fail closed: an allowlist entry without a reason is a hard error ───
mkdir -p "$TMPROOT/scanner" "$TMPROOT/clean"
cp "$SCAN" "$TMPROOT/scanner/secret-scan.sh"
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/scanner/secret-patterns.tsv"
printf '*\t*.md\twhatever\n' > "$TMPROOT/scanner/secret-allowlist.tsv"
echo "placeholder" > "$TMPROOT/clean/ok.md"
expect_exit 2 "allowlist entry with no reason fails closed" bash "$TMPROOT/scanner/secret-scan.sh" --path "$TMPROOT/clean" --quiet
# ── Verdict ───────────────────────────────────────────────────────────────
echo ""
if [ "$FAIL" -gt 0 ]; then
echo "❌ secret-scan self-test FAILED — $PASS passed, $FAIL failed"
exit 1
fi
echo "✅ secret-scan self-test passed ($PASS cases)"
-209
View File
@@ -1,209 +0,0 @@
#!/usr/bin/env python3
"""Behavioural tests for zulip-monitor.sh kagentz C1/C3 legs and the Result verdict.
WHY THIS FILE EXISTS: the 2026-09-19 kagentz A2A outage was correctly detected
by the monitor (C1 returned 000, the run logged an issue) but the lane's own
summarization in ops.status wrote "OK" with the note "A2A server DOWN — expected
(no credentials configured)". The optimistic verdict came from the lane, not the
script. The fix adds a C3 public-access-path leg and makes the Result line say
"INCIDENT" when issues are found, so the lane can quote it verbatim.
CONTRACT UNDER TEST:
* C1 (A2A liveness, no credential): 000 → INCIDENT.
* C3 (public access path, no credential): 502 → INCIDENT, 000 → INCIDENT,
200/302/401 → alive.
* Result verdict: when ISSUES > 0, the log's Result line says "INCIDENT",
not just "issues found".
* Healthy control: C1 401 + C3 302 → 0 issues, "all healthy".
HOW: behavioural execution using the sandbox pattern already in this repo
(tests/test_mumuni_monitor_removal.py). The sandbox copies the shipped monitor
verbatim, rewrites only its LOG constant, and runs it with stub ssh/curl on
PATH. Each test asserts from the run's own log/verdict, not from file text.
Usage: python3 -m pytest tests/test_zulip_kagentz_legs.py
"""
from __future__ import annotations
import os
import pathlib
import stat
import subprocess
ROOT = pathlib.Path(__file__).resolve().parents[1]
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
# ── Stub ssh: answers Tanko and Agent Zero probes by env vars ──────────
SSH_STUB = r"""#!/usr/bin/env bash
# Stub ssh: record the target host, then answer by host + remote command.
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
host=""
for a in "$@"; do
case "$a" in
*@192.168.*) host="${a##*@}" ;;
esac
done
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
cmd="${*: -1}"
case "$host" in
192.168.68.15)
case "$cmd" in
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
*curl*) printf '%s' "$TANKO_HTTP" ;;
esac ;;
192.168.68.14)
case "$cmd" in
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "${AZ_A2A_EXIT:-0}" ;;
esac ;;
*)
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
esac
exit 0
"""
# ── Stub curl: serves Zulip server, Abiba health, and C3 public URL ───
CURL_STUB = r"""#!/usr/bin/env bash
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
# C3 public URL probe (https://kagentz.sysloggh.net/). Record every call.
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
case "$*" in
*:9200/health*)
case " $* " in
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
*) printf '%s' "$PI_BODY" ;; # body probe
esac ;;
*server_settings*)
printf '%s' "$SERVER_HTTP" ;;
*kagentz.sysloggh.net*)
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
esac
exit 0
"""
def _write_exec(path: pathlib.Path, body: str) -> None:
path.write_text(body)
path.chmod(path.stat().st_mode
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
az_a2a_code="401", az_a2a_exit=0,
kagentz_public_code="302"):
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
Only the LOG constant is rewritten (to keep the run inside the worktree).
Everything else — legs, labels, notify logic — is the shipped script.
"""
sandbox = tmp_path / "sandbox"
bindir = sandbox / "bin"
record = sandbox / "record"
bindir.mkdir(parents=True)
record.mkdir()
_write_exec(bindir / "ssh", SSH_STUB)
_write_exec(bindir / "curl", CURL_STUB)
source = ZULIP_MONITOR.read_text()
log_line = 'LOG="/root/zulip-health-monitor.log"'
assert log_line in source, "LOG constant moved — update the sandbox harness"
log_path = sandbox / "zulip-health-monitor.log"
script = sandbox / "zulip-monitor.sh"
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
env = dict(os.environ)
env.update({
"PATH": f"{bindir}:{env['PATH']}",
"RECORD_DIR": str(record),
"TANKO_SVC": tanko_svc,
"TANKO_HTTP": tanko_http,
"AZ_A2A_CODE": az_a2a_code,
"AZ_A2A_EXIT": str(az_a2a_exit),
"PI_HTTP": "200",
"PI_BODY": CONNECTED_FIXTURE.read_text(),
"SERVER_HTTP": "200",
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
})
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
capture_output=True, text=True)
return proc, record, log_path
# ── Required behavioural cases ─────────────────────────────────────────
def test_c3_502_is_incident(tmp_path):
"""C3 public leg returns 502 → the run's verdict is an INCIDENT,
the C3 line names the 502, and the run is not summarised as healthy."""
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="502")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# The C3 line names the 502.
assert "kagentz C3: ❌ public URL 502 (upstream refused)" in log
# The verdict is an INCIDENT, not healthy.
assert "Result: 🔴 INCIDENT" in log
assert "all healthy" not in log
# The run is not summarised as healthy.
assert "✅ 0 issues" not in log
def test_c3_000_is_incident(tmp_path):
"""C3 public leg returns 000 → INCIDENT."""
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="000")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# The C3 line reports the connection failure.
assert "kagentz C3: ❌ public URL down (HTTP 000)" in log
# The verdict is an INCIDENT.
assert "Result: 🔴 INCIDENT" in log
assert "all healthy" not in log
assert "✅ 0 issues" not in log
def test_healthy_control_c1_401_c3_302(tmp_path):
"""Healthy control: C1 401 plus C3 302 → 0 issues and a healthy verdict,
proving the new leg cannot cry wolf."""
proc, record, log_path = _run_monitor(tmp_path,
az_a2a_code="401",
kagentz_public_code="302")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# Both legs report alive.
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
# Zero issues, healthy verdict.
assert "Result: ✅ 0 issues (all healthy)" in log
# No INCIDENT.
assert "INCIDENT" not in log
# No notify fired for kagentz.
assert "kagentz public URL" not in proc.stdout
assert "kagentz A2A server" not in proc.stdout
def test_c1_000_is_incident(tmp_path):
"""C1 returns 000 → INCIDENT, keeping the leg that actually caught this
outage covered behaviourally."""
proc, record, log_path = _run_monitor(tmp_path,
az_a2a_code="000",
az_a2a_exit=7,
kagentz_public_code="302")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# The C1 line reports the A2A down.
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
# The verdict is an INCIDENT (even though C3 is healthy).
assert "Result: 🔴 INCIDENT" in log
assert "all healthy" not in log
# The notify fired for the A2A down.
assert "kagentz A2A server DOWN" in proc.stdout
+19 -42
View File
@@ -1,9 +1,9 @@
---
kind: responsibility
name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.4.0
version: 3.3.0
runtime_contract: 2
agent: abiba
report_only_agents:
@@ -30,7 +30,7 @@ session start.
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200`
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
- **Relay access** via RA-H OS MCP for alert delivery
@@ -129,10 +129,10 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
### Liveness rule (scoped)
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN.
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above).
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
@@ -469,10 +469,9 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
> monitor issued a restart for something that could not start, posting a false
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
> liveness/response and public-path access only, and a probe must never restart
> a platform.
> liveness only, and a probe must never restart a platform.
**C1: A2A Server Health (no credential needed)**
**C1: A2A Server Health**
```bash
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
@@ -480,9 +479,9 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
```
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down (INCIDENT). Any other status → running but unexpected: log/report it, never restart.
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
**C2: A2A Response Verification (requires LITELLM_KEY)**
**C2: A2A Response Verification**
```bash
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
@@ -492,30 +491,15 @@ ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
```
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set (this is a credential issue, NOT a server-down incident).
**C3: Public Access Path (no credential needed)**
```bash
# Probes the public URL that NetBird proxies to the agent-zero container.
# This is the captain's point of view: if the captain can't reach it, it's down.
# 200/302/401 = alive, 502 = proxy's "upstream refused" page (INCIDENT),
# connection failed (000) = INCIDENT. Never restarts anything.
curl -s -o /dev/null --connect-timeout 10 --max-time 15 -w '%{http_code}' https://kagentz.sysloggh.net/
```
Expected: `200` (Agent Zero login page), `302` (redirect), or `401` (auth-gated) = alive. `502` = NetBird proxy's "upstream refused" page (the container's port 80 is not listening) = INCIDENT. `000` (connection failed) = INCIDENT. Any other status → running but unexpected: log/report it, never restart.
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
**Platform C Actions**
| Condition | Action |
|-----------|--------|
| C1 A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
| C1 A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
| C2 A2A returns 401 | Check `LITELLM_KEY` is set (credential issue, NOT server-down) |
| C3 public URL returns `502` (upstream refused) | Alert only — the container's port 80 is not listening; investigate the agent-zero container |
| C3 public URL returns `000` (connection failed) | Alert only — the public path is down; investigate the NetBird proxy or the container |
| C3 public URL returns a status other than `200`/`302`/`401` | Log/report as a warning — reported, never healed on |
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
### Step 5: Global Checks
@@ -529,19 +513,12 @@ If any bot processes >50 bot-originated messages in 15min → warning.
### Step 6: Compile and Report
1. Run `scripts/zulip-monitor.sh` and take its final `Result:` line as the
authoritative run verdict. The verdict line is either
`Result: ✅ 0 issues (all healthy)` or
`Result: 🔴 INCIDENT — N issue(s) found`.
2. Quote that `Result:` line verbatim in the status report. When it says
`INCIDENT`, the run MUST be reported as an incident — never summarised as
OK/healthy and never annotated as "expected".
3. Compile all platform checks and severity
4. Determine `overall_severity` from worst per-agent severity
5. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
6. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
7. If any agent critical or >2 degraded: send relay message to user
8. Update `last_check` timestamp in `### Maintains` snapshot
1. Compile all platform checks and severity
2. Determine `overall_severity` from worst per-agent severity
3. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
4. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
5. If any agent critical or >2 degraded: send relay message to user
6. Update `last_check` timestamp in `### Maintains` snapshot
### Restart Debounce