Compare commits
24
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5ab5de4704 | ||
|
|
8a5cba8515 | ||
|
|
20f882412f | ||
|
|
30b2fe3fdc | ||
|
|
5112c566c8 | ||
|
|
83307eb9b2 | ||
|
|
8245716286 | ||
|
|
cfb6c03572 | ||
|
|
85f70f65bc | ||
|
|
a820b3f7dd | ||
|
|
0b92ab17b1 | ||
|
|
9edefe036e | ||
|
|
dd6e1e8b22 | ||
|
|
c712d4faf0 | ||
|
|
7f62f19c24 | ||
|
|
57bfe7e06a | ||
|
|
9100ea3326 | ||
|
|
0f26119859 | ||
|
|
5c1c8d7c19 | ||
|
|
8a2ea2d0d7 | ||
|
|
dae8d14880 | ||
|
|
39209c7ac9 | ||
|
|
8ed3b9c606 | ||
|
|
6c616a9e58 |
@@ -1 +1,2 @@
|
||||
__pycache__/
|
||||
state/host-disk-bands.json
|
||||
|
||||
@@ -0,0 +1,135 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: agent-health-check
|
||||
description: >
|
||||
Consolidated agent health verification for the LiteLLM + GPU + Zulip +
|
||||
gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours
|
||||
via cron and on-demand via "run contract: agent-health-check". Verifies:
|
||||
LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway
|
||||
liveness, CT liveness, gateway log health, config YAML integrity,
|
||||
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
|
||||
title: Agent Health Check — Consolidated
|
||||
version: 1.0.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
---
|
||||
|
||||
# Agent Health Check
|
||||
|
||||
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
|
||||
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only.
|
||||
|
||||
**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
|
||||
|
||||
## Requires
|
||||
|
||||
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
|
||||
- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
|
||||
- **Python 3** for script execution
|
||||
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
|
||||
|
||||
## Maintains
|
||||
|
||||
- last_check: timestamp — When the last full diagnostic ran
|
||||
- overall_severity: "healthy" | "degraded" | "critical"
|
||||
- liteLLM_keys: map of agent → key validity
|
||||
- gpu_ports: map of host → port conflict status
|
||||
- agents: map of agent → streaming health + gateway liveness
|
||||
- ct_liveness: map of CT → active status
|
||||
- config_integrity: map of config file → valid/invalid
|
||||
|
||||
## Execution
|
||||
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool
|
||||
calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Run the consolidated health check script
|
||||
python3 /root/scripts/agent-health-check.py --json
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the script
|
||||
executed from** so a stale-consumer report is distinguishable from a real fault
|
||||
at read time. Summarize actual results from each check. Apply the standing probe
|
||||
rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
|
||||
|
||||
**Expected output**: JSON with `overall_severity` field. If `healthy`, report
|
||||
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
||||
failures and their severity.
|
||||
|
||||
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
|
||||
MUST include one clause per check leg, in every state: healthy, degraded/warn,
|
||||
skipped, or failed. A missing leg must never look the same as a healthy leg.
|
||||
Required legs and their templates in every state:
|
||||
|
||||
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
|
||||
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
|
||||
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
|
||||
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
|
||||
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
|
||||
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
|
||||
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
|
||||
- The GPU leg has six non-healthy states the code can produce:
|
||||
(i) `gpu-unreachable:{host}` — SSH probe failed;
|
||||
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
|
||||
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
|
||||
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
|
||||
(v) unit active, /health body contains "error" — error response;
|
||||
(vi) unit active, /health body unrecognised — unknown health.
|
||||
In every case the failing host and reason must be named.
|
||||
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
|
||||
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
|
||||
- skipped: `CTs: SKIPPED (SSH access unavailable)`
|
||||
- `Vault secrets: 3/3 present`
|
||||
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
|
||||
- skipped: `Vault secrets: SKIPPED (vault not configured)`
|
||||
|
||||
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
|
||||
long as the host is identifiable from context; the full `probe-failed: <target>
|
||||
<kind>` form is required when a leg reports a failure in the detail section.
|
||||
|
||||
### Probe Shape (per standing rules from 1150.msg)
|
||||
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||
service answered — report the code, never "down". A redirect is not a failure.
|
||||
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
2. **A failed probe is never a service verdict.** Print
|
||||
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
||||
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
||||
then report.
|
||||
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
||||
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
||||
(retried at 25s: also timeout)" is actionable.
|
||||
|
||||
## Strategies
|
||||
|
||||
### When LiteLLM keys are invalid
|
||||
Report the specific agent + key name. Do not attempt to fix — credential
|
||||
rotation is a separate operation.
|
||||
|
||||
### When GPU port conflicts are detected
|
||||
Report the conflicting ports and processes. Do not kill processes — that's a
|
||||
destructive action requiring captain approval.
|
||||
|
||||
### When gateway liveness is degraded
|
||||
Report the specific CT + gateway status. Do not restart unless the restart
|
||||
debounce window has passed.
|
||||
|
||||
### When CT liveness is down
|
||||
Report the specific CT. Do not restart — that's a destructive action.
|
||||
|
||||
### When config YAML is invalid
|
||||
Report the specific file + parse error. Do not fix — that's a config change.
|
||||
|
||||
### When gateway log health is degraded
|
||||
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
|
||||
|
||||
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
|
||||
|
||||
## Continuity
|
||||
|
||||
- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
|
||||
- **On `agent-health` command**: Run on-demand and report to user
|
||||
- **On critical alert**: Escalate to relay message immediately
|
||||
@@ -16,8 +16,8 @@ litellm.exceptions.AuthenticationError: OpenrouterException -
|
||||
```
|
||||
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
|
||||
|
||||
**Old Key**: `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`
|
||||
**New Key**: `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`
|
||||
**Old Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
**New Key**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
|
||||
### 2. Telegram Bot Conflict (CRITICAL)
|
||||
@@ -48,14 +48,14 @@ McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 sec
|
||||
```bash
|
||||
# Container .env update
|
||||
sudo docker exec agent-zero bash -c '
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
|
||||
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»|" /a0/usr/.env
|
||||
'
|
||||
```
|
||||
|
||||
**Verification**:
|
||||
```bash
|
||||
curl -s https://openrouter.ai/api/v1/auth/key \
|
||||
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
|
||||
-H "Authorization: Bearer «vault: agents/production OPENROUTER_API_KEY»" | python3 -m json.tool
|
||||
```
|
||||
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
|
||||
|
||||
@@ -140,7 +140,7 @@ Added section:
|
||||
|
||||
| Component | Status | Details |
|
||||
|-----------|--------|---------|
|
||||
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
|
||||
| **OpenRouter Key** | ✅ Valid | `«vault: agents/production OPENROUTER_API_KEY»` user verified, |
|
||||
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
|
||||
| **MCP Services** | ✅ Working | No timeouts after key fix |
|
||||
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
|
||||
|
||||
@@ -54,7 +54,7 @@ description: >
|
||||
```
|
||||
|
||||
4. **Return status**
|
||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-synthetic...", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
|
||||
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
|
||||
- If vault secret is missing: `{ vault_synced: false }`
|
||||
|
||||
@@ -89,8 +89,8 @@ description: >
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| **Key Prefix** | `sk-or-v1-0af3f3` |
|
||||
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
|
||||
| **Key Prefix** | `«vault: agents/production OPENROUTER_API_KEY»` |
|
||||
| **Full Key** | `«vault: agents/production OPENROUTER_API_KEY»` (in vault + /a0/usr/.env) |
|
||||
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
|
||||
| **Free Tier** | No |
|
||||
| **Monthly Usage** | 0 (as of 2026-09-01) |
|
||||
@@ -101,7 +101,7 @@ description: >
|
||||
|
||||
| Date | Action | Notes |
|
||||
|------|--------|-------|
|
||||
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||
| 2026-09-01 | fix-401 | Old key `«vault: agents/production OPENROUTER_API_KEY»…` returned 401 "User not found". Replaced with new key `«vault: agents/production OPENROUTER_API_KEY»…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
|
||||
|
||||
## Infrastructure References
|
||||
|
||||
|
||||
@@ -65,7 +65,7 @@ Host filesystems have their own risk profile and their own bands. A host root ne
|
||||
|
||||
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
||||
|
||||
**State lives in a small JSON state file:** `state/host-disk-bands.json`, keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
|
||||
**State lives in a small JSON state file:** the scanner resolves it to an absolute path from the script's location — `$(dirname "$0")/state/host-disk-bands.json` (i.e., the `state/` directory next to the `scripts/` directory in the prose-contracts repo). The state file is gitignored runtime state — the scanner creates it on first run. Keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
|
||||
|
||||
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
|
||||
|
||||
|
||||
@@ -101,7 +101,7 @@ auxiliary:
|
||||
fallback_providers:
|
||||
- provider: deepseek
|
||||
base_url: https://api.deepseek.com
|
||||
api_key: sk-b7d9... # ← hardcoded OK (external)
|
||||
api_key: sk-synthetic-external-example # ← hardcoded OK (external, synthetic example)
|
||||
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
|
||||
```
|
||||
|
||||
@@ -109,7 +109,7 @@ fallback_providers:
|
||||
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
|
||||
model:
|
||||
provider: harness
|
||||
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION: hardcoded key
|
||||
api_key: sk-synthetic-example-12345 # ← RULE VIOLATION: hardcoded key (synthetic example)
|
||||
|
||||
model:
|
||||
provider: harness
|
||||
@@ -139,6 +139,16 @@ scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### ACCEPTABLE PATTERN
|
||||
Agent keys live in `.env` or `.env.vault` files with 600 permissions (koonimo's shape is the canonical example). A plaintext key inside a `config.yaml` or any `config.yaml.bak-*` file is a violation — the backup files are not part of the runtime credential path and are not watched by the scanner, so a key in them is stale clutter that a future reader can mistake for a working key.
|
||||
|
||||
**Fix procedure** (when a backup file is found with a plaintext key):
|
||||
1. Move the file out of the scanned tree (e.g., `mv /root/.hermes/config.yaml.bak-* /root/hermes-config-backups/`) — do NOT delete the file, just move it so the scanner pattern no longer matches.
|
||||
2. Re-run the reachability check to confirm COMPLIANT.
|
||||
3. Report the before/after check output and the commands you ran.
|
||||
|
||||
**Rationale**: Moving the file preserves history without leaving a credential where a scanner trips over it. Deleting the file loses the historical context. Keeping it in place means the next scan will report it as a finding and waste time re-deciding.
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
@@ -172,7 +182,7 @@ grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
|
||||
|
||||
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
||||
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /root/.config/systemd/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-…' /root/.config/systemd/ 2>/dev/null
|
||||
|
||||
# 3. Verify running process env matches dedicated key
|
||||
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
||||
|
||||
@@ -181,8 +181,8 @@ description: >
|
||||
**Stirling-PDF** (deployed 2026-07-03, Authentik SSO 2026-07-03):
|
||||
- URL: `https://pdf.sysloggh.net` (public) / `http://192.168.68.7:8989` (direct)
|
||||
- Swagger: `http://192.168.68.7:8989/swagger-ui.html`
|
||||
- Admin credentials: `admin` / `kakashi20stirling`
|
||||
- API key: `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88`
|
||||
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_USER»` / `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||
- API key: `«vault: infrastructure/production STIRLING_API_KEY»`
|
||||
- Authentik OAuth2: configured but disabled (requires paid Server license). Ready to enable: set `SECURITY_OAUTH2_ENABLED=true` + `SECURITY_LOGINMETHOD=all`
|
||||
- Compose: `/opt/home_stack/docker-compose.yml`
|
||||
- Control script: `/opt/home_stack/infra-control.sh`
|
||||
@@ -636,7 +636,7 @@ monitor, or integration breaks.
|
||||
```bash
|
||||
# Full cluster status
|
||||
PVE="https://minipve.sysloggh.net"
|
||||
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
AUTH="Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
|
||||
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
|
||||
|
||||
# Docker health from Abiba
|
||||
|
||||
@@ -144,8 +144,8 @@ through its agent wrapper.
|
||||
safety net for vault outage or token revocation. Must be kept in sync on rotation.
|
||||
Example:
|
||||
```bash
|
||||
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
|
||||
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
|
||||
MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY»
|
||||
MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY»
|
||||
```
|
||||
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
|
||||
```ini
|
||||
@@ -191,7 +191,7 @@ through its agent wrapper.
|
||||
### Tanko migration (COMPLETED 2026-07-17)
|
||||
|
||||
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
|
||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
|
||||
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-synthetic-tanko-example…`)
|
||||
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
|
||||
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
|
||||
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
|
||||
@@ -271,13 +271,13 @@ not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update ma
|
||||
UI bootstrap, model selection) is built around OpenRouter's native authentication.
|
||||
|
||||
**Key Storage:**
|
||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
|
||||
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=«vault: agents/production OPENROUTER_API_KEY»…`)
|
||||
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
|
||||
- **Fallback**: The container's .env is the primary source; vault sync is optional
|
||||
(unlike fleet agents which require vault injection)
|
||||
|
||||
**Current Key (2026-09-01):**
|
||||
- **Prefix**: `sk-or-v1-0af3f3…`
|
||||
- **Prefix**: `«vault: agents/production OPENROUTER_API_KEY»`
|
||||
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
|
||||
- **Plan**: Paid (not free tier)
|
||||
- **Usage**: 0 (as of 2026-09-01)
|
||||
|
||||
@@ -116,7 +116,7 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (hardcoded key → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM (deprecated key, no live usage).
|
||||
|
||||
## Maintains
|
||||
|
||||
|
||||
@@ -122,10 +122,8 @@ def _fail(key, agent_name=None):
|
||||
|
||||
|
||||
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
|
||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||
|
||||
# Fallback: if no env token, read the shared vault token file
|
||||
if not INFISICAL_TOKEN:
|
||||
# Fallback: read the shared vault token file
|
||||
_token_path = os.path.expanduser("~/.infisical-token")
|
||||
if os.path.isfile(_token_path):
|
||||
try:
|
||||
@@ -133,6 +131,7 @@ if not INFISICAL_TOKEN:
|
||||
INFISICAL_TOKEN = _f.read().strip()
|
||||
except (OSError, UnicodeDecodeError):
|
||||
pass
|
||||
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
|
||||
|
||||
# ── Helpers ──────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@@ -16,14 +16,16 @@ from email.mime.text import MIMEText
|
||||
from email.mime.multipart import MIMEMultipart
|
||||
|
||||
PVE = "https://192.168.68.12:8006"
|
||||
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
||||
AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
|
||||
|
||||
# ── Shared credentials —─
|
||||
|
||||
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_KEY = "cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_KEY}"
|
||||
ZULIP_API_KEY = os.environ.get("ZULIP_API_KEY", "")
|
||||
if not ZULIP_API_KEY:
|
||||
raise SystemExit("ZULIP_API_KEY not set — refusing to run with no credential")
|
||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_API_KEY}"
|
||||
|
||||
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||
LITELLM_BACKEND = "192.168.68.116"
|
||||
@@ -669,7 +671,10 @@ def send_email(html_content, subject_prefix=""):
|
||||
msg.attach(MIMEText(html_content, "html"))
|
||||
|
||||
try:
|
||||
EMAIL_PASSWORD = "rgbuomwcydxwbszd"
|
||||
EMAIL_PASSWORD = os.environ.get("EMAIL_PASSWORD") or os.environ.get("SMTP_PASSWORD") or os.environ.get("MAIL_PASSWORD")
|
||||
if not EMAIL_PASSWORD:
|
||||
print("EMAIL_PASSWORD not set — refusing to send email", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
GMAIL_EMAIL = "jtabiri@gmail.com"
|
||||
|
||||
server = smtplib.SMTP("smtp.gmail.com", 587)
|
||||
|
||||
+236
-6
@@ -153,6 +153,25 @@ GPU_HOSTS = [
|
||||
CONNECT_TIMEOUT = 5
|
||||
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
|
||||
|
||||
# Host filesystem thresholds (from contract)
|
||||
HOST_THRESHOLDS = {
|
||||
"WARN": 85,
|
||||
"AMBER": 90,
|
||||
"RED": 95,
|
||||
}
|
||||
|
||||
# State file path (absolute, so execution context doesn't matter)
|
||||
STATE_FILE = pathlib.Path(__file__).resolve().parent.parent / "state" / "host-disk-bands.json"
|
||||
|
||||
# PVE nodes to probe for host filesystems
|
||||
HOST_NODES = [
|
||||
{"hostname": "acerpve", "ip": "192.168.68.9"},
|
||||
{"hostname": "amdpve", "ip": "192.168.68.15"},
|
||||
{"hostname": "storepve", "ip": "192.168.68.6"},
|
||||
{"hostname": "minipve", "ip": "192.168.68.12"},
|
||||
{"hostname": "ocupve", "ip": "192.168.68.5"},
|
||||
]
|
||||
|
||||
|
||||
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
|
||||
"""Run a command and return (exit_code, stdout, stderr)."""
|
||||
@@ -298,6 +317,151 @@ def scan_fleet() -> list[dict]:
|
||||
return results
|
||||
|
||||
|
||||
def classify_band(usage_pct: float) -> str:
|
||||
"""Classify a percentage into a band."""
|
||||
if usage_pct >= HOST_THRESHOLDS["RED"]:
|
||||
return "HOST-RED"
|
||||
elif usage_pct >= HOST_THRESHOLDS["AMBER"]:
|
||||
return "HOST-AMBER"
|
||||
elif usage_pct >= HOST_THRESHOLDS["WARN"]:
|
||||
return "HOST-WARN"
|
||||
else:
|
||||
return "GREEN"
|
||||
|
||||
|
||||
def probe_host_filesystems() -> tuple[list[dict], dict[str, str]]:
|
||||
"""Probe host filesystems on all PVE nodes.
|
||||
|
||||
Returns:
|
||||
- List of host filesystem results
|
||||
- Dict of volume_key -> current_band (for state file)
|
||||
"""
|
||||
results = []
|
||||
current_bands = {}
|
||||
|
||||
for node in HOST_NODES:
|
||||
ip = node["ip"]
|
||||
hostname = node["hostname"]
|
||||
|
||||
# Probe df for host filesystems
|
||||
probe_cmd = f'ssh {SSH_OPTS} root@{ip} "df -hP / /media/* tank 2>/dev/null | tail -n +2"'
|
||||
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
|
||||
|
||||
if exit_code != 0:
|
||||
results.append({
|
||||
"target": f"{hostname} ({ip})",
|
||||
"hostname": hostname,
|
||||
"ip": ip,
|
||||
"reachable": False,
|
||||
"volumes": [],
|
||||
"probe_cmd": probe_cmd,
|
||||
"failure_kind": "ssh-error",
|
||||
})
|
||||
continue
|
||||
|
||||
# Parse df output and classify each volume
|
||||
volumes = []
|
||||
for line in stdout.splitlines():
|
||||
parts = line.split()
|
||||
if len(parts) < 6:
|
||||
continue
|
||||
|
||||
dev, size, used, avail, pct_str, mount = parts[:6]
|
||||
pct = float(pct_str.rstrip("%"))
|
||||
band = classify_band(pct)
|
||||
|
||||
# Volume type classification
|
||||
if mount.startswith("/media/"):
|
||||
vol_type = "media"
|
||||
elif mount == "/" or "pve" in dev:
|
||||
vol_type = "host-root"
|
||||
elif mount == "tank" or "tank" in mount:
|
||||
vol_type = "pbs-datastore"
|
||||
else:
|
||||
vol_type = "other"
|
||||
|
||||
# Volume key for state file (host/volume)
|
||||
volume_key = f"{hostname}/{mount}"
|
||||
current_bands[volume_key] = band
|
||||
|
||||
volumes.append({
|
||||
"mount": mount,
|
||||
"device": dev,
|
||||
"size": size,
|
||||
"used": used,
|
||||
"avail": avail,
|
||||
"pct": pct,
|
||||
"band": band,
|
||||
"type": vol_type,
|
||||
})
|
||||
|
||||
results.append({
|
||||
"target": f"{hostname} ({ip})",
|
||||
"hostname": hostname,
|
||||
"ip": ip,
|
||||
"reachable": True,
|
||||
"volumes": volumes,
|
||||
"probe_cmd": probe_cmd,
|
||||
"failure_kind": None,
|
||||
})
|
||||
|
||||
return results, current_bands
|
||||
|
||||
|
||||
def read_state_file() -> Optional[dict[str, str]]:
|
||||
"""Read the state file if it exists."""
|
||||
if not STATE_FILE.exists():
|
||||
return None
|
||||
try:
|
||||
with open(STATE_FILE) as f:
|
||||
return json.load(f)
|
||||
except (json.JSONDecodeError, IOError) as e:
|
||||
print(f"⚠️ State file exists but unreadable: {e}", file=sys.stderr)
|
||||
return {}
|
||||
|
||||
|
||||
def write_state_file(bands: dict[str, str]) -> None:
|
||||
"""Write the state file."""
|
||||
STATE_FILE.parent.mkdir(parents=True, exist_ok=True)
|
||||
try:
|
||||
with open(STATE_FILE, "w") as f:
|
||||
json.dump(bands, f, indent=2)
|
||||
except IOError as e:
|
||||
print(f"⚠️ State file write failed: {e}", file=sys.stderr)
|
||||
|
||||
|
||||
def detect_transitions(current_bands: dict[str, str], prior_bands: Optional[dict[str, str]]) -> list[dict]:
|
||||
"""Detect band transitions (current vs. prior)."""
|
||||
if prior_bands is None:
|
||||
# First run — no transitions, just establish baseline
|
||||
return []
|
||||
|
||||
transitions = []
|
||||
# Check for volumes that moved to a higher band (escalation)
|
||||
for volume, current_band in current_bands.items():
|
||||
prior_band = prior_bands.get(volume, "GREEN")
|
||||
|
||||
# Band ordering: GREEN < HOST-WARN < HOST-AMBER < HOST-RED
|
||||
band_order = {"GREEN": 0, "HOST-WARN": 1, "HOST-AMBER": 2, "HOST-RED": 3}
|
||||
|
||||
if band_order[current_band] > band_order[prior_band]:
|
||||
transitions.append({
|
||||
"type": "escalation",
|
||||
"volume": volume,
|
||||
"from": prior_band,
|
||||
"to": current_band,
|
||||
})
|
||||
elif band_order[current_band] < band_order[prior_band]:
|
||||
transitions.append({
|
||||
"type": "recovery",
|
||||
"volume": volume,
|
||||
"from": prior_band,
|
||||
"to": current_band,
|
||||
})
|
||||
|
||||
return transitions
|
||||
|
||||
|
||||
def render_results(results: list[dict]) -> str:
|
||||
"""Render scan results in human-readable format."""
|
||||
lines = []
|
||||
@@ -316,22 +480,88 @@ def render_results(results: list[dict]) -> str:
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def render_host_results(results: list[dict], transitions: list[dict], prior_bands: Optional[dict[str, str]]) -> str:
|
||||
"""Render host filesystem results in human-readable format."""
|
||||
lines = []
|
||||
lines.append("")
|
||||
lines.append("=== Host Filesystem Bands ===")
|
||||
lines.append("")
|
||||
|
||||
# Render transitions first (they're the actionable alerts)
|
||||
if prior_bands is None:
|
||||
lines.append(" (first run — recording baseline, no alerts)")
|
||||
elif not transitions:
|
||||
lines.append(" (no band changes since last scan)")
|
||||
else:
|
||||
for t in transitions:
|
||||
volume, from_band, to_band = t["volume"], t["from"], t["to"]
|
||||
if t["type"] == "escalation":
|
||||
lines.append(f" ⚠️ {volume}: {from_band} → {to_band} (ESCALATION)")
|
||||
else:
|
||||
lines.append(f" ✅ {volume}: {from_band} → {to_band} (RECOVERY)")
|
||||
|
||||
# Render all volumes with their bands
|
||||
lines.append("")
|
||||
for node_result in results:
|
||||
if not node_result["reachable"]:
|
||||
lines.append(f" ❌ {node_result['target']}: UNREACHABLE ({node_result['failure_kind']})")
|
||||
continue
|
||||
|
||||
lines.append(f" {node_result['target']}:")
|
||||
for vol in node_result["volumes"]:
|
||||
lines.append(f" {vol['mount']} ({vol['type']}): {vol['pct']}% ({vol['used']}/{vol['size']}, {vol['avail']} free) -> {vol['band']}")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
import argparse
|
||||
|
||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
|
||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests and host filesystems.")
|
||||
ap.add_argument("--json", action="store_true", help="machine-readable output")
|
||||
ap.add_argument("--hosts-only", action="store_true", help="scan host filesystems only")
|
||||
ap.add_argument("--guests-only", action="store_true", help="scan guests only (skip host filesystems)")
|
||||
args = ap.parse_args()
|
||||
|
||||
results = scan_fleet()
|
||||
# Scan guests (unless --hosts-only)
|
||||
guest_results = []
|
||||
if not args.hosts_only:
|
||||
guest_results = scan_fleet()
|
||||
|
||||
# Scan host filesystems (unless --guests-only)
|
||||
host_results = []
|
||||
current_bands = {}
|
||||
if not args.guests_only:
|
||||
host_results, current_bands = probe_host_filesystems()
|
||||
|
||||
# Read prior state and detect transitions
|
||||
prior_bands = read_state_file()
|
||||
transitions = detect_transitions(current_bands, prior_bands)
|
||||
|
||||
# Write new state
|
||||
write_state_file(current_bands)
|
||||
else:
|
||||
prior_bands = None
|
||||
transitions = []
|
||||
|
||||
if args.json:
|
||||
print(json.dumps(results, indent=2))
|
||||
# JSON output
|
||||
output = {
|
||||
"guests": guest_results,
|
||||
"hosts": host_results,
|
||||
"transitions": transitions,
|
||||
"prior_bands": prior_bands,
|
||||
}
|
||||
print(json.dumps(output, indent=2))
|
||||
else:
|
||||
print(render_results(results))
|
||||
# Human-readable output
|
||||
if guest_results:
|
||||
print(render_results(guest_results))
|
||||
|
||||
if host_results:
|
||||
print(render_host_results(host_results, transitions, prior_bands))
|
||||
|
||||
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
|
||||
# (a probe error means the probe itself failed, not just that the guest was unreachable)
|
||||
# Exit 0 if all probed (reachable or not), 1 if any probe error
|
||||
return 0
|
||||
|
||||
|
||||
|
||||
Executable
+193
@@ -0,0 +1,193 @@
|
||||
#!/bin/bash
|
||||
# infrastructure-monitoring.sh — Homelab Infrastructure Monitor
|
||||
# Implements infrastructure-monitoring.prose.md v4
|
||||
#
|
||||
# Legs: Grafana, Prometheus, LiteLLM, PVE API (5 nodes), GPU exporters,
|
||||
# Docker Stats, PVE Exporter, PM2
|
||||
#
|
||||
# Design:
|
||||
# - Every target, port, path, and expected status is defined in code
|
||||
# - Any HTTP status (200/301/302/401/403/404) = ALIVE
|
||||
# - Only connection failures (000/timeout) = probe-failed
|
||||
# - PVE API uses -k flag (self-signed certs)
|
||||
# - Docker Stats and PVE Exporter bind to 127.0.0.1, probed via SSH
|
||||
# - Non-zero exit with named failures
|
||||
# - No "OK" summary when any leg failed
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
# ── Configuration ───────────────────────────────────────────────────────────
|
||||
# Documented targets (from infrastructure-monitoring.prose.md)
|
||||
# Change these in ONE place; tests assert against these values
|
||||
|
||||
GRAFANA_HOST="192.168.68.116"
|
||||
GRAFANA_PORT="3001"
|
||||
GRAFANA_PATH="/api/health"
|
||||
GRAFANA_EXPECTED="200|302"
|
||||
|
||||
PROMETHEUS_HOST="192.168.68.116"
|
||||
PROMETHEUS_PORT="9090"
|
||||
PROMETHEUS_PATH="/-/healthy"
|
||||
PROMETHEUS_EXPECTED="200"
|
||||
|
||||
LITELLM_HOST="192.168.68.116"
|
||||
LITELLM_PORT="4000"
|
||||
LITELLM_PATH="/"
|
||||
LITELLM_EXPECTED="200|401"
|
||||
|
||||
# PVE API: probe REAL PVE nodes, never the monitoring host (CT 116)
|
||||
PVE_NODES=("192.168.68.9" "192.168.68.12" "192.168.68.6" "192.168.68.15" "192.168.68.5")
|
||||
PVE_API_PORT="8006"
|
||||
PVE_API_SCHEME="https"
|
||||
PVE_API_EXPECTED="200|401"
|
||||
PVE_API_USE_K="1" # self-signed certs
|
||||
|
||||
# GPU exporters
|
||||
GPU_HOSTS=("192.168.68.8" "192.168.68.110" "192.168.68.15")
|
||||
GPU_PORT="9400"
|
||||
GPU_PATH="/metrics"
|
||||
GPU_EXPECTED="200"
|
||||
|
||||
# Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116
|
||||
DOCKER_STATS_HOST="192.168.68.116"
|
||||
DOCKER_STATS_PORT="9323"
|
||||
DOCKER_STATS_PATH="/"
|
||||
DOCKER_STATS_EXPECTED="200|404"
|
||||
|
||||
PVE_EXPORTER_HOST="192.168.68.116"
|
||||
PVE_EXPORTER_PORT="9324"
|
||||
PVE_EXPORTER_PATH="/"
|
||||
PVE_EXPORTER_EXPECTED="200|404"
|
||||
|
||||
# PM2 (CT 100)
|
||||
PM2_HOST="192.168.68.24"
|
||||
PM2_EXPECTED="online"
|
||||
|
||||
# ── Probe Functions ─────────────────────────────────────────────────────────
|
||||
|
||||
# probe_http <host> <port> <path> <expected_pattern> [use_k] [ssh_host] [scheme]
|
||||
# Returns: 0 if any HTTP status matches, 1 if probe-failed
|
||||
probe_http() {
|
||||
local host="$1" port="$2" path="$3" expected="$4" use_k="${5:-}" ssh_host="${6:-}"
|
||||
local scheme="${7:-http}"; local url="${scheme}://${host}:${port}${path}"
|
||||
local curl_opts=(-s -o /dev/null -w '%{http_code}' --max-time 10)
|
||||
local code=""
|
||||
|
||||
if [ -n "$use_k" ]; then
|
||||
curl_opts+=(-k)
|
||||
fi
|
||||
|
||||
if [ -n "$ssh_host" ]; then
|
||||
# Probe via SSH to the host where the service binds to 127.0.0.1
|
||||
code=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --max-time 10 ${use_k:+-k} ${scheme}://127.0.0.1:${port}${path}" 2>/dev/null) || code="000"
|
||||
else
|
||||
code=$(curl "${curl_opts[@]}" "$url" 2>/dev/null) || code="000"
|
||||
fi
|
||||
|
||||
# Clean up the code
|
||||
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||
[ -n "$code" ] || code="000"
|
||||
|
||||
# Check if code matches expected pattern
|
||||
if echo "$code" | grep -qE "^(${expected})$"; then
|
||||
return 0
|
||||
else
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
|
||||
# ── Main ────────────────────────────────────────────────────────────────────
|
||||
|
||||
FAILED=()
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
echo "=== Infrastructure Monitoring — $TIMESTAMP ==="
|
||||
|
||||
# Grafana
|
||||
if probe_http "$GRAFANA_HOST" "$GRAFANA_PORT" "$GRAFANA_PATH" "$GRAFANA_EXPECTED"; then
|
||||
echo " ✅ Grafana: alive"
|
||||
else
|
||||
echo " 🔴 Grafana: probe-failed: ${GRAFANA_HOST}:${GRAFANA_PORT} (expected ${GRAFANA_EXPECTED})"
|
||||
FAILED+=("grafana")
|
||||
fi
|
||||
|
||||
# Prometheus
|
||||
if probe_http "$PROMETHEUS_HOST" "$PROMETHEUS_PORT" "$PROMETHEUS_PATH" "$PROMETHEUS_EXPECTED"; then
|
||||
echo " ✅ Prometheus: alive"
|
||||
else
|
||||
echo " 🔴 Prometheus: probe-failed: ${PROMETHEUS_HOST}:${PROMETHEUS_PORT} (expected ${PROMETHEUS_EXPECTED})"
|
||||
FAILED+=("prometheus")
|
||||
fi
|
||||
|
||||
# LiteLLM
|
||||
if probe_http "$LITELLM_HOST" "$LITELLM_PORT" "$LITELLM_PATH" "$LITELLM_EXPECTED"; then
|
||||
echo " ✅ LiteLLM: alive"
|
||||
else
|
||||
echo " 🔴 LiteLLM: probe-failed: ${LITELLM_HOST}:${LITELLM_PORT} (expected ${LITELLM_EXPECTED})"
|
||||
FAILED+=("litellm")
|
||||
fi
|
||||
|
||||
# PVE API (5 nodes)
|
||||
PVE_FAILED=()
|
||||
for node in "${PVE_NODES[@]}"; do
|
||||
if probe_http "$node" "$PVE_API_PORT" "/" "$PVE_API_EXPECTED" "$PVE_API_USE_K" "" "https"; then
|
||||
echo " ✅ PVE API ${node}: alive"
|
||||
else
|
||||
echo " 🔴 PVE API ${node}: probe-failed: ${node}:${PVE_API_PORT} (expected ${PVE_API_EXPECTED})"
|
||||
PVE_FAILED+=("$node")
|
||||
fi
|
||||
done
|
||||
if [ ${#PVE_FAILED[@]} -gt 0 ]; then
|
||||
FAILED+=("pve-api: ${PVE_FAILED[*]}")
|
||||
fi
|
||||
|
||||
# GPU exporters
|
||||
GPU_FAILED=()
|
||||
for host in "${GPU_HOSTS[@]}"; do
|
||||
if probe_http "$host" "$GPU_PORT" "$GPU_PATH" "$GPU_EXPECTED"; then
|
||||
echo " ✅ GPU ${host}: alive"
|
||||
else
|
||||
echo " 🔴 GPU ${host}: probe-failed: ${host}:${GPU_PORT} (expected ${GPU_EXPECTED})"
|
||||
GPU_FAILED+=("$host")
|
||||
fi
|
||||
done
|
||||
if [ ${#GPU_FAILED[@]} -gt 0 ]; then
|
||||
FAILED+=("gpu: ${GPU_FAILED[*]}")
|
||||
fi
|
||||
|
||||
# Docker Stats (localhost via SSH)
|
||||
if probe_http "$DOCKER_STATS_HOST" "$DOCKER_STATS_PORT" "$DOCKER_STATS_PATH" "$DOCKER_STATS_EXPECTED" "" "$DOCKER_STATS_HOST"; then
|
||||
echo " ✅ Docker Stats: alive"
|
||||
else
|
||||
echo " 🔴 Docker Stats: probe-failed: ${DOCKER_STATS_HOST}:${DOCKER_STATS_PORT} (expected ${DOCKER_STATS_EXPECTED})"
|
||||
FAILED+=("docker-stats")
|
||||
fi
|
||||
|
||||
# PVE Exporter (localhost via SSH)
|
||||
if probe_http "$PVE_EXPORTER_HOST" "$PVE_EXPORTER_PORT" "$PVE_EXPORTER_PATH" "$PVE_EXPORTER_EXPECTED" "" "$PVE_EXPORTER_HOST"; then
|
||||
echo " ✅ PVE Exporter: alive"
|
||||
else
|
||||
echo " 🔴 PVE Exporter: probe-failed: ${PVE_EXPORTER_HOST}:${PVE_EXPORTER_PORT} (expected ${PVE_EXPORTER_EXPECTED})"
|
||||
FAILED+=("pve-exporter")
|
||||
fi
|
||||
|
||||
# PM2
|
||||
PM2_OUTPUT=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${PM2_HOST}" \
|
||||
"pm2 list 2>/dev/null | grep -c 'online'" 2>/dev/null) || PM2_OUTPUT="0"
|
||||
if [ "$PM2_OUTPUT" -gt 0 ]; then
|
||||
echo " ✅ PM2: ${PM2_OUTPUT} processes online"
|
||||
else
|
||||
echo " 🔴 PM2: probe-failed: no processes online on ${PM2_HOST}"
|
||||
FAILED+=("pm2")
|
||||
fi
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
|
||||
echo ""
|
||||
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||
echo " ✅ All legs OK"
|
||||
exit 0
|
||||
else
|
||||
echo " 🔴 FAILED legs: ${FAILED[*]}"
|
||||
exit 1
|
||||
fi
|
||||
@@ -197,16 +197,16 @@ def check_admin_key_list():
|
||||
# Try to parse the response
|
||||
try:
|
||||
data = json.loads(stdout)
|
||||
# Response is a dict with "keys" field
|
||||
# Response is a dict with "keys" (paginated list) and "total_count" fields
|
||||
if isinstance(data, dict) and "keys" in data:
|
||||
key_count = len(data["keys"])
|
||||
key_count = data.get("total_count", len(data["keys"]))
|
||||
elif isinstance(data, list):
|
||||
key_count = len(data)
|
||||
else:
|
||||
key_count = 0
|
||||
if key_count == 0:
|
||||
return "Admin Key List", False, "admin-call-failed (empty response)"
|
||||
return "Admin Key List", True, str(key_count) + " keys"
|
||||
return "Admin Key List", True, str(key_count) + " total (" + str(len(data.get("keys", []) if isinstance(data, dict) else data)) + " on page 1)" if isinstance(data, dict) else str(key_count) + " total"
|
||||
except Exception as e:
|
||||
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
|
||||
|
||||
|
||||
@@ -7,9 +7,11 @@
|
||||
# agent leg is retired — see the note after the Tanko leg.
|
||||
set -euo pipefail
|
||||
|
||||
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
|
||||
# Never fall back to a literal key
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:?ZULIP_API_KEY not set — refusing to run with no credential}"
|
||||
ZULIP_SITE="https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
||||
OWNER_ZULIP_ID="9"
|
||||
|
||||
|
||||
@@ -27,12 +29,12 @@ notify() {
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||
> /dev/null 2>&1 \
|
||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
|
||||
@@ -41,7 +43,7 @@ notify() {
|
||||
# ── Global: Zulip Server ──
|
||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||
https://chat.sysloggh.net/api/v1/server_settings \
|
||||
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null) || SERVER_CODE="000"
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
|
||||
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
|
||||
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
|
||||
if [ "$SERVER_CODE" != "200" ]; then
|
||||
|
||||
@@ -27,7 +27,7 @@ Agent (Hermes/pi) ──curl + X-API-Key──► Stirling-PDF (:8989) ──►
|
||||
| Base URL | `http://192.168.68.7:8989` |
|
||||
| Auth Method | API Key (header) |
|
||||
| Header Name | `X-API-Key` |
|
||||
| API Key | `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88` |
|
||||
| API Key | `«vault: infrastructure/production STIRLING_API_KEY»` |
|
||||
| Key Source | `SECURITY_CUSTOMGLOBALAPIKEY` in `/opt/home_stack/docker-compose.yml` |
|
||||
| Swagger | `http://192.168.68.7:8989/swagger-ui.html` |
|
||||
| Health | `http://192.168.68.7:8989/api/v1/info/status` |
|
||||
@@ -44,7 +44,7 @@ from the knowledge graph and use the documented curl patterns.
|
||||
Direct bash invocations:
|
||||
```bash
|
||||
curl -X POST "http://192.168.68.7:8989/api/v1/split-pdf" \
|
||||
-H "X-API-Key: adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88" \
|
||||
-H "X-API-Key: «vault: infrastructure/production STIRLING_API_KEY»" \
|
||||
-F "fileInput=@/path/to/file.pdf" \
|
||||
-F "pageNumbers=1,2,3" \
|
||||
-o /tmp/output.zip
|
||||
|
||||
@@ -0,0 +1,165 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
test_infra_monitoring.py — Tests for infrastructure-monitoring.sh
|
||||
|
||||
Tests:
|
||||
1. Port drift detection: asserts each probed port matches the documented value
|
||||
2. PVE API probe targets: verifies we probe real PVE nodes, not the monitoring host
|
||||
3. TLS failure labeling: verifies we use -k for self-signed certs
|
||||
4. Happy path and broken target (already done via shell test)
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
|
||||
SCRIPT_PATH = "/root/abiba-workspace/projects/prose-contracts/scripts/infra-monitoring.sh"
|
||||
|
||||
def run_script():
|
||||
"""Run the monitoring script and return output + exit code."""
|
||||
result = subprocess.run(
|
||||
["bash", SCRIPT_PATH],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=60
|
||||
)
|
||||
return result.stdout, result.stderr, result.returncode
|
||||
|
||||
def test_happy_path():
|
||||
"""Test that all documented targets are probed and healthy."""
|
||||
print("=== Test 1: Happy Path ===")
|
||||
stdout, stderr, returncode = run_script()
|
||||
print(stdout)
|
||||
|
||||
# Verify exit code is 0
|
||||
if returncode != 0:
|
||||
print(f"❌ FAILED: Expected exit code 0, got {returncode}")
|
||||
return False
|
||||
|
||||
# Verify all legs passed
|
||||
if "All legs OK" not in stdout:
|
||||
print(f"❌ FAILED: Expected 'All legs OK' in output")
|
||||
return False
|
||||
|
||||
print("✅ PASSED: Happy path works")
|
||||
return True
|
||||
|
||||
def test_port_drift_detection():
|
||||
"""Test that port drift from documented values is detected."""
|
||||
print("\n=== Test 2: Port Drift Detection ===")
|
||||
|
||||
# Create a modified version with wrong port
|
||||
import shutil
|
||||
test_script = SCRIPT_PATH.replace("scripts/", "scripts/test-drift-")
|
||||
shutil.copy(SCRIPT_PATH, test_script)
|
||||
|
||||
# Change Grafana port from 3001 to 3099
|
||||
with open(test_script, 'r') as f:
|
||||
content = f.read()
|
||||
content = content.replace('GRAFANA_PORT="3001"', 'GRAFANA_PORT="3099"')
|
||||
|
||||
with open(test_script, 'w') as f:
|
||||
f.write(content)
|
||||
|
||||
# Run the modified script
|
||||
stdout, stderr, returncode = run_script()
|
||||
|
||||
# Clean up
|
||||
os.remove(test_script)
|
||||
|
||||
print(stdout)
|
||||
|
||||
# Verify:
|
||||
# 1. Script exited non-zero
|
||||
if returncode == 0:
|
||||
print(f"❌ FAILED: Expected non-zero exit code for drift, got {returncode}")
|
||||
return False
|
||||
|
||||
# 2. Output mentions probe-failed for grafana
|
||||
if "probe-failed" not in stdout.lower():
|
||||
print(f"❌ FAILED: Expected 'probe-failed' in output for Grafana")
|
||||
return False
|
||||
|
||||
# 3. Output mentions the wrong port
|
||||
if "192.168.68.116:3099" not in stdout:
|
||||
print(f"❌ FAILED: Expected '192.168.68.116:3099' in output")
|
||||
return False
|
||||
|
||||
print("✅ PASSED: Port drift detection works")
|
||||
return True
|
||||
|
||||
def test_pve_api_targets():
|
||||
"""Test that PVE API probes the real nodes, not the monitoring host."""
|
||||
print("\n=== Test 3: PVE API Targets ===")
|
||||
|
||||
stdout, stderr, returncode = run_script()
|
||||
|
||||
# Verify we're NOT probing the monitoring host (192.168.68.116) for PVE API
|
||||
if ":116:8006" in stdout or "192.168.68.116:8006" in stdout:
|
||||
print(f"❌ FAILED: Should not probe monitoring host (192.168.68.116) for PVE API")
|
||||
return False
|
||||
|
||||
# Verify we're probing real PVE nodes
|
||||
expected_nodes = ["192.168.68.9", "192.168.68.12", "192.168.68.6", "192.168.68.15", "192.168.68.5"]
|
||||
found_nodes = [node for node in expected_nodes if f"{node}:8006" in stdout]
|
||||
|
||||
if len(found_nodes) != 5:
|
||||
print(f"❌ FAILED: Expected to probe all 5 PVE nodes, found {len(found_nodes)}")
|
||||
print(f" Found: {found_nodes}")
|
||||
return False
|
||||
|
||||
print("✅ PASSED: PVE API probes correct targets")
|
||||
return True
|
||||
|
||||
def test_tls_handling():
|
||||
"""Test that PVE API uses -k flag for self-signed certs."""
|
||||
print("\n=== Test 4: TLS Handling ===")
|
||||
|
||||
# Check the script for -k flag usage
|
||||
with open(SCRIPT_PATH, 'r') as f:
|
||||
content = f.read()
|
||||
|
||||
# Verify -k is used for PVE API
|
||||
if 'use_k' not in content or 'PVE_API_USE_K' not in content:
|
||||
print(f"❌ FAILED: Script should use -k flag for PVE API (self-signed certs)")
|
||||
return False
|
||||
|
||||
# Verify PVE API probes use -k
|
||||
if 'PVE_API_USE_K="1"' not in content:
|
||||
print(f"❌ FAILED: PVE_API_USE_K should be set to 1")
|
||||
return False
|
||||
|
||||
print("✅ PASSED: TLS handling configured correctly")
|
||||
return True
|
||||
|
||||
def main():
|
||||
"""Run all tests."""
|
||||
tests = [
|
||||
("Happy Path", test_happy_path),
|
||||
("Port Drift Detection", test_port_drift_detection),
|
||||
("PVE API Targets", test_pve_api_targets),
|
||||
("TLS Handling", test_tls_handling),
|
||||
]
|
||||
|
||||
results = []
|
||||
for name, test_func in tests:
|
||||
try:
|
||||
result = test_func()
|
||||
results.append((name, result))
|
||||
except Exception as e:
|
||||
print(f"\n❌ EXCEPTION in {name}: {e}")
|
||||
results.append((name, False))
|
||||
|
||||
print("\n" + "=" * 60)
|
||||
print("SUMMARY")
|
||||
print("=" * 60)
|
||||
for name, result in results:
|
||||
status = "✅ PASSED" if result else "❌ FAILED"
|
||||
print(f"{name}: {status}")
|
||||
|
||||
all_passed = all(r for _, r in results)
|
||||
sys.exit(0 if all_passed else 1)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user