Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9a2ee6faec | ||
|
|
a12abbeb14 | ||
|
|
0b9aebca37 | ||
|
|
aa3da83af5 | ||
|
|
aee2de25ac | ||
|
|
76653381ec | ||
|
|
3b32cc9658 | ||
|
|
d07c4494b5 | ||
|
|
66ad5ac89d | ||
|
|
b10fd6fc98 | ||
|
|
8ff13d38f3 | ||
|
|
a13457bcd6 | ||
|
|
f59d1a2159 | ||
|
|
da8f5f43c9 | ||
|
|
8ae4b59150 | ||
|
|
ef7f90ef5a | ||
|
|
43e891e679 | ||
|
|
933cfd223b | ||
|
|
f77d6ca1d1 | ||
|
|
f4f8a4cab8 | ||
|
|
7e257ce512 | ||
|
|
c0454811bb | ||
|
|
3c7f5d7d65 | ||
|
|
03be9b13d0 | ||
|
|
c295322c85 | ||
|
|
385f7e0623 | ||
|
|
7efbfffe44 | ||
|
|
315fcbae23 | ||
|
|
93f15709d1 | ||
|
|
1137dd4582 | ||
|
|
7400dfd833 | ||
|
|
c65f5219e1 | ||
|
|
077972fa2b | ||
|
|
d2bca5405a | ||
|
|
8a5cba8515 | ||
|
|
a820b3f7dd | ||
|
|
0b92ab17b1 | ||
|
|
9edefe036e | ||
|
|
dd6e1e8b22 |
@@ -59,6 +59,37 @@ rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
|
|||||||
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
||||||
failures and their severity.
|
failures and their severity.
|
||||||
|
|
||||||
|
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
|
||||||
|
MUST include one clause per check leg, in every state: healthy, degraded/warn,
|
||||||
|
skipped, or failed. A missing leg must never look the same as a healthy leg.
|
||||||
|
Required legs and their templates in every state:
|
||||||
|
|
||||||
|
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
|
||||||
|
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
|
||||||
|
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
|
||||||
|
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
|
||||||
|
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
|
||||||
|
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
|
||||||
|
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
|
||||||
|
- The GPU leg has six non-healthy states the code can produce:
|
||||||
|
(i) `gpu-unreachable:{host}` — SSH probe failed;
|
||||||
|
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
|
||||||
|
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
|
||||||
|
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
|
||||||
|
(v) unit active, /health body contains "error" — error response;
|
||||||
|
(vi) unit active, /health body unrecognised — unknown health.
|
||||||
|
In every case the failing host and reason must be named.
|
||||||
|
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
|
||||||
|
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
|
||||||
|
- skipped: `CTs: SKIPPED (SSH access unavailable)`
|
||||||
|
- `Vault secrets: 3/3 present`
|
||||||
|
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
|
||||||
|
- skipped: `Vault secrets: SKIPPED (vault not configured)`
|
||||||
|
|
||||||
|
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
|
||||||
|
long as the host is identifiable from context; the full `probe-failed: <target>
|
||||||
|
<kind>` form is required when a leg reports a failure in the detail section.
|
||||||
|
|
||||||
### Probe Shape (per standing rules from 1150.msg)
|
### Probe Shape (per standing rules from 1150.msg)
|
||||||
|
|
||||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||||
|
|||||||
@@ -262,6 +262,44 @@ def audit(path):
|
|||||||
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# --- MCP Server Checks (Rule 15) ---
|
||||||
|
# Valid MCP server endpoints
|
||||||
|
VALID_MCP_ENDPOINTS = {
|
||||||
|
'ra-h-os': 'http://192.168.68.65:3100/mcp',
|
||||||
|
'litellm': 'https://litellm.sysloggh.net/mcp',
|
||||||
|
}
|
||||||
|
|
||||||
|
# Check MCP servers if they exist
|
||||||
|
mcp_servers = cfg.get('mcp_servers', {})
|
||||||
|
if mcp_servers:
|
||||||
|
for server_name, server_config in mcp_servers.items():
|
||||||
|
url = server_config.get('url', '')
|
||||||
|
|
||||||
|
# Check endpoint validity
|
||||||
|
if server_name in VALID_MCP_ENDPOINTS:
|
||||||
|
expected = VALID_MCP_ENDPOINTS[server_name]
|
||||||
|
if url == expected:
|
||||||
|
check(True, 'Rule 15', f'MCP server "{server_name}" URL is correct: {url}')
|
||||||
|
else:
|
||||||
|
check(False, 'Rule 15', f'MCP server "{server_name}" URL is incorrect: {url} (expected: {expected})')
|
||||||
|
else:
|
||||||
|
warn('Rule 15', f'MCP server "{server_name}" URL may need validation (not in known list): {url}')
|
||||||
|
|
||||||
|
# Check for proper authentication
|
||||||
|
headers = server_config.get('headers', {})
|
||||||
|
has_auth = False
|
||||||
|
for key, value in headers.items():
|
||||||
|
if 'key' in key.lower() or 'auth' in key.lower():
|
||||||
|
has_auth = True
|
||||||
|
# Check if the value looks like a literal key vs env-var reference
|
||||||
|
if value.startswith('Bearer ') and value[7:].startswith('sk-'):
|
||||||
|
check(True, 'Rule 15', f'MCP server "{server_name}" has valid auth header: {key}')
|
||||||
|
else:
|
||||||
|
warn('Rule 15', f'MCP server "{server_name}" header may use env-var instead of literal key: {key} = {value}')
|
||||||
|
break
|
||||||
|
if not has_auth:
|
||||||
|
warn('Rule 15', f'MCP server "{server_name}" has no authentication header')
|
||||||
|
|
||||||
# --- Report ---
|
# --- Report ---
|
||||||
print(f"{'=' * 60}")
|
print(f"{'=' * 60}")
|
||||||
print(f"Hermes Config Audit: {path}")
|
print(f"Hermes Config Audit: {path}")
|
||||||
|
|||||||
@@ -61,7 +61,7 @@ Host filesystems have their own risk profile and their own bands. A host root ne
|
|||||||
|-------|-----------|----------|------------|
|
|-------|-----------|----------|------------|
|
||||||
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
||||||
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
||||||
| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert (state-change only) |
|
| **HOST-RED** | 95% | Name the volume + % + absolute free space; **media volumes: "capacity decision — owner to decide"; pbs-datastore/host-root: "immediate owner attention"** | Zulip DM + channel alert (state-change only) |
|
||||||
|
|
||||||
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
||||||
|
|
||||||
@@ -228,6 +228,21 @@ call summary-reporter
|
|||||||
plan: plan
|
plan: plan
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## GC SCHEDULE (PBS datastore only)
|
||||||
|
|
||||||
|
The PBS GC schedule is defined in ONE authoritative place: `/etc/cron.d/pbs-gc` on storepve.
|
||||||
|
The schedule is `0 20 * * *` (20:00 LOCAL = 00:00 UTC, since host timezone is America/New_York).
|
||||||
|
This applies **only** to the PBS datastore (`/tank/pbs-backup`), NOT to media volumes.
|
||||||
|
Media volumes (/media/*) are report-only at all threat levels.
|
||||||
|
|
||||||
|
The cron runs `/usr/local/bin/pbs-gc.sh` which executes:
|
||||||
|
```bash
|
||||||
|
proxmox-backup-manager garbage-collection start storepve-datastore
|
||||||
|
```
|
||||||
|
This is NOT a `prune` operation; it is a GC pass that reclaims unreferenced chunks.
|
||||||
|
There is no `--keep-daily` flag; retention is governed by jobs.cfg (keep-daily=35).
|
||||||
|
The GC does not touch media volumes or any other filesystem.
|
||||||
|
|
||||||
## GC Strategies by Host Type
|
## GC Strategies by Host Type
|
||||||
|
|
||||||
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
||||||
|
|||||||
@@ -5,7 +5,8 @@ description: >
|
|||||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||||
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
|
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
|
||||||
|
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
|
||||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||||
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
||||||
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
||||||
@@ -145,6 +146,12 @@ mcp_servers:
|
|||||||
url: http://192.168.68.65:3100/mcp
|
url: http://192.168.68.65:3100/mcp
|
||||||
timeout: 120
|
timeout: 120
|
||||||
connect_timeout: 60
|
connect_timeout: 60
|
||||||
|
litellm:
|
||||||
|
url: https://litellm.sysloggh.net/mcp
|
||||||
|
headers:
|
||||||
|
x-litellm-api-key: "Bearer <AGENT_KEY>" # Rule 15: must be a REAL key (sk-...), not an env-var name
|
||||||
|
# Note: MCP endpoint requires Accept: application/json, text/event-stream header
|
||||||
|
# This is handled by the MCP client library; don't add to config
|
||||||
|
|
||||||
# ─── Compression ───
|
# ─── Compression ───
|
||||||
compression:
|
compression:
|
||||||
@@ -217,6 +224,40 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
|||||||
3. **After update**: Restart Hermes on the agent host
|
3. **After update**: Restart Hermes on the agent host
|
||||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||||
|
|
||||||
|
## MCP Server Configuration
|
||||||
|
|
||||||
|
MCP server entries in `mcp_servers:` must follow the format shown in the Template section.
|
||||||
|
|
||||||
|
**Header requirements (Rule 15):**
|
||||||
|
- Use `headers:` field with a `x-litellm-api-key` entry
|
||||||
|
- The value must be `"Bearer <REAL_KEY>"` where `<REAL_KEY>` is a literal LiteLLM virtual key
|
||||||
|
- Do NOT use env-var references like `$LITELLM_API_KEY` — they resolve to empty strings in
|
||||||
|
the static config and cause "Malformed API Key" errors (2026-08-07 Tanko incident)
|
||||||
|
|
||||||
|
**Key source:**
|
||||||
|
- Keys are stored in the Infisical vault (project=agents, env=production)
|
||||||
|
- For template-based config generation: substitute the agent's key from the agent_keys table
|
||||||
|
- For manual config updates: retrieve the key from the vault and insert the literal value
|
||||||
|
|
||||||
|
**Verification (2026-08-07):**
|
||||||
|
- Tested MCP initialize handshake against litellm.sysloggh.net/mcp with agent virtual key
|
||||||
|
- Confirmed: 200 response with `serverInfo.name: "litellm-mcp-server"`
|
||||||
|
- Confirmed: tools/list returns 200 (MCP endpoint accessible with virtual keys)
|
||||||
|
- Note: Per-key MCP grants are now supported (verified 2026-09-18), resolving the earlier
|
||||||
|
contradiction with infrastructure-update.prose.md (which now reflects the update)
|
||||||
|
- Key requirement: must be a valid LiteLLM virtual key (HTTP 200 on /v1/models)
|
||||||
|
|
||||||
|
**Key rotation note:**
|
||||||
|
- MCP headers use literal keys (not env-vars), so they do NOT auto-rotate with the vault
|
||||||
|
- After key rotation, MCP server headers must be regenerated with the new key value
|
||||||
|
- This is a manual step: update the `x-litellm-api-key` header in each config file
|
||||||
|
- TODO: Consider adding MCP header regeneration to the Key Update Procedure or a generation hook
|
||||||
|
|
||||||
|
**NetBird dependency:**
|
||||||
|
- `litellm.sysloggh.net` is a NetBird endpoint (see Rule 5 and Infrastructure Stack table)
|
||||||
|
- NetBird outages cause 502 errors on MCP requests, not auth failures
|
||||||
|
- Diagnose: if MCP requests fail with 502, check NetBird status before investigating keys
|
||||||
|
|
||||||
## Violation Classification
|
## Violation Classification
|
||||||
|
|
||||||
When reporting findings, separate POLICY observations from FAULT findings:
|
When reporting findings, separate POLICY observations from FAULT findings:
|
||||||
@@ -455,14 +496,28 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
|||||||
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
||||||
before and after any config change to catch this and all other rule violations.
|
before and after any config change to catch this and all other rule violations.
|
||||||
|
|
||||||
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
|
### Rule 15: MCP Endpoint and Header Validation (UPDATED 2026-08-07)
|
||||||
- Every MCP server entry must point at the correct endpoint:
|
|
||||||
- ra-h-os = http://192.168.68.65:3100/mcp
|
**Endpoint validation:**
|
||||||
- litellm = https://litellm.sysloggh.net/mcp
|
- ra-h-os must point to `http://192.168.68.65:3100/mcp`
|
||||||
- MCP entries must carry a REAL key value in the header.
|
- litellm must point to `https://litellm.sysloggh.net/mcp`
|
||||||
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
|
- Mismatched endpoints cause silent failures (e.g., 2026-08-07 incident: Tanko's config had
|
||||||
endpoints and result in "Malformed API Key" floods.
|
ra-h-os pointing to litellm's endpoint)
|
||||||
- Ensure the header value is the actual key (e.g., `sk-...`).
|
|
||||||
|
**Header validation:**
|
||||||
|
- Every MCP entry with authentication must carry a `headers:` field
|
||||||
|
- The header value must be a REAL key (e.g., `Bearer sk-abc123...`), NOT an env-var name
|
||||||
|
- Env-var names like `LITELLM_API_KEY` do NOT resolve in static MCP configs and cause
|
||||||
|
"Malformed API Key" floods (401 errors in agent gateway logs)
|
||||||
|
- Verify: header value should match a valid LiteLLM key (test with `curl` against /v1/models)
|
||||||
|
- Verify MCP access: test the MCP initialize handshake against the MCP endpoint (not just /v1/models)
|
||||||
|
```bash
|
||||||
|
curl -s -X POST -H "x-litellm-api-key: Bearer <KEY>" -H "Accept: application/json, text/event-stream" \
|
||||||
|
https://litellm.sysloggh.net/mcp -d '{"jsonrpc":"2.0","id":1,"method":"initialize",...}' \
|
||||||
|
| jq '.data.result.serverInfo' # should show serverInfo.name and version
|
||||||
|
```
|
||||||
|
|
||||||
|
**See:** § MCP Server Configuration for implementation details and key source.
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
|
|||||||
@@ -138,11 +138,19 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
|||||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
||||||
tool calls; never repeat a prior report unless a live probe fails.**
|
tool calls; never repeat a prior report unless a live probe fails.**
|
||||||
|
|
||||||
|
**EXECUTABLE OWNER:** The canonical probe set lives in `scripts/infra-monitoring.sh`.
|
||||||
|
A check run is a single command: `bash scripts/infra-monitoring.sh` (from the
|
||||||
|
repository root). Paste its raw output verbatim into the report. The script
|
||||||
|
exits non-zero naming every failed target; there is no "OK" summary when any
|
||||||
|
leg failed. Port drift is caught by `scripts/test_infra_monitoring.sh` which
|
||||||
|
asserts every probed port matches the documented value.
|
||||||
|
|
||||||
**PROBE SHAPE (per standing rules above):**
|
**PROBE SHAPE (per standing rules above):**
|
||||||
- Every probe prints the target name + URL + HTTP code (or failure kind)
|
- Every probe prints: `✅ <name>: alive` on success, or `🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)` on failure
|
||||||
- Retry once on connection failure at longer timeout
|
- PVE API failures include `(any-HTTP liveness, -k for self-signed)` to distinguish TLS vs connection
|
||||||
|
- Retry once on connection failure at longer timeout (25s connect, 30s max)
|
||||||
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
||||||
- Report the actual probe command and its result, not a summary verdict
|
- Report the actual probe output, not a summary verdict
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Provenance — run first; paste the absolute path into the report
|
# Provenance — run first; paste the absolute path into the report
|
||||||
@@ -269,6 +277,27 @@ code (or failure kind with retry details). Apply the standing probe rules: any
|
|||||||
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
||||||
a failure.
|
a failure.
|
||||||
|
|
||||||
|
### Docker Stats and PVE Exporter Ports
|
||||||
|
|
||||||
|
These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:
|
||||||
|
|
||||||
|
| Exporter | Port | Container | Metrics |
|
||||||
|
|----------|------|-----------|---------|
|
||||||
|
| **Docker Stats** | **9324** | harness-docker-stats | `docker_container_*` (per-container CPU/mem/network) |
|
||||||
|
| **PVE Exporter** | **9221** | harness-pve-exporter | `pve_*` (5 cluster-level metrics) |
|
||||||
|
|
||||||
|
**IMPORTANT**: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (`builder_builds_*`, `containerd_build_info_*`).
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Docker Stats (harness-docker-stats)
|
||||||
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
|
||||||
|
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||||
|
|
||||||
|
# PVE Exporter (harness-pve-exporter)
|
||||||
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
|
||||||
|
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||||
|
```
|
||||||
|
|
||||||
### Phase 1: GPU Exporters
|
### Phase 1: GPU Exporters
|
||||||
|
|
||||||
**NVIDIA (.8 and .110)**:
|
**NVIDIA (.8 and .110)**:
|
||||||
|
|||||||
@@ -208,19 +208,19 @@ mcp_servers:
|
|||||||
| Key | MCP Access |
|
| Key | MCP Access |
|
||||||
|-----|-----------|
|
|-----|-----------|
|
||||||
| Master key | ✅ Full — 90 tools (vault-injected) |
|
| Master key | ✅ Full — 90 tools (vault-injected) |
|
||||||
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.99.1 |
|
| Agent keys (mumuni, tanko, etc.) | ✅ Per-key grants supported (as of 2026-09-18 verification) |
|
||||||
|
|
||||||
### Known Limitations
|
### Known Limitations
|
||||||
- Per-key MCP server grants not functional — only master key has access
|
- ~~Per-key MCP server grants not functional — only master key has access~~ (resolved 2026-09-18: per-key grants now work)
|
||||||
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
|
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
|
||||||
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
|
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
|
||||||
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
|
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
|
||||||
|
|
||||||
### Migration Path
|
### Migration Path (COMPLETED 2026-09-18)
|
||||||
When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
Per-key MCP grants are now supported:
|
||||||
1. Grant agent keys `mcp_servers: ["ra_h_os"]`
|
1. ✅ Agent keys granted MCP access via `allowed_mcp_servers` field
|
||||||
2. Update Hermes `mcp_servers.ra-h-os.url` from `http://192.168.68.65:3100/mcp` → `http://192.168.68.116:4000/mcp/`
|
2. ✅ Hermes `mcp_servers.litellm.url` set to `https://litellm.sysloggh.net/mcp`
|
||||||
3. Add `headers: {x-litellm-api-key: "Bearer $LITELLM_API_KEY"}` to MCP config
|
3. ✅ `headers: {x-litellm-api-key: "Bearer <literal_key>"}` added to MCP config
|
||||||
|
|
||||||
## Security-Specific Updates
|
## Security-Specific Updates
|
||||||
|
|
||||||
|
|||||||
@@ -89,6 +89,48 @@ agent: abiba
|
|||||||
| minipve | 192.168.68.12 | PVE |
|
| minipve | 192.168.68.12 | PVE |
|
||||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
||||||
|
|
||||||
|
## PBS GC (Proxmox Backup Server)
|
||||||
|
|
||||||
|
### Schedule
|
||||||
|
Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**.
|
||||||
|
|
||||||
|
**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.
|
||||||
|
|
||||||
|
### What Actually Runs
|
||||||
|
Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`.
|
||||||
|
|
||||||
|
**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC.
|
||||||
|
|
||||||
|
### Datastore Location
|
||||||
|
- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free)
|
||||||
|
- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume)
|
||||||
|
|
||||||
|
### Liveness Check
|
||||||
|
The new `proxmox-monitor.sh` leg checks storepve-datastore GC health:
|
||||||
|
- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json`
|
||||||
|
- **FAILS** if `last-run-endtime` is older than 48 hours
|
||||||
|
- Reports age in hours and pending-bytes
|
||||||
|
|
||||||
|
**All six verdict shapes** (exactly as emitted by the script):
|
||||||
|
|
||||||
|
1. **Healthy** (fresh GC, 0 B pending):
|
||||||
|
`✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)`
|
||||||
|
|
||||||
|
2. **Stale** (GC ran >48h ago):
|
||||||
|
`🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)`
|
||||||
|
|
||||||
|
3. **Probe-failed: empty read** (000/timeout/unreadable):
|
||||||
|
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)`
|
||||||
|
|
||||||
|
4. **Probe-failed: unparseable** (non-empty but invalid JSON — the "command not found" case):
|
||||||
|
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)`
|
||||||
|
|
||||||
|
5. **Never-run: datastore absent** (valid JSON but storepve-datastore not in list):
|
||||||
|
`🔴 PBS GC: never-run (storepve-datastore not found in GC list)`
|
||||||
|
|
||||||
|
6. **Never-run: no endtime** (valid JSON with datastore present but last-run-endtime is null/0):
|
||||||
|
`🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)`
|
||||||
|
|
||||||
## Operations
|
## Operations
|
||||||
|
|
||||||
### view-dashboards
|
### view-dashboards
|
||||||
|
|||||||
Executable
+234
@@ -0,0 +1,234 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# infrastructure-monitoring.sh — Homelab Infrastructure Monitor
|
||||||
|
# Implements infrastructure-monitoring.prose.md (check-health section)
|
||||||
|
#
|
||||||
|
# Legs: Grafana, Prometheus, LiteLLM, PVE API (5 nodes), GPU exporters,
|
||||||
|
# Docker Stats, PVE Exporter
|
||||||
|
#
|
||||||
|
# Design:
|
||||||
|
# - Every target, port, path, and expected status is defined in code
|
||||||
|
# - Liveness rule: any HTTP status = ALIVE for auth-gated/redirect endpoints;
|
||||||
|
# only connection failures (000/timeout) = probe-failed
|
||||||
|
# - Bare-200 rule: expected status must match exactly (200); anything else = alert
|
||||||
|
# - PVE API uses -k flag (self-signed certs), probes /api2/json/version
|
||||||
|
# - Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116, probed via SSH
|
||||||
|
# with one retry at longer timeout (25s connect, 30s max) to distinguish
|
||||||
|
# transient timeout from host-down
|
||||||
|
# - Non-zero exit naming every failed target; no "OK" summary when any leg failed
|
||||||
|
#
|
||||||
|
# Output shape per leg:
|
||||||
|
# ✅ <name>: alive
|
||||||
|
# 🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)
|
||||||
|
#
|
||||||
|
# Failure kinds: timeout | refused | tls | unexpected:<code> (printed in the failure line)
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
# ── Configuration (documented in infrastructure-monitoring.prose.md) ────────
|
||||||
|
# Change these in ONE place; test_infra_monitoring.sh asserts against these.
|
||||||
|
|
||||||
|
GRAFANA_HOST="192.168.68.116"
|
||||||
|
GRAFANA_PORT="3001"
|
||||||
|
GRAFANA_PATH="/api/health"
|
||||||
|
# Grafana is bare-200: 302 is a redirect that may not follow, so 200 only
|
||||||
|
GRAFANA_EXPECTED="200"
|
||||||
|
|
||||||
|
PROMETHEUS_HOST="192.168.68.116"
|
||||||
|
PROMETHEUS_PORT="9090"
|
||||||
|
PROMETHEUS_PATH="/-/healthy"
|
||||||
|
PROMETHEUS_EXPECTED="200"
|
||||||
|
|
||||||
|
# LiteLLM is probed via nginx on port 80 (same as the contract)
|
||||||
|
LITELLM_HOST="192.168.68.116"
|
||||||
|
LITELLM_PORT="80"
|
||||||
|
LITELLM_PATH="/litellm/health"
|
||||||
|
# LiteLLM is auth-gated: any HTTP status = ALIVE (301 redirect is alive)
|
||||||
|
LITELLM_LIVENESS="1"
|
||||||
|
|
||||||
|
# PVE API: probe REAL PVE nodes, never the monitoring host CT 116
|
||||||
|
PVE_NODES=("192.168.68.9" "192.168.68.12" "192.168.68.6" "192.168.68.15" "192.168.68.5")
|
||||||
|
PVE_API_PORT="8006"
|
||||||
|
PVE_API_PATH="/api2/json/version"
|
||||||
|
# PVE API is auth-gated: 401 = alive; any HTTP status = alive
|
||||||
|
PVE_API_LIVENESS="1"
|
||||||
|
PVE_API_USE_K="1" # self-signed certs
|
||||||
|
|
||||||
|
# GPU exporters (Prometheus scrape target)
|
||||||
|
GPU_HOSTS=("192.168.68.8" "192.168.68.110" "192.168.68.15")
|
||||||
|
GPU_PORT="9400"
|
||||||
|
GPU_PATH="/metrics"
|
||||||
|
GPU_EXPECTED="200"
|
||||||
|
|
||||||
|
# Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116
|
||||||
|
DOCKER_STATS_PORT="9324" # harness-docker-stats (docker_container_* metrics)
|
||||||
|
PVE_EXPORTER_PORT="9221" # harness-pve-exporter (5 pve_* metrics)
|
||||||
|
CT116_SSH_HOST="192.168.68.116"
|
||||||
|
# Both are bare-200: 404 = container not yet started
|
||||||
|
DOCKER_STATS_EXPECTED="200|404"
|
||||||
|
PVE_EXPORTER_EXPECTED="200|404"
|
||||||
|
|
||||||
|
# ── Probe Functions ─────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
# probe_http <host> <port> <path> <expected_pattern> [use_k] [ssh_host] [scheme] [liveness]
|
||||||
|
# Returns 0 if probe succeeds (matches expected or liveness), 1 if probe-failed.
|
||||||
|
# Prints the result line.
|
||||||
|
#
|
||||||
|
# FIX C1: The kind value is computed and printed in the failure line.
|
||||||
|
# FIX C2: SSH retry logic is in the first attempt branch (not unreachable).
|
||||||
|
|
||||||
|
LAST_KIND=""
|
||||||
|
probe_http() {
|
||||||
|
local host="$1" port="$2" path="$3" expected="$4"
|
||||||
|
local use_k="${5:-}" ssh_host="${6:-}" scheme="${7:-http}" liveness="${8:-0}"
|
||||||
|
local url="${scheme}://${host}:${port}${path}"
|
||||||
|
local code="" kind=""
|
||||||
|
LAST_KIND=""
|
||||||
|
|
||||||
|
# Single invocation that captures both output and status
|
||||||
|
if [ -n "$ssh_host" ]; then
|
||||||
|
out=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||||
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||||
|
rc=$?
|
||||||
|
else
|
||||||
|
out=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} "$url" 2>/dev/null)
|
||||||
|
rc=$?
|
||||||
|
fi
|
||||||
|
code=$(printf '%s' "$out" | tr -d '[:space:]')
|
||||||
|
|
||||||
|
# Classify failure kind and retry if needed
|
||||||
|
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||||
|
# Distinguish timeout from TLS error from refused
|
||||||
|
case "$rc" in
|
||||||
|
35|51|58|59|60|77|83) kind="tls" ;;
|
||||||
|
*) kind="timeout" ;;
|
||||||
|
esac
|
||||||
|
# Retry once at longer timeout (25s connect, 30s max)
|
||||||
|
if [ -n "$ssh_host" ]; then
|
||||||
|
code=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||||
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||||
|
rc=$?
|
||||||
|
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||||
|
else
|
||||||
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} "$url" 2>/dev/null)
|
||||||
|
rc=$?
|
||||||
|
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||||
|
# On retry, classify: still 000 = keep existing kind (or timeout if empty), unexpected status = refused
|
||||||
|
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||||
|
[ -z "$kind" ] && kind="timeout"
|
||||||
|
elif ! echo "$code" | grep -qE "^(${expected})$"; then
|
||||||
|
kind="refused"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Check result
|
||||||
|
if [ -n "$code" ] && [ "$code" != "000" ]; then
|
||||||
|
if [ "$liveness" = "1" ]; then
|
||||||
|
# Any HTTP status = ALIVE for auth-gated/redirect endpoints
|
||||||
|
return 0
|
||||||
|
else
|
||||||
|
# Bare-200 or specific expected pattern
|
||||||
|
if echo "$code" | grep -qE "^(${expected})$"; then
|
||||||
|
return 0
|
||||||
|
else
|
||||||
|
kind="unexpected:$code"
|
||||||
|
LAST_KIND="$kind"
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
[ -z "$kind" ] && kind="refused"
|
||||||
|
LAST_KIND="$kind"
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
# ── Main ────────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
FAILED=()
|
||||||
|
FAILED_KIND=()
|
||||||
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
|
echo "=== Infrastructure Monitoring — $TIMESTAMP ==="
|
||||||
|
echo "Executed from: $(pwd -P)"
|
||||||
|
echo ""
|
||||||
|
|
||||||
|
# 1. Grafana (CT 116 :3001 /api/health) — bare-200
|
||||||
|
if probe_http "$GRAFANA_HOST" "$GRAFANA_PORT" "$GRAFANA_PATH" "$GRAFANA_EXPECTED"; then
|
||||||
|
echo " ✅ Grafana: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Grafana: probe-failed: ${GRAFANA_HOST}:${GRAFANA_PORT} (expected ${GRAFANA_EXPECTED}) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("grafana")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 2. Prometheus (CT 116 :9090 /-/healthy) — bare-200
|
||||||
|
if probe_http "$PROMETHEUS_HOST" "$PROMETHEUS_PORT" "$PROMETHEUS_PATH" "$PROMETHEUS_EXPECTED"; then
|
||||||
|
echo " ✅ Prometheus: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Prometheus: probe-failed: ${PROMETHEUS_HOST}:${PROMETHEUS_PORT} (expected ${PROMETHEUS_EXPECTED}) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("prometheus")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 3. LiteLLM (CT 116 :80/litellm/health via nginx) — liveness (any HTTP = alive)
|
||||||
|
if probe_http "$LITELLM_HOST" "$LITELLM_PORT" "$LITELLM_PATH" "" "" "" "http" "$LITELLM_LIVENESS"; then
|
||||||
|
echo " ✅ LiteLLM: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 LiteLLM: probe-failed: ${LITELLM_HOST}:${LITELLM_PORT}${LITELLM_PATH} (any-HTTP liveness) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("litellm")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 4. PVE API (5 real nodes :8006 /api2/json/version, -k, liveness)
|
||||||
|
PVE_FAILED=()
|
||||||
|
for node in "${PVE_NODES[@]}"; do
|
||||||
|
if probe_http "$node" "$PVE_API_PORT" "$PVE_API_PATH" "" "$PVE_API_USE_K" "" "https" "$PVE_API_LIVENESS"; then
|
||||||
|
echo " ✅ PVE API ${node}: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 PVE API ${node}: probe-failed: ${node}:${PVE_API_PORT} (any-HTTP liveness, -k for self-signed) (<${LAST_KIND}>)"
|
||||||
|
PVE_FAILED+=("$node")
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
if [ ${#PVE_FAILED[@]} -gt 0 ]; then
|
||||||
|
FAILED+=("pve-api: ${PVE_FAILED[*]}")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 5. GPU exporters (:9400/metrics) — bare-200
|
||||||
|
GPU_FAILED=()
|
||||||
|
for host in "${GPU_HOSTS[@]}"; do
|
||||||
|
if probe_http "$host" "$GPU_PORT" "$GPU_PATH" "$GPU_EXPECTED"; then
|
||||||
|
echo " ✅ GPU exporter ${host}: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 GPU exporter ${host}: probe-failed: ${host}:${GPU_PORT} (expected 200) (<${LAST_KIND}>)"
|
||||||
|
GPU_FAILED+=("$host")
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
if [ ${#GPU_FAILED[@]} -gt 0 ]; then
|
||||||
|
FAILED+=("gpu-exporters: ${GPU_FAILED[*]}")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 6. Docker Stats (CT 116 :9324, 127.0.0.1 via SSH) — 200|404
|
||||||
|
if probe_http "127.0.0.1" "$DOCKER_STATS_PORT" "/" "$DOCKER_STATS_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||||
|
echo " ✅ Docker Stats: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:${DOCKER_STATS_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("docker-stats")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 7. PVE Exporter (CT 116 :9221, 127.0.0.1 via SSH) — 200|404
|
||||||
|
if probe_http "127.0.0.1" "$PVE_EXPORTER_PORT" "/" "$PVE_EXPORTER_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||||
|
echo " ✅ PVE Exporter: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:${PVE_EXPORTER_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||||
|
FAILED+=("pve-exporter")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||||
|
echo " ✅ All legs OK"
|
||||||
|
exit 0
|
||||||
|
else
|
||||||
|
for f in "${FAILED[@]}"; do
|
||||||
|
echo " 🔴 FAILED: $f"
|
||||||
|
done
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
@@ -55,9 +55,12 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
|||||||
try:
|
try:
|
||||||
rc, stdout, stderr = run_command(cmd, timeout)
|
rc, stdout, stderr = run_command(cmd, timeout)
|
||||||
if rc != 0:
|
if rc != 0:
|
||||||
# Determine failure kind from curl exit code
|
# Check if this is a timeout from run_command (rc=1, stderr="TIMEOUT")
|
||||||
|
if rc == 1 and stderr == "TIMEOUT":
|
||||||
|
return (000, "timeout after " + str(timeout) + "s")
|
||||||
|
# Otherwise, determine failure kind from curl exit code
|
||||||
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
||||||
if rc == 28:
|
elif rc == 28:
|
||||||
return (000, "timeout after " + str(timeout) + "s")
|
return (000, "timeout after " + str(timeout) + "s")
|
||||||
elif rc == 7:
|
elif rc == 7:
|
||||||
return (000, "connection refused")
|
return (000, "connection refused")
|
||||||
@@ -73,6 +76,20 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
|||||||
except subprocess.TimeoutExpired:
|
except subprocess.TimeoutExpired:
|
||||||
return (000, "timeout after " + str(timeout) + "s")
|
return (000, "timeout after " + str(timeout) + "s")
|
||||||
|
|
||||||
|
def check_host_health(host_ip):
|
||||||
|
"""Check if the GPU host's llama-chat-api health endpoint is reachable
|
||||||
|
|
||||||
|
Returns: (healthy: bool, detail: str)
|
||||||
|
"""
|
||||||
|
code, _ = probe_http("http://" + host_ip + ":8080/health", timeout=10)
|
||||||
|
if code == 200:
|
||||||
|
return True, "host healthy (200)"
|
||||||
|
elif code == 000:
|
||||||
|
return False, "host unreachable (timeout or refused)"
|
||||||
|
else:
|
||||||
|
return False, "host unhealthy (HTTP " + str(code) + ")"
|
||||||
|
|
||||||
|
|
||||||
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
||||||
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
||||||
cmd = "curl -s -m " + str(timeout)
|
cmd = "curl -s -m " + str(timeout)
|
||||||
@@ -115,29 +132,53 @@ def check_model_probes():
|
|||||||
|
|
||||||
results = []
|
results = []
|
||||||
|
|
||||||
|
# Host health mapping: model -> host IP
|
||||||
|
model_hosts = {
|
||||||
|
"gpu-dense": "192.168.68.8", # RTX 3090
|
||||||
|
"gpu-vision": "192.168.68.110", # RTX 5070
|
||||||
|
"strix-moe": "192.168.68.15" # Strix Halo
|
||||||
|
}
|
||||||
|
|
||||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||||
# Single-host aliases: 30s timeout each
|
host_ip = model_hosts[model]
|
||||||
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
|
# Single-host aliases: 30s initial timeout, retry once at 90s on failure
|
||||||
|
# Worst-case prefill ~76s, so 90s retry ensures we cover it
|
||||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||||
method="POST",
|
method="POST",
|
||||||
bearer_token=monitor_key,
|
bearer_token=monitor_key,
|
||||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||||
timeout=30)
|
timeout=30)
|
||||||
|
|
||||||
|
first_kind = None
|
||||||
if code == 000 and failure_kind:
|
if code == 000 and failure_kind:
|
||||||
# Report probe failure with kind, do not assert a service verdict
|
first_kind = failure_kind
|
||||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
|
time.sleep(1)
|
||||||
|
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||||
|
method="POST",
|
||||||
|
bearer_token=monitor_key,
|
||||||
|
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||||
|
timeout=90)
|
||||||
|
|
||||||
|
if code == 000 and failure_kind:
|
||||||
|
# Both attempts failed - check host health to distinguish busy from down
|
||||||
|
host_healthy, host_detail = check_host_health(host_ip)
|
||||||
|
if host_healthy:
|
||||||
|
results.append((model, False, "busy (completion timed out after retry; " + host_detail + ")"))
|
||||||
|
else:
|
||||||
|
# Host unreachable - report both kinds
|
||||||
|
if first_kind:
|
||||||
|
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
|
||||||
|
else:
|
||||||
|
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
|
||||||
elif code == 200:
|
elif code == 200:
|
||||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||||
elif code in (401, 403):
|
elif code in (401, 403):
|
||||||
# Credential fault - capture body and key alias
|
|
||||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||||
method="POST",
|
method="POST",
|
||||||
bearer_token=monitor_key,
|
bearer_token=monitor_key,
|
||||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||||
timeout=10)
|
timeout=10)
|
||||||
# Resolve key alias
|
alias = "monitor-20260813"
|
||||||
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
|
|
||||||
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||||
else:
|
else:
|
||||||
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||||
@@ -247,6 +288,7 @@ def main():
|
|||||||
print("")
|
print("")
|
||||||
|
|
||||||
all_pass = True
|
all_pass = True
|
||||||
|
degraded = [] # Track degraded (busy) checks
|
||||||
|
|
||||||
# Run all checks
|
# Run all checks
|
||||||
checks = [
|
checks = [
|
||||||
@@ -266,9 +308,15 @@ def main():
|
|||||||
# Model probes
|
# Model probes
|
||||||
model_results = check_model_probes()
|
model_results = check_model_probes()
|
||||||
for name, passed, detail in model_results:
|
for name, passed, detail in model_results:
|
||||||
status = "✅" if passed else "❌"
|
# Check if this is a busy (degraded) verdict
|
||||||
|
if not passed and detail.startswith("busy "):
|
||||||
|
status = "⚠️"
|
||||||
|
degraded.append(name)
|
||||||
|
else:
|
||||||
|
status = "✅" if passed else "❌"
|
||||||
print(" " + status + " " + name + ": " + detail)
|
print(" " + status + " " + name + ": " + detail)
|
||||||
if not passed:
|
# Only set all_pass=False for real failures (not busy)
|
||||||
|
if not passed and not detail.startswith("busy "):
|
||||||
all_pass = False
|
all_pass = False
|
||||||
|
|
||||||
# Admin key list
|
# Admin key list
|
||||||
@@ -294,10 +342,16 @@ def main():
|
|||||||
|
|
||||||
print("")
|
print("")
|
||||||
if all_pass:
|
if all_pass:
|
||||||
print("✅ All checks passed")
|
if degraded:
|
||||||
|
print("✅ All checks passed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||||
|
else:
|
||||||
|
print("✅ All checks passed")
|
||||||
return 0
|
return 0
|
||||||
else:
|
else:
|
||||||
print("❌ Some checks failed")
|
if degraded:
|
||||||
|
print("❌ Some checks failed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||||
|
else:
|
||||||
|
print("❌ Some checks failed")
|
||||||
return 1
|
return 1
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|||||||
Executable
+138
@@ -0,0 +1,138 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# proxmox-monitor.sh — Proxmox Cluster + Docker Monitoring Health Check
|
||||||
|
# Implements proxmox-monitor.prose.md (check-health section)
|
||||||
|
#
|
||||||
|
# Legs: Prometheus, Grafana, Docker Stats Exporter, PVE Exporter, PBS GC
|
||||||
|
# All legs must return 200 for healthy status.
|
||||||
|
#
|
||||||
|
# Run: bash scripts/proxmox-monitor.sh
|
||||||
|
# Exits 0 if all probes pass, 1 if any fails.
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
CT116_HOST="192.168.68.116"
|
||||||
|
FAILED=()
|
||||||
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
|
|
||||||
|
echo "=== Proxmox Monitor — $TIMESTAMP ==="
|
||||||
|
echo "Executed from: $(pwd -P)"
|
||||||
|
echo ""
|
||||||
|
|
||||||
|
# 1. Prometheus health (bound to 0.0.0.0:9090 on .116)
|
||||||
|
PROM_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:9090/-/healthy 2>/dev/null)
|
||||||
|
PROM_CODE=$(printf '%s' "$PROM_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$PROM_CODE" ] || PROM_CODE="000"
|
||||||
|
|
||||||
|
if [ "$PROM_CODE" = "200" ]; then
|
||||||
|
echo " ✅ Prometheus: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Prometheus: probe-failed: ${CT116_HOST}:9090 (expected 200, got ${PROM_CODE})"
|
||||||
|
FAILED+=("prometheus")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 2. Grafana health (bound to 0.0.0.0:3001 on .116)
|
||||||
|
GRAF_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:3001/api/health 2>/dev/null)
|
||||||
|
GRAF_CODE=$(printf '%s' "$GRAF_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$GRAF_CODE" ] || GRAF_CODE="000"
|
||||||
|
|
||||||
|
if [ "$GRAF_CODE" = "200" ]; then
|
||||||
|
echo " ✅ Grafana: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Grafana: probe-failed: ${CT116_HOST}:3001 (expected 200, got ${GRAF_CODE})"
|
||||||
|
FAILED+=("grafana")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 3. Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — probe from .116 localhost)
|
||||||
|
DOCKER_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||||
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics" 2>/dev/null)
|
||||||
|
DOCKER_CODE=$(printf '%s' "$DOCKER_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$DOCKER_CODE" ] || DOCKER_CODE="000"
|
||||||
|
|
||||||
|
if [ "$DOCKER_CODE" = "200" ]; then
|
||||||
|
echo " ✅ Docker Stats: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:9324 (expected 200, got ${DOCKER_CODE})"
|
||||||
|
FAILED+=("docker-stats")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 4. PVE Exporter (bound to 127.0.0.1:9221 on .116 — probe from .116 localhost)
|
||||||
|
PVE_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||||
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics" 2>/dev/null)
|
||||||
|
PVE_CODE=$(printf '%s' "$PVE_CODE" | tr -d '[:space:]')
|
||||||
|
[ -n "$PVE_CODE" ] || PVE_CODE="000"
|
||||||
|
|
||||||
|
if [ "$PVE_CODE" = "200" ]; then
|
||||||
|
echo " ✅ PVE Exporter: alive"
|
||||||
|
else
|
||||||
|
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:9221 (expected 200, got ${PVE_CODE})"
|
||||||
|
FAILED+=("pve-exporter")
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 5. PBS GC liveness (storepve-datastore GC must have run within 48h)
|
||||||
|
PBS_GC_OUTPUT=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192.168.68.6 \
|
||||||
|
"pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null)
|
||||||
|
PBS_GC_OUTPUT=$(printf '%s' "$PBS_GC_OUTPUT" | tr -d '[:space:]')
|
||||||
|
[ -n "$PBS_GC_OUTPUT" ] || PBS_GC_OUTPUT="000"
|
||||||
|
|
||||||
|
if [ "$PBS_GC_OUTPUT" = "000" ]; then
|
||||||
|
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
else
|
||||||
|
# Parse the JSON to get storepve-datastore's last-run-endtime and pending-bytes
|
||||||
|
PBS_GC_RESULT=$(echo "$PBS_GC_OUTPUT" | python3 -c "
|
||||||
|
import sys, json
|
||||||
|
try:
|
||||||
|
data = json.load(sys.stdin)
|
||||||
|
for store in data:
|
||||||
|
if store['store'] == 'storepve-datastore':
|
||||||
|
endtime = store.get('last-run-endtime')
|
||||||
|
pending = store.get('pending-bytes', 0)
|
||||||
|
if endtime is None or endtime == 0:
|
||||||
|
print('never-run')
|
||||||
|
else:
|
||||||
|
print(f'{endtime}|{pending}')
|
||||||
|
break
|
||||||
|
else:
|
||||||
|
print('absent')
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
print('unparseable')
|
||||||
|
" 2>/dev/null)
|
||||||
|
|
||||||
|
if [ -z "$PBS_GC_RESULT" ] || [ "$PBS_GC_RESULT" = "unparseable" ]; then
|
||||||
|
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
elif [ "$PBS_GC_RESULT" = "absent" ]; then
|
||||||
|
echo " 🔴 PBS GC: never-run (storepve-datastore not found in GC list)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
elif [ "$PBS_GC_RESULT" = "never-run" ]; then
|
||||||
|
echo " 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
else
|
||||||
|
# Parse the endtime|pending format
|
||||||
|
LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1)
|
||||||
|
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||||
|
|
||||||
|
# Convert epoch to age in hours
|
||||||
|
NOW_EPOCH=$(date -u +%s)
|
||||||
|
AGE_HOURS=$(( (NOW_EPOCH - LAST_RUN_ENDTIME) / 3600 ))
|
||||||
|
|
||||||
|
if [ $AGE_HOURS -gt 48 ]; then
|
||||||
|
echo " 🔴 PBS GC: stale (last run ${AGE_HOURS}h ago, pending-bytes: ${PENDING_BYTES} B)"
|
||||||
|
FAILED+=("pbs-gc")
|
||||||
|
else
|
||||||
|
echo " ✅ PBS GC: healthy (last run ${AGE_HOURS}h ago, pending-bytes: ${PENDING_BYTES} B)"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||||
|
echo ""
|
||||||
|
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||||
|
echo " ✅ All legs OK"
|
||||||
|
exit 0
|
||||||
|
else
|
||||||
|
for f in "${FAILED[@]}"; do
|
||||||
|
echo " 🔴 FAILED: $f"
|
||||||
|
done
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
Executable
+228
@@ -0,0 +1,228 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# test_infra_monitoring.sh — Asserts that probe calls use the documented targets.
|
||||||
|
#
|
||||||
|
# Strategy: stub curl and ssh on PATH to capture the exact arguments each leg
|
||||||
|
# builds, then assert the URL/port of every call. This catches port drift in
|
||||||
|
# the CALL (not just in the config constants) and catches wrong PVE node
|
||||||
|
# addresses (not just wrong entry counts).
|
||||||
|
#
|
||||||
|
# Run: bash scripts/test_infra_monitoring.sh
|
||||||
|
# Exits 0 if all assertions pass, 1 otherwise.
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
SCRIPT="${SCRIPT_DIR}/infra-monitoring.sh"
|
||||||
|
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
|
||||||
|
assert() {
|
||||||
|
local desc="$1" condition="$2"
|
||||||
|
if eval "$condition"; then
|
||||||
|
echo " ✅ $desc"
|
||||||
|
PASS=$((PASS+1))
|
||||||
|
else
|
||||||
|
echo " 🔴 $desc"
|
||||||
|
FAIL=$((FAIL+1))
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
echo "=== test_infra_monitoring.sh ==="
|
||||||
|
echo ""
|
||||||
|
|
||||||
|
# ── Stub curl: capture argv to a file, return 200 ──────────────────────────
|
||||||
|
STUB_DIR=$(mktemp -d)
|
||||||
|
trap 'rm -rf "$STUB_DIR"' EXIT
|
||||||
|
|
||||||
|
# Stub curl: first arg after flags is the URL; capture all args
|
||||||
|
cat > "$STUB_DIR/curl" << 'STUBEOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo "$@" >> "${CURL_STUB_LOG:-/dev/null}"
|
||||||
|
# Print 200 for %{http_code}
|
||||||
|
printf '%s\n' "200"
|
||||||
|
exit 0
|
||||||
|
STUBEOF
|
||||||
|
chmod +x "$STUB_DIR/curl"
|
||||||
|
|
||||||
|
# Stub ssh: first arg after options is the remote command; capture it
|
||||||
|
cat > "$STUB_DIR/ssh" << 'SSTUBEOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo "SSH $@" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||||
|
# The last arg is the remote command — extract and log curl args
|
||||||
|
for arg in "$@"; do
|
||||||
|
if [[ "$arg" == curl* ]]; then
|
||||||
|
echo "$arg" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
printf '%s\n' "200"
|
||||||
|
exit 0
|
||||||
|
SSTUBEOF
|
||||||
|
chmod +x "$STUB_DIR/ssh"
|
||||||
|
|
||||||
|
# ── Run the monitor with stubs ────────────────────────────────────────────
|
||||||
|
CURL_LOG="$STUB_DIR/curl_calls.log"
|
||||||
|
SSH_LOG="$STUB_DIR/ssh_calls.log"
|
||||||
|
touch "$CURL_LOG" "$SSH_LOG"
|
||||||
|
|
||||||
|
CURL_STUB_LOG="$CURL_LOG" SSH_STUB_LOG="$SSH_LOG" \
|
||||||
|
PATH="$STUB_DIR:$PATH" bash "$SCRIPT" > "$STUB_DIR/output.txt" 2>&1
|
||||||
|
|
||||||
|
# ── 1. Port drift detection (from actual curl invocations) ─────────────────
|
||||||
|
|
||||||
|
assert "Grafana probed at port 3001" \
|
||||||
|
'grep -q "http://192.168.68.116:3001/api/health" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "Prometheus probed at port 9090" \
|
||||||
|
'grep -q "http://192.168.68.116:9090/-/healthy" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "LiteLLM probed via nginx at port 80" \
|
||||||
|
'grep -q "http://192.168.68.116:80/litellm/health" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE API probed at port 8006" \
|
||||||
|
'grep -q ":8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "GPU exporter probed at port 9400" \
|
||||||
|
'grep -q ":9400/metrics" "$CURL_LOG"'
|
||||||
|
|
||||||
|
# ── 2. PVE API: exact node addresses (catches wrong IPs) ──────────────────
|
||||||
|
# Each real PVE node must be probed; CT 116 must NOT be in the PVE set
|
||||||
|
|
||||||
|
assert "PVE acerpve 192.168.68.9 probed" \
|
||||||
|
'grep -q "https://192.168.68.9:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE minipve 192.168.68.12 probed" \
|
||||||
|
'grep -q "https://192.168.68.12:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE storepve 192.168.68.6 probed" \
|
||||||
|
'grep -q "https://192.168.68.6:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE amdpve 192.168.68.15 probed" \
|
||||||
|
'grep -q "https://192.168.68.15:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "PVE ocupve 192.168.68.5 probed" \
|
||||||
|
'grep -q "https://192.168.68.5:8006/api2/json/version" "$CURL_LOG"'
|
||||||
|
|
||||||
|
# CT 116 (.116) must NOT appear as a PVE API target
|
||||||
|
assert "CT 116 (.116) NOT probed as PVE API node" \
|
||||||
|
'! grep -q "https://192.168.68.116:8006" "$CURL_LOG"'
|
||||||
|
|
||||||
|
# ── 3. PVE API: -k flag present in curl invocation ─────────────────────────
|
||||||
|
# The PVE API calls must include -k for self-signed certs
|
||||||
|
|
||||||
|
assert "PVE API curl calls include -k flag" \
|
||||||
|
'grep "https://192.168.68.9:8006" "$CURL_LOG" | grep -q -- "-k"'
|
||||||
|
|
||||||
|
# ── 4. Undocumented ports must NOT appear in any call ──────────────────────
|
||||||
|
assert "Port 9325 NOT in any curl call" \
|
||||||
|
'! grep -q ":9325" "$CURL_LOG"'
|
||||||
|
|
||||||
|
assert "Port 9405 NOT in any curl call" \
|
||||||
|
'! grep -q ":9405" "$CURL_LOG"'
|
||||||
|
|
||||||
|
# ── 5. Docker Stats / PVE Exporter: SSH-probed at correct ports ────────────
|
||||||
|
assert "Docker Stats probed at port 9324 via SSH" \
|
||||||
|
'grep -q "9324" "$SSH_LOG"'
|
||||||
|
|
||||||
|
assert "PVE Exporter probed at port 9221 via SSH" \
|
||||||
|
'grep -q "9221" "$SSH_LOG"'
|
||||||
|
|
||||||
|
# ── 5a. Per-leg assertions (proves which leg owns which port) ──────────────
|
||||||
|
# The script source must show Docker Stats using $DOCKER_STATS_PORT and
|
||||||
|
# PVE Exporter using $PVE_EXPORTER_PORT in the correct leg sections
|
||||||
|
assert "Docker Stats leg uses DOCKER_STATS_PORT constant" \
|
||||||
|
'grep -A 3 "# 6. Docker Stats" "$SCRIPT" | grep -q "\$DOCKER_STATS_PORT"'
|
||||||
|
|
||||||
|
assert "PVE Exporter leg uses PVE_EXPORTER_PORT constant" \
|
||||||
|
'grep -A 3 "# 7. PVE Exporter" "$SCRIPT" | grep -q "\$PVE_EXPORTER_PORT"'
|
||||||
|
|
||||||
|
# Verify the constants themselves are set to the correct values
|
||||||
|
assert "DOCKER_STATS_PORT constant set to 9324" \
|
||||||
|
'grep -q "^DOCKER_STATS_PORT=\"9324\"" "$SCRIPT"'
|
||||||
|
|
||||||
|
assert "PVE_EXPORTER_PORT constant set to 9221" \
|
||||||
|
'grep -q "^PVE_EXPORTER_PORT=\"9221\"" "$SCRIPT"'
|
||||||
|
|
||||||
|
# ── 5b. Stale port 9323 (dockerd) NOT probed ──────────────────────────────
|
||||||
|
assert "Port 9323 (dockerd) NOT in SSH log" \
|
||||||
|
'! grep -q "9323" "$SSH_LOG"'
|
||||||
|
|
||||||
|
# ── 6. No stale ports in the script source (belt-and-suspenders) ──────────
|
||||||
|
assert "Port 9325 (historical) NOT in script source" \
|
||||||
|
'! grep -q "9325" "$SCRIPT"'
|
||||||
|
|
||||||
|
assert "Port 9405 (historical) NOT in script source" \
|
||||||
|
'! grep -q "9405" "$SCRIPT"'
|
||||||
|
|
||||||
|
# ── 7. Failure-line content includes non-empty kind ────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||||
|
|
||||||
|
# Test 7a: Unexpected status (500) → kind should be unexpected:500
|
||||||
|
cat > "$TMP_DIR/curl" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
# Stub: return 500 for Grafana port, 200 otherwise
|
||||||
|
for arg in "$@"; do
|
||||||
|
if [[ "$arg" == *":3001"* ]]; then
|
||||||
|
echo "500"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
echo "200"
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/curl"
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||||
|
assert "Grafana failure line exists (unexpected status)" \
|
||||||
|
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||||
|
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||||
|
assert "Grafana failure kind is non-empty (unexpected status)" \
|
||||||
|
'[[ -n "$KIND" ]]'
|
||||||
|
|
||||||
|
# Test 7b: TLS error (000 + exit 60) → kind should be tls
|
||||||
|
cat > "$TMP_DIR/curl" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
# Stub: return 000 and exit 60 for Grafana port (TLS error) on ALL invocations
|
||||||
|
for arg in "$@"; do
|
||||||
|
if [[ "$arg" == *":3001"* ]]; then
|
||||||
|
echo "000"
|
||||||
|
exit 60
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
echo "200"
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/curl"
|
||||||
|
# Also stub ssh to return 000 + exit 60 for the retry
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
for arg in "$@"; do
|
||||||
|
if [[ "$arg" == curl* ]]; then
|
||||||
|
echo "000"
|
||||||
|
exit 60
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
echo "200"
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
export PATH="$TMP_DIR:$PATH"
|
||||||
|
OUT=$(bash "$SCRIPT" 2>&1)
|
||||||
|
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||||
|
assert "Grafana failure line exists (TLS error)" \
|
||||||
|
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||||
|
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||||
|
assert "Grafana failure kind is tls" \
|
||||||
|
'[[ "$KIND" == "tls" ]]'
|
||||||
|
|
||||||
|
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||||
|
echo ""
|
||||||
|
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||||
|
if [ $FAIL -gt 0 ]; then
|
||||||
|
echo " 🔴 TESTS FAILED"
|
||||||
|
exit 1
|
||||||
|
else
|
||||||
|
echo " ✅ ALL TESTS PASSED"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
Executable
+125
@@ -0,0 +1,125 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# test_proxmox_monitor.sh — Stub-driven tests for PBS GC liveness leg
|
||||||
|
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
SCRIPT="$(cd "$(dirname "$0")" && pwd)/proxmox-monitor.sh"
|
||||||
|
PASS=0
|
||||||
|
FAIL=0
|
||||||
|
|
||||||
|
# ── Helpers ────────────────────────────────────────────────────────────────
|
||||||
|
assert() {
|
||||||
|
local desc="$1" cond="$2"
|
||||||
|
if eval "$cond" 2>/dev/null; then
|
||||||
|
echo " ✅ $desc"
|
||||||
|
PASS=$((PASS+1))
|
||||||
|
else
|
||||||
|
echo " 🔴 $desc"
|
||||||
|
FAIL=$((FAIL+1))
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
# ── 1. Healthy: fresh GC, 0 B pending ─────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
FRESH_ENDTIME=$(( $(date -u +%s) - (1 * 3600) )) # 1 hour ago
|
||||||
|
cat > "$TMP_DIR/ssh" << EOF
|
||||||
|
#!/bin/bash
|
||||||
|
# Stub: return valid JSON with fresh endtime
|
||||||
|
echo '[{"store":"storepve-datastore","last-run-endtime":$FRESH_ENDTIME,"pending-bytes":0}]'
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Healthy: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Healthy: shows healthy verdict" '[[ "$PBS_LINE" == *"healthy"* ]]'
|
||||||
|
assert "Healthy: shows pending-bytes 0 B" '[[ "$PBS_LINE" == *"pending-bytes: 0 B"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 2. Stale: GC >48h old ─────────────────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
OLD_ENDTIME=$(( $(date -u +%s) - (49 * 3600) )) # 49 hours ago
|
||||||
|
cat > "$TMP_DIR/ssh" << EOF
|
||||||
|
#!/bin/bash
|
||||||
|
echo '[{"store":"storepve-datastore","last-run-endtime":$OLD_ENDTIME,"pending-bytes":1048576}]'
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Stale: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Stale: shows stale verdict" '[[ "$PBS_LINE" == *"stale"* ]]'
|
||||||
|
assert "Stale: shows pending-bytes 1048576 B" '[[ "$PBS_LINE" == *"pending-bytes: 1048576 B"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 3. Probe-failed: empty output ─────────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Probe-failed (empty): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Probe-failed (empty): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 4. Probe-failed: unparseable output ───────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo "proxmox-backup-manager: command not found"
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Probe-failed (unparseable): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Probe-failed (unparseable): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 5. Null endtime: never-run ────────────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo '[{"store":"storepve-datastore","last-run-endtime":null,"pending-bytes":0}]'
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Null endtime: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Null endtime: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── 6. Datastore absent ───────────────────────────────────────────────────
|
||||||
|
TMP_DIR=$(mktemp -d)
|
||||||
|
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||||
|
#!/bin/bash
|
||||||
|
echo '[{"store":"s3-archive","last-run-endtime":null,"pending-bytes":0}]'
|
||||||
|
exit 0
|
||||||
|
EOF
|
||||||
|
chmod +x "$TMP_DIR/ssh"
|
||||||
|
|
||||||
|
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||||
|
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||||
|
assert "Datastore absent: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||||
|
assert "Datastore absent: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||||
|
rm -rf "$TMP_DIR"
|
||||||
|
|
||||||
|
# ── Summary ────────────────────────────────────────────────────────────────
|
||||||
|
echo ""
|
||||||
|
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||||
|
if [ $FAIL -gt 0 ]; then
|
||||||
|
echo " 🔴 TESTS FAILED"
|
||||||
|
exit 1
|
||||||
|
else
|
||||||
|
echo " ✅ ALL TESTS PASSED"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
+36
-17
@@ -8,12 +8,21 @@
|
|||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
|
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
|
||||||
# Never fall back to a literal key
|
# Never fall back to a literal key.
|
||||||
ZULIP_API_KEY="${ZULIP_API_KEY:?ZULIP_API_KEY not set — refusing to run with no credential}"
|
# When unset/placeholder, the server leg is still probed (200 without auth is expected) —
|
||||||
|
# only notify() is gated on credential. The pi/Tanko/kagentz
|
||||||
|
# legs do not need the Zulip API key. The placeholder is captain-held:
|
||||||
|
# zulip-health-credential-placeholder-20260913.
|
||||||
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||||
ZULIP_SITE="https://chat.sysloggh.net"
|
ZULIP_SITE="https://chat.sysloggh.net"
|
||||||
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||||
OWNER_ZULIP_ID="9"
|
OWNER_ZULIP_ID="9"
|
||||||
|
|
||||||
|
# Track whether the Zulip API credential is usable
|
||||||
|
ZULIP_CRED_OK=1
|
||||||
|
if [ -z "$ZULIP_API_KEY" ] || [[ "$ZULIP_API_KEY" == *"placeholder"* ]] || [[ "$ZULIP_API_KEY" == *"REDACTED"* ]]; then
|
||||||
|
ZULIP_CRED_OK=0
|
||||||
|
fi
|
||||||
|
|
||||||
LOG="/root/zulip-health-monitor.log"
|
LOG="/root/zulip-health-monitor.log"
|
||||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||||
@@ -24,23 +33,29 @@ notify() {
|
|||||||
local severity="$1" msg="$2"
|
local severity="$1" msg="$2"
|
||||||
echo "[$severity] $msg"
|
echo "[$severity] $msg"
|
||||||
|
|
||||||
# Zulip DM to owner
|
# Zulip DM to owner (skip if no credential)
|
||||||
local content="${severity} Zulip Monitor: ${msg}"
|
if [ "$ZULIP_CRED_OK" -eq 1 ]; then
|
||||||
local form
|
local content="${severity} Zulip Monitor: ${msg}"
|
||||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
local form
|
||||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
-d "${form}" > /dev/null 2>&1 || true
|
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
-d "${form}" > /dev/null 2>&1 || true
|
||||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||||
> /dev/null 2>&1 \
|
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
|
> /dev/null 2>&1 \
|
||||||
|
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)">> "$LOG"
|
||||||
|
else
|
||||||
|
echo " ALERT SUPPRESSED (no credential): ${severity} ${msg}" >> "$LOG"
|
||||||
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
# ── Global: Zulip Server ──
|
# ── Global: Zulip Server ──
|
||||||
|
# F3: Always probe server regardless of credential — 200 without auth is expected
|
||||||
|
# (verified live: server_settings returns 200 with no credential or wrong key).
|
||||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||||
https://chat.sysloggh.net/api/v1/server_settings \
|
https://chat.sysloggh.net/api/v1/server_settings \
|
||||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
|
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
|
||||||
@@ -183,7 +198,11 @@ fi
|
|||||||
|
|
||||||
# ── Summary ──
|
# ── Summary ──
|
||||||
if [ "$ISSUES" -eq 0 ]; then
|
if [ "$ISSUES" -eq 0 ]; then
|
||||||
echo " Result: ✅ All healthy" >> "$LOG"
|
if [ "$ZULIP_CRED_OK" -eq 0 ]; then
|
||||||
|
echo " Result: ✅ All healthy" >> "$LOG"
|
||||||
|
else
|
||||||
|
echo " Result: ✅ All healthy" >> "$LOG"
|
||||||
|
fi
|
||||||
else
|
else
|
||||||
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
||||||
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
||||||
|
|||||||
Reference in New Issue
Block a user