Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8f1e5eebc4 | ||
|
|
bb1b65340e | ||
|
|
9e87927444 | ||
|
|
b0683e9566 | ||
|
|
dd11c8f14f | ||
|
|
0732eed329 | ||
|
|
59ed7cdbf7 | ||
|
|
5d70bbf25b | ||
|
|
ccc916d1ec | ||
|
|
48fb263d4b | ||
|
|
64790ebd19 | ||
|
|
6fa0a255df | ||
|
|
c72436b406 | ||
|
|
a17379676d | ||
|
|
63990b84f7 | ||
|
|
7a5ddb46a9 | ||
|
|
568fec2efa | ||
|
|
641a52c6da | ||
|
|
58033f39c5 | ||
|
|
6c59988e7f | ||
|
|
9a2ee6faec | ||
|
|
e71ded3c8c | ||
|
|
a12abbeb14 | ||
|
|
0b9aebca37 | ||
|
|
fa458afa26 | ||
|
|
aa3da83af5 | ||
|
|
aee2de25ac | ||
|
|
76653381ec | ||
|
|
3b32cc9658 | ||
|
|
d07c4494b5 | ||
|
|
66ad5ac89d | ||
|
|
b10fd6fc98 | ||
|
|
8ff13d38f3 | ||
|
|
a13457bcd6 | ||
|
|
f59d1a2159 | ||
|
|
da8f5f43c9 | ||
|
|
8ae4b59150 | ||
|
|
ef7f90ef5a | ||
|
|
43e891e679 | ||
|
|
933cfd223b | ||
|
|
f77d6ca1d1 | ||
|
|
f4f8a4cab8 | ||
|
|
7e257ce512 | ||
|
|
c0454811bb | ||
|
|
3c7f5d7d65 | ||
|
|
03be9b13d0 | ||
|
|
c295322c85 | ||
|
|
385f7e0623 | ||
|
|
7efbfffe44 | ||
|
|
315fcbae23 | ||
|
|
93f15709d1 | ||
|
|
1137dd4582 | ||
|
|
7400dfd833 | ||
|
|
c65f5219e1 | ||
|
|
077972fa2b | ||
|
|
d2bca5405a | ||
|
|
8a5cba8515 | ||
|
|
a820b3f7dd | ||
|
|
0b92ab17b1 | ||
|
|
9edefe036e | ||
|
|
dd6e1e8b22 |
@@ -71,6 +71,18 @@ jobs:
|
||||
git fetch origin "${{ gitea.ref }}" --depth=50
|
||||
git checkout "${{ gitea.sha }}"
|
||||
|
||||
- name: Committed-credential scan (secret guard)
|
||||
run: |
|
||||
# Fails the build on a credential-shaped string in the tree. Patterns
|
||||
# live in scripts/secret-patterns.tsv; the only tolerated literal
|
||||
# examples are in scripts/secret-allowlist.tsv, each with a reason.
|
||||
# Do not turn this into a warning: a warning in a stream nobody reads
|
||||
# is how six live credentials sat in this repo for weeks.
|
||||
bash scripts/secret-scan.sh
|
||||
|
||||
- name: Secret guard self-test
|
||||
run: bash tests/test_secret_scan.sh
|
||||
|
||||
- name: Structure + regression + consistency lint
|
||||
run: bash scripts/prose-lint.sh
|
||||
|
||||
|
||||
@@ -51,6 +51,12 @@ Two incidents taught us this:
|
||||
- `/grafana/` nginx route — was reverted Jul 2, must not reappear
|
||||
- `CT 122` or `CT 123` as CT ID labels — don't exist in the cluster
|
||||
- These rules are hardcoded in `scripts/prose-lint.sh`
|
||||
- **Committed-credential guard:** `scripts/secret-scan.sh` FAILS the build on
|
||||
credential-shaped strings (patterns in `scripts/secret-patterns.tsv`, prose
|
||||
included). Tolerated literals are listed one-per-example with a reason in
|
||||
`scripts/secret-allowlist.tsv`; never allowlist a live credential. It runs in
|
||||
the CI lint job, in `scripts/prose-lint.sh`, and via
|
||||
`bash scripts/secret-scan.sh --staged` before committing.
|
||||
|
||||
### Stage 3 — AI Review
|
||||
- Diff is sent to `syslog-auto` model via LiteLLM
|
||||
|
||||
@@ -59,6 +59,37 @@ rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
|
||||
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
||||
failures and their severity.
|
||||
|
||||
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
|
||||
MUST include one clause per check leg, in every state: healthy, degraded/warn,
|
||||
skipped, or failed. A missing leg must never look the same as a healthy leg.
|
||||
Required legs and their templates in every state:
|
||||
|
||||
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
|
||||
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
|
||||
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
|
||||
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
|
||||
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
|
||||
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
|
||||
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
|
||||
- The GPU leg has six non-healthy states the code can produce:
|
||||
(i) `gpu-unreachable:{host}` — SSH probe failed;
|
||||
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
|
||||
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
|
||||
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
|
||||
(v) unit active, /health body contains "error" — error response;
|
||||
(vi) unit active, /health body unrecognised — unknown health.
|
||||
In every case the failing host and reason must be named.
|
||||
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
|
||||
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
|
||||
- skipped: `CTs: SKIPPED (SSH access unavailable)`
|
||||
- `Vault secrets: 3/3 present`
|
||||
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
|
||||
- skipped: `Vault secrets: SKIPPED (vault not configured)`
|
||||
|
||||
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
|
||||
long as the host is identifiable from context; the full `probe-failed: <target>
|
||||
<kind>` form is required when a leg reports a failure in the detail section.
|
||||
|
||||
### Probe Shape (per standing rules from 1150.msg)
|
||||
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||
|
||||
+54
-6
@@ -114,12 +114,22 @@ def audit(path):
|
||||
)
|
||||
|
||||
# --- Rule 5: Main Config Base URL ---
|
||||
expected_base = "http://192.168.68.116/v1"
|
||||
check(
|
||||
model.get("base_url") == expected_base,
|
||||
"Rule 5",
|
||||
f"model.base_url must be {expected_base} (got {model.get('base_url')!r}) — /v1 not /litellm/v1",
|
||||
)
|
||||
# Canonical internal base (hermes-key-enforcement.prose.md:19/38/57/84/91/96)
|
||||
# and public base (serves /v1 only, per 2026-09-19 probe from CT 116).
|
||||
# The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only.
|
||||
# FAIL anything else (do not widen to accept any path ending in /v1).
|
||||
# Internal /v1 is non-canonical but working (authenticated via nginx), so WARN not FAIL.
|
||||
canonical_internal = "http://192.168.68.116/litellm/v1"
|
||||
public_host = "https://litellm.sysloggh.net/v1"
|
||||
non_canonical_internal = "http://192.168.68.116/v1"
|
||||
allowed_bases = (canonical_internal, public_host)
|
||||
actual_base = model.get("base_url")
|
||||
if actual_base in allowed_bases:
|
||||
check(True, "Rule 5", f"model.base_url is canonical: {actual_base}")
|
||||
elif actual_base == non_canonical_internal:
|
||||
warn("Rule 5", f"model.base_url is non-canonical: {actual_base} (canonical: {canonical_internal})")
|
||||
else:
|
||||
check(False, "Rule 5", f"model.base_url must be one of {allowed_bases} (got {actual_base!r})")
|
||||
|
||||
# --- Rule 6: max_tokens Is Required ---
|
||||
check(
|
||||
@@ -262,6 +272,44 @@ def audit(path):
|
||||
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
||||
)
|
||||
|
||||
# --- MCP Server Checks (Rule 15) ---
|
||||
# Valid MCP server endpoints
|
||||
VALID_MCP_ENDPOINTS = {
|
||||
'ra-h-os': 'http://192.168.68.65:3100/mcp',
|
||||
'litellm': 'https://litellm.sysloggh.net/mcp',
|
||||
}
|
||||
|
||||
# Check MCP servers if they exist
|
||||
mcp_servers = cfg.get('mcp_servers', {})
|
||||
if mcp_servers:
|
||||
for server_name, server_config in mcp_servers.items():
|
||||
url = server_config.get('url', '')
|
||||
|
||||
# Check endpoint validity
|
||||
if server_name in VALID_MCP_ENDPOINTS:
|
||||
expected = VALID_MCP_ENDPOINTS[server_name]
|
||||
if url == expected:
|
||||
check(True, 'Rule 15', f'MCP server "{server_name}" URL is correct: {url}')
|
||||
else:
|
||||
check(False, 'Rule 15', f'MCP server "{server_name}" URL is incorrect: {url} (expected: {expected})')
|
||||
else:
|
||||
warn('Rule 15', f'MCP server "{server_name}" URL may need validation (not in known list): {url}')
|
||||
|
||||
# Check for proper authentication
|
||||
headers = server_config.get('headers', {})
|
||||
has_auth = False
|
||||
for key, value in headers.items():
|
||||
if 'key' in key.lower() or 'auth' in key.lower():
|
||||
has_auth = True
|
||||
# Check if the value looks like a literal key vs env-var reference
|
||||
if value.startswith('Bearer ') and value[7:].startswith('sk-'):
|
||||
check(True, 'Rule 15', f'MCP server "{server_name}" has valid auth header: {key}')
|
||||
else:
|
||||
warn('Rule 15', f'MCP server "{server_name}" header may use env-var instead of literal key: {key} = {value}')
|
||||
break
|
||||
if not has_auth:
|
||||
warn('Rule 15', f'MCP server "{server_name}" has no authentication header')
|
||||
|
||||
# --- Report ---
|
||||
print(f"{'=' * 60}")
|
||||
print(f"Hermes Config Audit: {path}")
|
||||
|
||||
+55
-1
@@ -38,6 +38,7 @@ index:
|
||||
by_category:
|
||||
compliance:
|
||||
- hermes-key-enforcement
|
||||
- litellm-api-keys
|
||||
- hermes-config-template
|
||||
- hermes-agent-baseline
|
||||
monitoring:
|
||||
@@ -101,6 +102,7 @@ index:
|
||||
proxmox:
|
||||
- proxmox-monitor
|
||||
litellm:
|
||||
- litellm-api-keys
|
||||
- litellm-health
|
||||
- litellm-self-heal
|
||||
memory:
|
||||
@@ -628,7 +630,7 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 3.3.0
|
||||
version: 3.4.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/15 * * * *'
|
||||
@@ -1869,6 +1871,58 @@ contracts:
|
||||
drift_alerts: []
|
||||
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||
- name: litellm-api-keys
|
||||
file: litellm-api-keys.prose.md
|
||||
kind: function
|
||||
category: compliance
|
||||
sensitivity: critical
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: on_demand
|
||||
cadence: null
|
||||
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
timeout: 120
|
||||
requires: []
|
||||
protocol:
|
||||
- Load contract from prose-contracts/main
|
||||
- Retrieve master key from Infisical (project=infrastructure env=production)
|
||||
- Read live key-scoped model roster from CT 116 /v1/models
|
||||
- Create/rotate/verify the requested agent key with an EXPLICIT models list
|
||||
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
|
||||
verification:
|
||||
postconditions:
|
||||
- check: standard agent key is local-only
|
||||
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
|
||||
expect: 0 cloud models
|
||||
- check: key exists with correct alias
|
||||
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
|
||||
expect: 200 with matching alias
|
||||
artifact: key creation/rotation report
|
||||
receipt:
|
||||
format: json
|
||||
storage: ~/.hermes/runs/litellm-api-keys/
|
||||
graph_node: true
|
||||
escalation:
|
||||
info:
|
||||
action: log_to_receipt
|
||||
notify: []
|
||||
warning:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
critical:
|
||||
action: relay_alert
|
||||
notify:
|
||||
- abiba
|
||||
- mumuni
|
||||
- ops
|
||||
|
||||
koby_report_only: true
|
||||
koby_host: "CT 111 (tdunna)"
|
||||
koby_ip: ".129"
|
||||
|
||||
@@ -61,7 +61,7 @@ Host filesystems have their own risk profile and their own bands. A host root ne
|
||||
|-------|-----------|----------|------------|
|
||||
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
||||
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
||||
| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert (state-change only) |
|
||||
| **HOST-RED** | 95% | Name the volume + % + absolute free space; **media volumes: "capacity decision — owner to decide"; pbs-datastore/host-root: "immediate owner attention"** | Zulip DM + channel alert (state-change only) |
|
||||
|
||||
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
||||
|
||||
@@ -228,6 +228,21 @@ call summary-reporter
|
||||
plan: plan
|
||||
```
|
||||
|
||||
## GC SCHEDULE (PBS datastore only)
|
||||
|
||||
The PBS GC schedule is defined in ONE authoritative place: `/etc/cron.d/pbs-gc` on storepve.
|
||||
The schedule is `0 20 * * *` (20:00 LOCAL = 00:00 UTC, since host timezone is America/New_York).
|
||||
This applies **only** to the PBS datastore (`/tank/pbs-backup`), NOT to media volumes.
|
||||
Media volumes (/media/*) are report-only at all threat levels.
|
||||
|
||||
The cron runs `/usr/local/bin/pbs-gc.sh` which executes:
|
||||
```bash
|
||||
proxmox-backup-manager garbage-collection start storepve-datastore
|
||||
```
|
||||
This is NOT a `prune` operation; it is a GC pass that reclaims unreferenced chunks.
|
||||
There is no `--keep-daily` flag; retention is governed by jobs.cfg (keep-daily=35).
|
||||
The GC does not touch media volumes or any other filesystem.
|
||||
|
||||
## GC Strategies by Host Type
|
||||
|
||||
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
||||
|
||||
@@ -5,7 +5,8 @@ description: >
|
||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
|
||||
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
|
||||
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
|
||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
||||
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
||||
@@ -145,6 +146,12 @@ mcp_servers:
|
||||
url: http://192.168.68.65:3100/mcp
|
||||
timeout: 120
|
||||
connect_timeout: 60
|
||||
litellm:
|
||||
url: https://litellm.sysloggh.net/mcp
|
||||
headers:
|
||||
x-litellm-api-key: "Bearer <AGENT_KEY>" # Rule 15: must be a REAL key (sk-...), not an env-var name
|
||||
# Note: MCP endpoint requires Accept: application/json, text/event-stream header
|
||||
# This is handled by the MCP client library; don't add to config
|
||||
|
||||
# ─── Compression ───
|
||||
compression:
|
||||
@@ -217,6 +224,40 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
||||
3. **After update**: Restart Hermes on the agent host
|
||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||
|
||||
## MCP Server Configuration
|
||||
|
||||
MCP server entries in `mcp_servers:` must follow the format shown in the Template section.
|
||||
|
||||
**Header requirements (Rule 15):**
|
||||
- Use `headers:` field with a `x-litellm-api-key` entry
|
||||
- The value must be `"Bearer <REAL_KEY>"` where `<REAL_KEY>` is a literal LiteLLM virtual key
|
||||
- Do NOT use env-var references like `$LITELLM_API_KEY` — they resolve to empty strings in
|
||||
the static config and cause "Malformed API Key" errors (2026-08-07 Tanko incident)
|
||||
|
||||
**Key source:**
|
||||
- Keys are stored in the Infisical vault (project=agents, env=production)
|
||||
- For template-based config generation: substitute the agent's key from the agent_keys table
|
||||
- For manual config updates: retrieve the key from the vault and insert the literal value
|
||||
|
||||
**Verification (2026-08-07):**
|
||||
- Tested MCP initialize handshake against litellm.sysloggh.net/mcp with agent virtual key
|
||||
- Confirmed: 200 response with `serverInfo.name: "litellm-mcp-server"`
|
||||
- Confirmed: tools/list returns 200 (MCP endpoint accessible with virtual keys)
|
||||
- Note: Per-key MCP grants are now supported (verified 2026-09-18), resolving the earlier
|
||||
contradiction with infrastructure-update.prose.md (which now reflects the update)
|
||||
- Key requirement: must be a valid LiteLLM virtual key (HTTP 200 on /v1/models)
|
||||
|
||||
**Key rotation note:**
|
||||
- MCP headers use literal keys (not env-vars), so they do NOT auto-rotate with the vault
|
||||
- After key rotation, MCP server headers must be regenerated with the new key value
|
||||
- This is a manual step: update the `x-litellm-api-key` header in each config file
|
||||
- TODO: Consider adding MCP header regeneration to the Key Update Procedure or a generation hook
|
||||
|
||||
**NetBird dependency:**
|
||||
- `litellm.sysloggh.net` is a NetBird endpoint (see Rule 5 and Infrastructure Stack table)
|
||||
- NetBird outages cause 502 errors on MCP requests, not auth failures
|
||||
- Diagnose: if MCP requests fail with 502, check NetBird status before investigating keys
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
@@ -455,14 +496,28 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
|
||||
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
|
||||
before and after any config change to catch this and all other rule violations.
|
||||
|
||||
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
|
||||
- Every MCP server entry must point at the correct endpoint:
|
||||
- ra-h-os = http://192.168.68.65:3100/mcp
|
||||
- litellm = https://litellm.sysloggh.net/mcp
|
||||
- MCP entries must carry a REAL key value in the header.
|
||||
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
|
||||
endpoints and result in "Malformed API Key" floods.
|
||||
- Ensure the header value is the actual key (e.g., `sk-...`).
|
||||
### Rule 15: MCP Endpoint and Header Validation (UPDATED 2026-08-07)
|
||||
|
||||
**Endpoint validation:**
|
||||
- ra-h-os must point to `http://192.168.68.65:3100/mcp`
|
||||
- litellm must point to `https://litellm.sysloggh.net/mcp`
|
||||
- Mismatched endpoints cause silent failures (e.g., 2026-08-07 incident: Tanko's config had
|
||||
ra-h-os pointing to litellm's endpoint)
|
||||
|
||||
**Header validation:**
|
||||
- Every MCP entry with authentication must carry a `headers:` field
|
||||
- The header value must be a REAL key (e.g., `Bearer sk-abc123...`), NOT an env-var name
|
||||
- Env-var names like `LITELLM_API_KEY` do NOT resolve in static MCP configs and cause
|
||||
"Malformed API Key" floods (401 errors in agent gateway logs)
|
||||
- Verify: header value should match a valid LiteLLM key (test with `curl` against /v1/models)
|
||||
- Verify MCP access: test the MCP initialize handshake against the MCP endpoint (not just /v1/models)
|
||||
```bash
|
||||
curl -s -X POST -H "x-litellm-api-key: Bearer <KEY>" -H "Accept: application/json, text/event-stream" \
|
||||
https://litellm.sysloggh.net/mcp -d '{"jsonrpc":"2.0","id":1,"method":"initialize",...}' \
|
||||
| jq '.data.result.serverInfo' # should show serverInfo.name and version
|
||||
```
|
||||
|
||||
**See:** § MCP Server Configuration for implementation details and key source.
|
||||
|
||||
## Execution
|
||||
|
||||
|
||||
@@ -16,7 +16,26 @@ author: Abiba (pi agent)
|
||||
|
||||
## Rule (One Sentence)
|
||||
|
||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with authenticated path `http://192.168.68.116/litellm/v1/responses` — hardcoded keys AND unauthenticated `/v1` direct access are both forbidden.**
|
||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
|
||||
|
||||
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
|
||||
|
||||
## Model Access Tiers (2026-09-20)
|
||||
|
||||
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
|
||||
|
||||
| Tier | Models | Who gets it |
|
||||
|------|--------|-------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
|
||||
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
|
||||
|
||||
**Rules:**
|
||||
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
|
||||
2. Agent keys MUST carry an **explicit local-only** `models` list.
|
||||
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
|
||||
4. The master key always bypasses scoping — it is admin-only, never for inference.
|
||||
|
||||
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
|
||||
|
||||
## Scope
|
||||
|
||||
@@ -38,8 +57,8 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
|
||||
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
|
||||
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
|
||||
|
||||
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid).
|
||||
Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**.
|
||||
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` non-canonical but working).
|
||||
The public host `https://litellm.sysloggh.net` serves `/v1` ONLY (404 on `/litellm/v1`).
|
||||
|
||||
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
|
||||
|
||||
@@ -113,7 +132,7 @@ model:
|
||||
|
||||
model:
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/v1 # ← RULE VIOLATION: unauthenticated path
|
||||
base_url: http://192.168.68.116/v1 # ← NON-CANONICAL but WORKING (authenticated via nginx, WARN not FAIL)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
```
|
||||
|
||||
|
||||
@@ -759,4 +759,4 @@ which was kill+nohup outside systemd) are banned by policy.
|
||||
| Script | Why Disabled |
|
||||
|--------|-------------|
|
||||
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
|
||||
| `zulip-monitor.sh` (Abiba) | Replaced by agent-health-check.py + PM2 auto-restart |
|
||||
| `zulip-monitor.sh` (Abiba) | Standalone cron replaced by agent-health-check.py + PM2 auto-restart; the script itself remains active as the execution step of the `zulip-health` responsibility contract (see `zulip-health.prose.md`) |
|
||||
|
||||
@@ -138,11 +138,19 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
||||
tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
**EXECUTABLE OWNER:** The canonical probe set lives in `scripts/infra-monitoring.sh`.
|
||||
A check run is a single command: `bash scripts/infra-monitoring.sh` (from the
|
||||
repository root). Paste its raw output verbatim into the report. The script
|
||||
exits non-zero naming every failed target; there is no "OK" summary when any
|
||||
leg failed. Port drift is caught by `scripts/test_infra_monitoring.sh` which
|
||||
asserts every probed port matches the documented value.
|
||||
|
||||
**PROBE SHAPE (per standing rules above):**
|
||||
- Every probe prints the target name + URL + HTTP code (or failure kind)
|
||||
- Retry once on connection failure at longer timeout
|
||||
- Every probe prints: `✅ <name>: alive` on success, or `🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)` on failure
|
||||
- PVE API failures include `(any-HTTP liveness, -k for self-signed)` to distinguish TLS vs connection
|
||||
- Retry once on connection failure at longer timeout (25s connect, 30s max)
|
||||
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
||||
- Report the actual probe command and its result, not a summary verdict
|
||||
- Report the actual probe output, not a summary verdict
|
||||
|
||||
```bash
|
||||
# Provenance — run first; paste the absolute path into the report
|
||||
@@ -269,6 +277,27 @@ code (or failure kind with retry details). Apply the standing probe rules: any
|
||||
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
||||
a failure.
|
||||
|
||||
### Docker Stats and PVE Exporter Ports
|
||||
|
||||
These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:
|
||||
|
||||
| Exporter | Port | Container | Metrics |
|
||||
|----------|------|-----------|---------|
|
||||
| **Docker Stats** | **9324** | harness-docker-stats | `docker_container_*` (per-container CPU/mem/network) |
|
||||
| **PVE Exporter** | **9221** | harness-pve-exporter | `pve_*` (5 cluster-level metrics) |
|
||||
|
||||
**IMPORTANT**: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (`builder_builds_*`, `containerd_build_info_*`).
|
||||
|
||||
```bash
|
||||
# Docker Stats (harness-docker-stats)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
|
||||
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||
|
||||
# PVE Exporter (harness-pve-exporter)
|
||||
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
|
||||
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
||||
```
|
||||
|
||||
### Phase 1: GPU Exporters
|
||||
|
||||
**NVIDIA (.8 and .110)**:
|
||||
|
||||
@@ -208,19 +208,19 @@ mcp_servers:
|
||||
| Key | MCP Access |
|
||||
|-----|-----------|
|
||||
| Master key | ✅ Full — 90 tools (vault-injected) |
|
||||
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.99.1 |
|
||||
| Agent keys (mumuni, tanko, etc.) | ✅ Per-key grants supported (as of 2026-09-18 verification) |
|
||||
|
||||
### Known Limitations
|
||||
- Per-key MCP server grants not functional — only master key has access
|
||||
- ~~Per-key MCP server grants not functional — only master key has access~~ (resolved 2026-09-18: per-key grants now work)
|
||||
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
|
||||
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
|
||||
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
|
||||
|
||||
### Migration Path
|
||||
When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
||||
1. Grant agent keys `mcp_servers: ["ra_h_os"]`
|
||||
2. Update Hermes `mcp_servers.ra-h-os.url` from `http://192.168.68.65:3100/mcp` → `http://192.168.68.116:4000/mcp/`
|
||||
3. Add `headers: {x-litellm-api-key: "Bearer $LITELLM_API_KEY"}` to MCP config
|
||||
### Migration Path (COMPLETED 2026-09-18)
|
||||
Per-key MCP grants are now supported:
|
||||
1. ✅ Agent keys granted MCP access via `allowed_mcp_servers` field
|
||||
2. ✅ Hermes `mcp_servers.litellm.url` set to `https://litellm.sysloggh.net/mcp`
|
||||
3. ✅ `headers: {x-litellm-api-key: "Bearer <literal_key>"}` added to MCP config
|
||||
|
||||
## Security-Specific Updates
|
||||
|
||||
|
||||
@@ -71,6 +71,10 @@ description: >
|
||||
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
||||
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
||||
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
||||
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
|
||||
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
|
||||
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
|
||||
In LiteLLM Community both silently grant access to EVERY model, including cloud.
|
||||
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
||||
- Return the new key
|
||||
5. **If action == "rotate"**:
|
||||
@@ -87,6 +91,66 @@ description: >
|
||||
- Confirm key alias matches agent_name in LiteLLM key list
|
||||
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
||||
|
||||
## Cloud Provider Consolidation (2026-09-20)
|
||||
|
||||
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
|
||||
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
|
||||
|
||||
### Provider map (per-account namespacing)
|
||||
|
||||
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
|
||||
stay separate:
|
||||
|
||||
| Prefix | Upstream | Auth | Vault secret |
|
||||
|--------|----------|------|--------------|
|
||||
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
|
||||
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
|
||||
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
|
||||
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
|
||||
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
|
||||
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
|
||||
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
|
||||
|
||||
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
|
||||
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
|
||||
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
|
||||
|
||||
### Access tiers (MUST be enforced per key)
|
||||
|
||||
| Tier | Model names | Granted to |
|
||||
|------|-------------|------------|
|
||||
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
|
||||
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
|
||||
|
||||
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
|
||||
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
|
||||
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
|
||||
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
|
||||
|
||||
### Creating the cloud-enabled key
|
||||
|
||||
```bash
|
||||
# ALWAYS read the live roster first (key-scoped):
|
||||
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
|
||||
| jq -r '.data[].id'
|
||||
|
||||
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
|
||||
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
|
||||
```
|
||||
|
||||
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
|
||||
any `<prefix>/` cloud model.
|
||||
|
||||
### Adding a new cloud provider
|
||||
|
||||
1. Add the upstream key to Infisical `infrastructure/production/root`.
|
||||
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
|
||||
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
|
||||
3. Restart the `harness-litellm` container.
|
||||
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
|
||||
captain approval.
|
||||
5. Update this table and the access-tier section.
|
||||
|
||||
## Production Vault Access Process (canonical, 2026-07-17)
|
||||
|
||||
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
||||
|
||||
+20
-10
@@ -2,10 +2,10 @@
|
||||
kind: responsibility
|
||||
name: pm2-self-heal
|
||||
description: >
|
||||
PM2 process health check for abiba-telegram, abiba-zulip, gitea-runner, and zulip-watchdog.
|
||||
gpu-monitor is systemd-managed (gpu-monitor.service), NOT PM2.
|
||||
gpu-watchdog is decommissioned and folded into gpu-monitor.service.
|
||||
gitea-runner is KEPT. abiba-zulip is KEPT (online for days).
|
||||
Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
|
||||
and auto-restarts any that are stopped or errored. Logs every action to
|
||||
Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary).
|
||||
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -16,6 +16,8 @@ description: >
|
||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||
- last_check: timestamp
|
||||
|
||||
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed.
|
||||
|
||||
|
||||
## Continuity
|
||||
|
||||
@@ -40,7 +42,7 @@ description: >
|
||||
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
||||
spoton incident). Alerts include the restart count.
|
||||
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
|
||||
four processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog). The spoton
|
||||
four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton
|
||||
reference above is historical context for the crash-loop guard, not a live process.
|
||||
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
||||
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
||||
@@ -51,17 +53,25 @@ description: >
|
||||
## Execution
|
||||
|
||||
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
||||
2. **Check abiba-telegram**:
|
||||
2. **Check abiba-telegram** (safe to auto-restart):
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
|
||||
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
||||
- If status is "online" → pass, log restarts count
|
||||
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
||||
- If restarts > 5 in last hour → alert owner with full diagnostics
|
||||
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
6. **Wait 5 min** → repeat from step 1
|
||||
4. **Check gitea-runner**:
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
5. **Check zulip-watchdog**:
|
||||
- If status is "online" → pass
|
||||
- If status is "stopped" or "errored" → apply Rule 1
|
||||
- If restarts > 5 → alert owner
|
||||
6. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
||||
7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
|
||||
8. **Wait 5 min** → repeat from step 1
|
||||
|
||||
## Example Output (when healthy)
|
||||
|
||||
|
||||
@@ -89,6 +89,48 @@ agent: abiba
|
||||
| minipve | 192.168.68.12 | PVE |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
||||
|
||||
## PBS GC (Proxmox Backup Server)
|
||||
|
||||
### Schedule
|
||||
Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**.
|
||||
|
||||
**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.
|
||||
|
||||
### What Actually Runs
|
||||
Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`.
|
||||
|
||||
**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC.
|
||||
|
||||
### Datastore Location
|
||||
- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free)
|
||||
- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume)
|
||||
|
||||
### Liveness Check
|
||||
The new `proxmox-monitor.sh` leg checks storepve-datastore GC health:
|
||||
- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json`
|
||||
- **FAILS** if `last-run-endtime` is older than 48 hours
|
||||
- Reports age in hours and pending-bytes
|
||||
|
||||
**All six verdict shapes** (exactly as emitted by the script):
|
||||
|
||||
1. **Healthy** (fresh GC, 0 B pending):
|
||||
`✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)`
|
||||
|
||||
2. **Stale** (GC ran >48h ago):
|
||||
`🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)`
|
||||
|
||||
3. **Probe-failed: empty read** (000/timeout/unreadable):
|
||||
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)`
|
||||
|
||||
4. **Probe-failed: unparseable** (non-empty but invalid JSON — the "command not found" case):
|
||||
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)`
|
||||
|
||||
5. **Never-run: datastore absent** (valid JSON but storepve-datastore not in list):
|
||||
`🔴 PBS GC: never-run (storepve-datastore not found in GC list)`
|
||||
|
||||
6. **Never-run: no endtime** (valid JSON with datastore present but last-run-endtime is null/0):
|
||||
`🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)`
|
||||
|
||||
## Operations
|
||||
|
||||
### view-dashboards
|
||||
|
||||
@@ -22,10 +22,12 @@ AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TO
|
||||
|
||||
ZULIP_SITE = "https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL = "abiba-bot@chat.sysloggh.net"
|
||||
ZULIP_API_KEY = os.environ.get("ZULIP_API_KEY", "")
|
||||
if not ZULIP_API_KEY:
|
||||
raise SystemExit("ZULIP_API_KEY not set — refusing to run with no credential")
|
||||
ZULIP_AUTH = f"{ZULIP_EMAIL}:{ZULIP_API_KEY}"
|
||||
# Note: /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential).
|
||||
# No Zulip API key is required for this call. If a future leg genuinely needs abiba-bot's key,
|
||||
# it must prove it with a 200 from /api/v1/users/me as abiba-bot and label itself degraded when it cannot.
|
||||
# Never fall back to the vault's shared ZULIP_API_KEY.
|
||||
ZULIP_AUTH = None
|
||||
DEGRADED_LEGS = []
|
||||
|
||||
LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||
LITELLM_BACKEND = "192.168.68.116"
|
||||
@@ -158,7 +160,7 @@ def collect():
|
||||
# ── Storage ──
|
||||
storages = pve_get("/api2/json/nodes/storepve/storage")
|
||||
report["storage"] = []
|
||||
for s in storages:
|
||||
for s in (storages or []):
|
||||
total = s.get("total",0) or 1
|
||||
used = s.get("used",0)
|
||||
pct = used/total*100
|
||||
@@ -673,8 +675,9 @@ def send_email(html_content, subject_prefix=""):
|
||||
try:
|
||||
EMAIL_PASSWORD = os.environ.get("EMAIL_PASSWORD") or os.environ.get("SMTP_PASSWORD") or os.environ.get("MAIL_PASSWORD")
|
||||
if not EMAIL_PASSWORD:
|
||||
print("EMAIL_PASSWORD not set — refusing to send email", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
print(" ⚠️ Degraded leg: credential-missing: EMAIL_PASSWORD (or SMTP_PASSWORD/MAIL_PASSWORD)", file=sys.stderr)
|
||||
DEGRADED_LEGS.append("credential-missing: EMAIL_PASSWORD")
|
||||
return True, "✅ Email leg degraded (no credential) — report still produced"
|
||||
GMAIL_EMAIL = "jtabiri@gmail.com"
|
||||
|
||||
server = smtplib.SMTP("smtp.gmail.com", 587)
|
||||
@@ -713,6 +716,17 @@ if __name__ == "__main__":
|
||||
print(f" {msg}")
|
||||
|
||||
# Show summary
|
||||
if DEGRADED_LEGS:
|
||||
print(f"\n⚠️ Degraded legs ({len(DEGRADED_LEGS)}):")
|
||||
for leg in DEGRADED_LEGS:
|
||||
print(f" - {leg}")
|
||||
else:
|
||||
print("\n✅ All legs fully credentialed")
|
||||
|
||||
# A failed send must exit non-zero; a degraded leg (no credential) must stay exit 0
|
||||
if not ok:
|
||||
sys.exit(1)
|
||||
|
||||
issues = sum(1 for i in ["red"] if report.get("zulip_ext", {}).get("connected") == False)
|
||||
print(f"\n📋 Summary:")
|
||||
print(f" Proxmox: {report['nodes_online']}/{report['node_count']} nodes online")
|
||||
|
||||
Executable
+234
@@ -0,0 +1,234 @@
|
||||
#!/bin/bash
|
||||
# infrastructure-monitoring.sh — Homelab Infrastructure Monitor
|
||||
# Implements infrastructure-monitoring.prose.md (check-health section)
|
||||
#
|
||||
# Legs: Grafana, Prometheus, LiteLLM, PVE API (5 nodes), GPU exporters,
|
||||
# Docker Stats, PVE Exporter
|
||||
#
|
||||
# Design:
|
||||
# - Every target, port, path, and expected status is defined in code
|
||||
# - Liveness rule: any HTTP status = ALIVE for auth-gated/redirect endpoints;
|
||||
# only connection failures (000/timeout) = probe-failed
|
||||
# - Bare-200 rule: expected status must match exactly (200); anything else = alert
|
||||
# - PVE API uses -k flag (self-signed certs), probes /api2/json/version
|
||||
# - Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116, probed via SSH
|
||||
# with one retry at longer timeout (25s connect, 30s max) to distinguish
|
||||
# transient timeout from host-down
|
||||
# - Non-zero exit naming every failed target; no "OK" summary when any leg failed
|
||||
#
|
||||
# Output shape per leg:
|
||||
# ✅ <name>: alive
|
||||
# 🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)
|
||||
#
|
||||
# Failure kinds: timeout | refused | tls | unexpected:<code> (printed in the failure line)
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
# ── Configuration (documented in infrastructure-monitoring.prose.md) ────────
|
||||
# Change these in ONE place; test_infra_monitoring.sh asserts against these.
|
||||
|
||||
GRAFANA_HOST="192.168.68.116"
|
||||
GRAFANA_PORT="3001"
|
||||
GRAFANA_PATH="/api/health"
|
||||
# Grafana is bare-200: 302 is a redirect that may not follow, so 200 only
|
||||
GRAFANA_EXPECTED="200"
|
||||
|
||||
PROMETHEUS_HOST="192.168.68.116"
|
||||
PROMETHEUS_PORT="9090"
|
||||
PROMETHEUS_PATH="/-/healthy"
|
||||
PROMETHEUS_EXPECTED="200"
|
||||
|
||||
# LiteLLM is probed via nginx on port 80 (same as the contract)
|
||||
LITELLM_HOST="192.168.68.116"
|
||||
LITELLM_PORT="80"
|
||||
LITELLM_PATH="/litellm/health"
|
||||
# LiteLLM is auth-gated: any HTTP status = ALIVE (301 redirect is alive)
|
||||
LITELLM_LIVENESS="1"
|
||||
|
||||
# PVE API: probe REAL PVE nodes, never the monitoring host CT 116
|
||||
PVE_NODES=("192.168.68.9" "192.168.68.12" "192.168.68.6" "192.168.68.15" "192.168.68.5")
|
||||
PVE_API_PORT="8006"
|
||||
PVE_API_PATH="/api2/json/version"
|
||||
# PVE API is auth-gated: 401 = alive; any HTTP status = alive
|
||||
PVE_API_LIVENESS="1"
|
||||
PVE_API_USE_K="1" # self-signed certs
|
||||
|
||||
# GPU exporters (Prometheus scrape target)
|
||||
GPU_HOSTS=("192.168.68.8" "192.168.68.110" "192.168.68.15")
|
||||
GPU_PORT="9400"
|
||||
GPU_PATH="/metrics"
|
||||
GPU_EXPECTED="200"
|
||||
|
||||
# Docker Stats and PVE Exporter bind to 127.0.0.1 on CT 116
|
||||
DOCKER_STATS_PORT="9324" # harness-docker-stats (docker_container_* metrics)
|
||||
PVE_EXPORTER_PORT="9221" # harness-pve-exporter (5 pve_* metrics)
|
||||
CT116_SSH_HOST="192.168.68.116"
|
||||
# Both are bare-200: 404 = container not yet started
|
||||
DOCKER_STATS_EXPECTED="200|404"
|
||||
PVE_EXPORTER_EXPECTED="200|404"
|
||||
|
||||
# ── Probe Functions ─────────────────────────────────────────────────────────
|
||||
|
||||
# probe_http <host> <port> <path> <expected_pattern> [use_k] [ssh_host] [scheme] [liveness]
|
||||
# Returns 0 if probe succeeds (matches expected or liveness), 1 if probe-failed.
|
||||
# Prints the result line.
|
||||
#
|
||||
# FIX C1: The kind value is computed and printed in the failure line.
|
||||
# FIX C2: SSH retry logic is in the first attempt branch (not unreachable).
|
||||
|
||||
LAST_KIND=""
|
||||
probe_http() {
|
||||
local host="$1" port="$2" path="$3" expected="$4"
|
||||
local use_k="${5:-}" ssh_host="${6:-}" scheme="${7:-http}" liveness="${8:-0}"
|
||||
local url="${scheme}://${host}:${port}${path}"
|
||||
local code="" kind=""
|
||||
LAST_KIND=""
|
||||
|
||||
# Single invocation that captures both output and status
|
||||
if [ -n "$ssh_host" ]; then
|
||||
out=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||
rc=$?
|
||||
else
|
||||
out=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 --max-time 15 ${use_k:+-k} "$url" 2>/dev/null)
|
||||
rc=$?
|
||||
fi
|
||||
code=$(printf '%s' "$out" | tr -d '[:space:]')
|
||||
|
||||
# Classify failure kind and retry if needed
|
||||
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||
# Distinguish timeout from TLS error from refused
|
||||
case "$rc" in
|
||||
35|51|58|59|60|77|83) kind="tls" ;;
|
||||
*) kind="timeout" ;;
|
||||
esac
|
||||
# Retry once at longer timeout (25s connect, 30s max)
|
||||
if [ -n "$ssh_host" ]; then
|
||||
code=$(ssh -o ConnectTimeout=5 -o BatchMode=yes "root@${ssh_host}" \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} ${url}" 2>/dev/null)
|
||||
rc=$?
|
||||
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||
else
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 --max-time 30 ${use_k:+-k} "$url" 2>/dev/null)
|
||||
rc=$?
|
||||
code=$(printf '%s' "$code" | tr -d '[:space:]')
|
||||
# On retry, classify: still 000 = keep existing kind (or timeout if empty), unexpected status = refused
|
||||
if [ -z "$code" ] || [ "$code" = "000" ]; then
|
||||
[ -z "$kind" ] && kind="timeout"
|
||||
elif ! echo "$code" | grep -qE "^(${expected})$"; then
|
||||
kind="refused"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# Check result
|
||||
if [ -n "$code" ] && [ "$code" != "000" ]; then
|
||||
if [ "$liveness" = "1" ]; then
|
||||
# Any HTTP status = ALIVE for auth-gated/redirect endpoints
|
||||
return 0
|
||||
else
|
||||
# Bare-200 or specific expected pattern
|
||||
if echo "$code" | grep -qE "^(${expected})$"; then
|
||||
return 0
|
||||
else
|
||||
kind="unexpected:$code"
|
||||
LAST_KIND="$kind"
|
||||
return 1
|
||||
fi
|
||||
fi
|
||||
else
|
||||
[ -z "$kind" ] && kind="refused"
|
||||
LAST_KIND="$kind"
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
|
||||
# ── Main ────────────────────────────────────────────────────────────────────
|
||||
|
||||
FAILED=()
|
||||
FAILED_KIND=()
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
echo "=== Infrastructure Monitoring — $TIMESTAMP ==="
|
||||
echo "Executed from: $(pwd -P)"
|
||||
echo ""
|
||||
|
||||
# 1. Grafana (CT 116 :3001 /api/health) — bare-200
|
||||
if probe_http "$GRAFANA_HOST" "$GRAFANA_PORT" "$GRAFANA_PATH" "$GRAFANA_EXPECTED"; then
|
||||
echo " ✅ Grafana: alive"
|
||||
else
|
||||
echo " 🔴 Grafana: probe-failed: ${GRAFANA_HOST}:${GRAFANA_PORT} (expected ${GRAFANA_EXPECTED}) (<${LAST_KIND}>)"
|
||||
FAILED+=("grafana")
|
||||
fi
|
||||
|
||||
# 2. Prometheus (CT 116 :9090 /-/healthy) — bare-200
|
||||
if probe_http "$PROMETHEUS_HOST" "$PROMETHEUS_PORT" "$PROMETHEUS_PATH" "$PROMETHEUS_EXPECTED"; then
|
||||
echo " ✅ Prometheus: alive"
|
||||
else
|
||||
echo " 🔴 Prometheus: probe-failed: ${PROMETHEUS_HOST}:${PROMETHEUS_PORT} (expected ${PROMETHEUS_EXPECTED}) (<${LAST_KIND}>)"
|
||||
FAILED+=("prometheus")
|
||||
fi
|
||||
|
||||
# 3. LiteLLM (CT 116 :80/litellm/health via nginx) — liveness (any HTTP = alive)
|
||||
if probe_http "$LITELLM_HOST" "$LITELLM_PORT" "$LITELLM_PATH" "" "" "" "http" "$LITELLM_LIVENESS"; then
|
||||
echo " ✅ LiteLLM: alive"
|
||||
else
|
||||
echo " 🔴 LiteLLM: probe-failed: ${LITELLM_HOST}:${LITELLM_PORT}${LITELLM_PATH} (any-HTTP liveness) (<${LAST_KIND}>)"
|
||||
FAILED+=("litellm")
|
||||
fi
|
||||
|
||||
# 4. PVE API (5 real nodes :8006 /api2/json/version, -k, liveness)
|
||||
PVE_FAILED=()
|
||||
for node in "${PVE_NODES[@]}"; do
|
||||
if probe_http "$node" "$PVE_API_PORT" "$PVE_API_PATH" "" "$PVE_API_USE_K" "" "https" "$PVE_API_LIVENESS"; then
|
||||
echo " ✅ PVE API ${node}: alive"
|
||||
else
|
||||
echo " 🔴 PVE API ${node}: probe-failed: ${node}:${PVE_API_PORT} (any-HTTP liveness, -k for self-signed) (<${LAST_KIND}>)"
|
||||
PVE_FAILED+=("$node")
|
||||
fi
|
||||
done
|
||||
if [ ${#PVE_FAILED[@]} -gt 0 ]; then
|
||||
FAILED+=("pve-api: ${PVE_FAILED[*]}")
|
||||
fi
|
||||
|
||||
# 5. GPU exporters (:9400/metrics) — bare-200
|
||||
GPU_FAILED=()
|
||||
for host in "${GPU_HOSTS[@]}"; do
|
||||
if probe_http "$host" "$GPU_PORT" "$GPU_PATH" "$GPU_EXPECTED"; then
|
||||
echo " ✅ GPU exporter ${host}: alive"
|
||||
else
|
||||
echo " 🔴 GPU exporter ${host}: probe-failed: ${host}:${GPU_PORT} (expected 200) (<${LAST_KIND}>)"
|
||||
GPU_FAILED+=("$host")
|
||||
fi
|
||||
done
|
||||
if [ ${#GPU_FAILED[@]} -gt 0 ]; then
|
||||
FAILED+=("gpu-exporters: ${GPU_FAILED[*]}")
|
||||
fi
|
||||
|
||||
# 6. Docker Stats (CT 116 :9324, 127.0.0.1 via SSH) — 200|404
|
||||
if probe_http "127.0.0.1" "$DOCKER_STATS_PORT" "/" "$DOCKER_STATS_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||
echo " ✅ Docker Stats: alive"
|
||||
else
|
||||
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:${DOCKER_STATS_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||
FAILED+=("docker-stats")
|
||||
fi
|
||||
|
||||
# 7. PVE Exporter (CT 116 :9221, 127.0.0.1 via SSH) — 200|404
|
||||
if probe_http "127.0.0.1" "$PVE_EXPORTER_PORT" "/" "$PVE_EXPORTER_EXPECTED" "" "$CT116_SSH_HOST"; then
|
||||
echo " ✅ PVE Exporter: alive"
|
||||
else
|
||||
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:${PVE_EXPORTER_PORT} (expected 200|404) (<${LAST_KIND}>)"
|
||||
FAILED+=("pve-exporter")
|
||||
fi
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
|
||||
echo ""
|
||||
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||
echo " ✅ All legs OK"
|
||||
exit 0
|
||||
else
|
||||
for f in "${FAILED[@]}"; do
|
||||
echo " 🔴 FAILED: $f"
|
||||
done
|
||||
exit 1
|
||||
fi
|
||||
@@ -55,9 +55,12 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
||||
try:
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
if rc != 0:
|
||||
# Determine failure kind from curl exit code
|
||||
# Check if this is a timeout from run_command (rc=1, stderr="TIMEOUT")
|
||||
if rc == 1 and stderr == "TIMEOUT":
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
# Otherwise, determine failure kind from curl exit code
|
||||
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
|
||||
if rc == 28:
|
||||
elif rc == 28:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
elif rc == 7:
|
||||
return (000, "connection refused")
|
||||
@@ -73,6 +76,20 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
|
||||
except subprocess.TimeoutExpired:
|
||||
return (000, "timeout after " + str(timeout) + "s")
|
||||
|
||||
def check_host_health(host_ip):
|
||||
"""Check if the GPU host's llama-chat-api health endpoint is reachable
|
||||
|
||||
Returns: (healthy: bool, detail: str)
|
||||
"""
|
||||
code, _ = probe_http("http://" + host_ip + ":8080/health", timeout=10)
|
||||
if code == 200:
|
||||
return True, "host healthy (200)"
|
||||
elif code == 000:
|
||||
return False, "host unreachable (timeout or refused)"
|
||||
else:
|
||||
return False, "host unhealthy (HTTP " + str(code) + ")"
|
||||
|
||||
|
||||
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
|
||||
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
|
||||
cmd = "curl -s -m " + str(timeout)
|
||||
@@ -115,29 +132,53 @@ def check_model_probes():
|
||||
|
||||
results = []
|
||||
|
||||
# Host health mapping: model -> host IP
|
||||
model_hosts = {
|
||||
"gpu-dense": "192.168.68.8", # RTX 3090
|
||||
"gpu-vision": "192.168.68.110", # RTX 5070
|
||||
"strix-moe": "192.168.68.15" # Strix Halo
|
||||
}
|
||||
|
||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||
# Single-host aliases: 30s timeout each
|
||||
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
|
||||
host_ip = model_hosts[model]
|
||||
# Single-host aliases: 30s initial timeout, retry once at 90s on failure
|
||||
# Worst-case prefill ~76s, so 90s retry ensures we cover it
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
|
||||
first_kind = None
|
||||
if code == 000 and failure_kind:
|
||||
# Report probe failure with kind, do not assert a service verdict
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
|
||||
first_kind = failure_kind
|
||||
time.sleep(1)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=90)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Both attempts failed - check host health to distinguish busy from down
|
||||
host_healthy, host_detail = check_host_health(host_ip)
|
||||
if host_healthy:
|
||||
results.append((model, False, "busy (completion timed out after retry; " + host_detail + ")"))
|
||||
else:
|
||||
# Host unreachable - report both kinds
|
||||
if first_kind:
|
||||
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
|
||||
else:
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
|
||||
elif code == 200:
|
||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
elif code in (401, 403):
|
||||
# Credential fault - capture body and key alias
|
||||
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
|
||||
timeout=10)
|
||||
# Resolve key alias
|
||||
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
|
||||
alias = "monitor-20260813"
|
||||
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
|
||||
else:
|
||||
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
@@ -247,6 +288,7 @@ def main():
|
||||
print("")
|
||||
|
||||
all_pass = True
|
||||
degraded = [] # Track degraded (busy) checks
|
||||
|
||||
# Run all checks
|
||||
checks = [
|
||||
@@ -266,9 +308,15 @@ def main():
|
||||
# Model probes
|
||||
model_results = check_model_probes()
|
||||
for name, passed, detail in model_results:
|
||||
status = "✅" if passed else "❌"
|
||||
# Check if this is a busy (degraded) verdict
|
||||
if not passed and detail.startswith("busy "):
|
||||
status = "⚠️"
|
||||
degraded.append(name)
|
||||
else:
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
if not passed:
|
||||
# Only set all_pass=False for real failures (not busy)
|
||||
if not passed and not detail.startswith("busy "):
|
||||
all_pass = False
|
||||
|
||||
# Admin key list
|
||||
@@ -294,10 +342,16 @@ def main():
|
||||
|
||||
print("")
|
||||
if all_pass:
|
||||
print("✅ All checks passed")
|
||||
if degraded:
|
||||
print("✅ All checks passed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||
else:
|
||||
print("✅ All checks passed")
|
||||
return 0
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
if degraded:
|
||||
print("❌ Some checks failed (" + str(len(degraded)) + " degraded: " + ", ".join(degraded) + ")")
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
return 1
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
#!/bin/bash
|
||||
# pm2-self-heal — hourly PM2 process check
|
||||
# Part of the pm2-self-heal prose contract
|
||||
# Alerts via Telegram (abiba-zulip decommissioned 2026-07-04)
|
||||
# Alerts via Telegram (primary) and Zulip DM (secondary, abiba-zulip restored 2026-08-03)
|
||||
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
|
||||
|
||||
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
|
||||
@@ -46,9 +46,75 @@ if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
|
||||
fi
|
||||
fi
|
||||
|
||||
# Check abiba-zulip (live Zulip bridge, heartbeating)
|
||||
ZULIP_LINE=$(echo "$STATUS" | grep "abiba-zulip")
|
||||
ZULIP_STATUS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
ZULIP_RESTARTS=$(echo "$ZULIP_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$ZULIP_STATUS" != "online" ]; then
|
||||
pm2 restart abiba-zulip > /dev/null 2>&1
|
||||
sleep 3
|
||||
ZULIP_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-zulip")
|
||||
ZULIP_STATUS2=$(echo "$ZULIP_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$ZULIP_STATUS2" = "online" ]; then
|
||||
msg="⚠️ abiba-zulip was **$ZULIP_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 abiba-zulip **failed restart** (was $ZULIP_STATUS, still $ZULIP_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$ZULIP_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ abiba-zulip has **$ZULIP_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Check gitea-runner
|
||||
GITEA_LINE=$(echo "$STATUS" | grep "gitea-runner")
|
||||
GITEA_STATUS=$(echo "$GITEA_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
GITEA_RESTARTS=$(echo "$GITEA_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$GITEA_STATUS" != "online" ]; then
|
||||
pm2 restart gitea-runner > /dev/null 2>&1
|
||||
sleep 3
|
||||
GITEA_LINE2=$(pm2 status --no-color 2>/dev/null | grep "gitea-runner")
|
||||
GITEA_STATUS2=$(echo "$GITEA_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$GITEA_STATUS2" = "online" ]; then
|
||||
msg="⚠️ gitea-runner was **$GITEA_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 gitea-runner **failed restart** (was $GITEA_STATUS, still $GITEA_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$GITEA_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ gitea-runner has **$GITEA_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Check zulip-watchdog
|
||||
WATCHDOG_LINE=$(echo "$STATUS" | grep "zulip-watchdog")
|
||||
WATCHDOG_STATUS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $10}' | xargs)
|
||||
WATCHDOG_RESTARTS=$(echo "$WATCHDOG_LINE" | awk -F'│' '{print $9}' | xargs)
|
||||
|
||||
if [ "$WATCHDOG_STATUS" != "online" ]; then
|
||||
pm2 restart zulip-watchdog > /dev/null 2>&1
|
||||
sleep 3
|
||||
WATCHDOG_LINE2=$(pm2 status --no-color 2>/dev/null | grep "zulip-watchdog")
|
||||
WATCHDOG_STATUS2=$(echo "$WATCHDOG_LINE2" | awk -F'│' '{print $10}' | xargs)
|
||||
if [ "$WATCHDOG_STATUS2" = "online" ]; then
|
||||
msg="⚠️ zulip-watchdog was **$WATCHDOG_STATUS** → restarted to online"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
else
|
||||
msg="🚨 zulip-watchdog **failed restart** (was $WATCHDOG_STATUS, still $WATCHDOG_STATUS2)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
elif [ "$WATCHDOG_RESTARTS" -gt 5 ]; then
|
||||
msg="⚠️ zulip-watchdog has **$WATCHDOG_RESTARTS** restarts (high count)"
|
||||
ALERTS="${ALERTS}${msg}\n"
|
||||
fi
|
||||
|
||||
# Log check
|
||||
{
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS alerts=${ALERTS:+yes}"
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] tel=$TEL_STATUS zulip=$ZULIP_STATUS gitea=$GITEA_STATUS watchdog=$WATCHDOG_STATUS alerts=${ALERTS:+yes}"
|
||||
[ -n "$ALERTS" ] && echo "$ALERTS"
|
||||
} >> "$LOG"
|
||||
|
||||
|
||||
+16
-1
@@ -135,7 +135,22 @@ fi
|
||||
|
||||
echo " Cross-contract: $WARNINGS total warnings across all checks"
|
||||
|
||||
# ── 4. Summary ──
|
||||
# ── 4. Committed-credential scan ──
|
||||
# The 2026-09-17 purge removed six live credentials that had sat in .md prose
|
||||
# and scripts for weeks. This step makes that class of commit FAIL the gate
|
||||
# instead of printing a warning. Patterns: scripts/secret-patterns.tsv.
|
||||
# Only deliberate synthetic examples may be listed in scripts/secret-allowlist.tsv,
|
||||
# each with a reason. Run `bash scripts/secret-scan.sh --staged` before committing.
|
||||
echo ""
|
||||
echo "── 4. Secret scan (committed credentials) ──"
|
||||
if bash scripts/secret-scan.sh; then
|
||||
echo " ✅ No committed credentials"
|
||||
else
|
||||
echo " ❌ COMMITTED CREDENTIAL DETECTED"
|
||||
FAILED=1
|
||||
fi
|
||||
|
||||
# ── 5. Summary ──
|
||||
echo ""
|
||||
echo "═══════════════════════════════════"
|
||||
if [ $FAILED -eq 1 ]; then
|
||||
|
||||
Executable
+138
@@ -0,0 +1,138 @@
|
||||
#!/bin/bash
|
||||
# proxmox-monitor.sh — Proxmox Cluster + Docker Monitoring Health Check
|
||||
# Implements proxmox-monitor.prose.md (check-health section)
|
||||
#
|
||||
# Legs: Prometheus, Grafana, Docker Stats Exporter, PVE Exporter, PBS GC
|
||||
# All legs must return 200 for healthy status.
|
||||
#
|
||||
# Run: bash scripts/proxmox-monitor.sh
|
||||
# Exits 0 if all probes pass, 1 if any fails.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
CT116_HOST="192.168.68.116"
|
||||
FAILED=()
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
|
||||
echo "=== Proxmox Monitor — $TIMESTAMP ==="
|
||||
echo "Executed from: $(pwd -P)"
|
||||
echo ""
|
||||
|
||||
# 1. Prometheus health (bound to 0.0.0.0:9090 on .116)
|
||||
PROM_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:9090/-/healthy 2>/dev/null)
|
||||
PROM_CODE=$(printf '%s' "$PROM_CODE" | tr -d '[:space:]')
|
||||
[ -n "$PROM_CODE" ] || PROM_CODE="000"
|
||||
|
||||
if [ "$PROM_CODE" = "200" ]; then
|
||||
echo " ✅ Prometheus: alive"
|
||||
else
|
||||
echo " 🔴 Prometheus: probe-failed: ${CT116_HOST}:9090 (expected 200, got ${PROM_CODE})"
|
||||
FAILED+=("prometheus")
|
||||
fi
|
||||
|
||||
# 2. Grafana health (bound to 0.0.0.0:3001 on .116)
|
||||
GRAF_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:3001/api/health 2>/dev/null)
|
||||
GRAF_CODE=$(printf '%s' "$GRAF_CODE" | tr -d '[:space:]')
|
||||
[ -n "$GRAF_CODE" ] || GRAF_CODE="000"
|
||||
|
||||
if [ "$GRAF_CODE" = "200" ]; then
|
||||
echo " ✅ Grafana: alive"
|
||||
else
|
||||
echo " 🔴 Grafana: probe-failed: ${CT116_HOST}:3001 (expected 200, got ${GRAF_CODE})"
|
||||
FAILED+=("grafana")
|
||||
fi
|
||||
|
||||
# 3. Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — probe from .116 localhost)
|
||||
DOCKER_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics" 2>/dev/null)
|
||||
DOCKER_CODE=$(printf '%s' "$DOCKER_CODE" | tr -d '[:space:]')
|
||||
[ -n "$DOCKER_CODE" ] || DOCKER_CODE="000"
|
||||
|
||||
if [ "$DOCKER_CODE" = "200" ]; then
|
||||
echo " ✅ Docker Stats: alive"
|
||||
else
|
||||
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:9324 (expected 200, got ${DOCKER_CODE})"
|
||||
FAILED+=("docker-stats")
|
||||
fi
|
||||
|
||||
# 4. PVE Exporter (bound to 127.0.0.1:9221 on .116 — probe from .116 localhost)
|
||||
PVE_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
||||
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics" 2>/dev/null)
|
||||
PVE_CODE=$(printf '%s' "$PVE_CODE" | tr -d '[:space:]')
|
||||
[ -n "$PVE_CODE" ] || PVE_CODE="000"
|
||||
|
||||
if [ "$PVE_CODE" = "200" ]; then
|
||||
echo " ✅ PVE Exporter: alive"
|
||||
else
|
||||
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:9221 (expected 200, got ${PVE_CODE})"
|
||||
FAILED+=("pve-exporter")
|
||||
fi
|
||||
|
||||
# 5. PBS GC liveness (storepve-datastore GC must have run within 48h)
|
||||
PBS_GC_OUTPUT=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192.168.68.6 \
|
||||
"pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null)
|
||||
PBS_GC_OUTPUT=$(printf '%s' "$PBS_GC_OUTPUT" | tr -d '[:space:]')
|
||||
[ -n "$PBS_GC_OUTPUT" ] || PBS_GC_OUTPUT="000"
|
||||
|
||||
if [ "$PBS_GC_OUTPUT" = "000" ]; then
|
||||
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
# Parse the JSON to get storepve-datastore's last-run-endtime and pending-bytes
|
||||
PBS_GC_RESULT=$(echo "$PBS_GC_OUTPUT" | python3 -c "
|
||||
import sys, json
|
||||
try:
|
||||
data = json.load(sys.stdin)
|
||||
for store in data:
|
||||
if store['store'] == 'storepve-datastore':
|
||||
endtime = store.get('last-run-endtime')
|
||||
pending = store.get('pending-bytes', 0)
|
||||
if endtime is None or endtime == 0:
|
||||
print('never-run')
|
||||
else:
|
||||
print(f'{endtime}|{pending}')
|
||||
break
|
||||
else:
|
||||
print('absent')
|
||||
except json.JSONDecodeError:
|
||||
print('unparseable')
|
||||
" 2>/dev/null)
|
||||
|
||||
if [ -z "$PBS_GC_RESULT" ] || [ "$PBS_GC_RESULT" = "unparseable" ]; then
|
||||
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)"
|
||||
FAILED+=("pbs-gc")
|
||||
elif [ "$PBS_GC_RESULT" = "absent" ]; then
|
||||
echo " 🔴 PBS GC: never-run (storepve-datastore not found in GC list)"
|
||||
FAILED+=("pbs-gc")
|
||||
elif [ "$PBS_GC_RESULT" = "never-run" ]; then
|
||||
echo " 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
# Parse the endtime|pending format
|
||||
LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1)
|
||||
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
||||
|
||||
# Convert epoch to age in hours
|
||||
NOW_EPOCH=$(date -u +%s)
|
||||
AGE_HOURS=$(( (NOW_EPOCH - LAST_RUN_ENDTIME) / 3600 ))
|
||||
|
||||
if [ $AGE_HOURS -gt 48 ]; then
|
||||
echo " 🔴 PBS GC: stale (last run ${AGE_HOURS}h ago, pending-bytes: ${PENDING_BYTES} B)"
|
||||
FAILED+=("pbs-gc")
|
||||
else
|
||||
echo " ✅ PBS GC: healthy (last run ${AGE_HOURS}h ago, pending-bytes: ${PENDING_BYTES} B)"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
if [ ${#FAILED[@]} -eq 0 ]; then
|
||||
echo " ✅ All legs OK"
|
||||
exit 0
|
||||
else
|
||||
for f in "${FAILED[@]}"; do
|
||||
echo " 🔴 FAILED: $f"
|
||||
done
|
||||
exit 1
|
||||
fi
|
||||
@@ -0,0 +1,55 @@
|
||||
# secret-allowlist.tsv — exceptions for scripts/secret-scan.sh, every entry with a reason.
|
||||
#
|
||||
# Format: <rule-id|*><TAB><path-glob><TAB><literal-substring><TAB><reason>
|
||||
# Blank lines and lines whose first field starts with '#' are ignored.
|
||||
# A finding is suppressed only when ALL THREE of rule, path and literal match:
|
||||
# * the rule id equals the finding's rule id, or is '*'
|
||||
# * the finding's repo-relative path matches <path-glob> (bash glob)
|
||||
# * the finding's line contains <literal-substring> verbatim
|
||||
# An entry whose reason is empty is a hard error (exit 2) — no silent exceptions.
|
||||
#
|
||||
# RULE: never allowlist a live credential, and never broaden an entry (rule '*',
|
||||
# a wide path glob, or a short generic literal) just to silence a finding.
|
||||
# If the finding is real, remove the credential from the file.
|
||||
#
|
||||
# Entries are one per deliberate synthetic example, so the file reads as an
|
||||
# audit trail of reviewed exceptions rather than a list of things to ignore.
|
||||
# Rule '*' is used only where the same literal is matched by more than one rule.
|
||||
#
|
||||
# ── The 2026-09-17 purge placeholders ─────────────────────────────────────
|
||||
# PR #112 replaced six live credentials with `«vault: <project>/<env> <SECRET>»`
|
||||
# references. Those references are safe by construction (they name where the
|
||||
# secret is read from), but they are listed here explicitly rather than being
|
||||
# filtered by a general "vault" rule, so a new occurrence still needs a
|
||||
# deliberate, reasoned entry.
|
||||
secret-assign litellm-api-keys.prose.md MUMUNI_LITELLM_API_KEY=«vault: agents/production LITELLM_API_KEY» 2026-09-17 purge: replaced the live Mumuni LiteLLM key with its vault reference; no literal credential.
|
||||
secret-assign litellm-api-keys.prose.md MUMUNI_ZULIP_API_KEY=«vault: agents/production ZULIP_API_KEY» 2026-09-17 purge: replaced the live Mumuni Zulip key with its vault reference; no literal credential.
|
||||
* infrastructure-control.prose.md PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
|
||||
cred-prose infrastructure-control.prose.md Admin credentials: 2026-09-17 purge: the Stirling admin user/password are two `«vault: ...»` references; no literal credential.
|
||||
* scripts/daily-infra-report.py PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN» 2026-09-17 purge: Proxmox API token is read from the vault; the line only names the vault path.
|
||||
secret-assign stirling-pdf-agent-access.prose.md «vault: infrastructure/production STIRLING_API_KEY» 2026-09-17 purge: Stirling PDF API key is read from the vault; the curl example only names the vault path.
|
||||
bearer-token agent-zero-fix-summary.md «vault: agents/production OPENROUTER_API_KEY» 2026-09-17 purge: OpenRouter key is read from the vault; the example curl only names the vault path.
|
||||
# ── Deliberate synthetic examples in contracts (not from the purge) ───────
|
||||
# These exist to teach the rule they illustrate. They are listed here so the
|
||||
# guard is never taught to skip the words "synthetic"/"example" — a fabricated
|
||||
# example is always an explicit exception, never a pattern-level exemption.
|
||||
* hermes-key-enforcement.prose.md sk-synthetic-external-example Rule 15 illustration of a hardcoded external key that is tolerated; fabricated, never a live key.
|
||||
* hermes-key-enforcement.prose.md sk-synthetic-example-12345 Rule 15 illustration of a forbidden hardcoded key; fabricated, never a live key.
|
||||
openai-key hermes-key-enforcement.prose.md sk-synthetic-litellm- Fabricated key name inside a `grep 'LITELLM_API_KEY=...'` example; not a live key.
|
||||
secret-assign hermes-key-enforcement.prose.md sk-NEW_KEY Placeholder standing for the rotated key in an `infisical secrets set` command; not a literal key.
|
||||
openrouter-key agent-zero-openrouter-key.prose.md sk-or-v1-synthetic Synthetic key prefix in the contract's example response; the real key is read from the vault.
|
||||
openai-key litellm-api-keys.prose.md sk-synthetic-tanko-example Fabricated key name in migration history prose; not a live key.
|
||||
openai-key litellm-self-heal.prose.md sk-syslog-local-master-key Deprecated local LiteLLM master key name documented as no-live-usage; kept for history, not a usable credential.
|
||||
# ── Redacted evidence, not a credential ──────────────────────────────────
|
||||
secret-assign docs/probe-drift-round2-evidence.md =sk-... Probe evidence records redacted key trailers (`sk-...x6uw`); the usable part of the key is not present.
|
||||
# ── tests/test_secret_scan.sh fixtures ───────────────────────────────────
|
||||
# The self-test plants these fabricated values into a TEMP tree, whose path no
|
||||
# entry here covers, so each still fails the guard when planted (see the test's
|
||||
# "... fails the guard" cases). They are listed only so the repo-wide scan of
|
||||
# the test file itself stays quiet.
|
||||
* tests/test_secret_scan.sh sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef Self-test fixture: fabricated OpenRouter-shaped key written to a temp tree; the guard must fail on it there.
|
||||
bearer-token tests/test_secret_scan.sh Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb Self-test fixture: fabricated Bearer token written to a temp tree; the guard must fail on it there.
|
||||
proxmox-token tests/test_secret_scan.sh PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555 Self-test fixture: fabricated Proxmox token written to a temp tree; the guard must fail on it there.
|
||||
private-key tests/test_secret_scan.sh -----BEGIN OPENSSH PRIVATE KEY----- Self-test fixture: fabricated PEM banner written to a temp tree; the guard must fail on it there.
|
||||
cred-prose tests/test_secret_scan.sh Admin credentials: Self-test fixture: fabricated prose credential line written to a temp tree; the guard must fail on it there.
|
||||
secret-assign tests/test_secret_scan.sh DB_PASSWORD=correct-horse-battery-staple Self-test fixture: fabricated password assignment written to a temp tree; the guard must fail on it there.
|
||||
|
Can't render this file because it contains an unexpected character in line 23 and column 25.
|
@@ -0,0 +1,22 @@
|
||||
# secret-patterns.tsv — checked-in pattern list for scripts/secret-scan.sh
|
||||
#
|
||||
# Format: <rule-id><TAB><POSIX ERE><TAB><description><TAB><check>
|
||||
# Blank lines and lines whose first field starts with '#' are ignored.
|
||||
# <check> is optional; the only value today is "value", which tells the scanner
|
||||
# to run the matched value through its inert-value classifier (see
|
||||
# value_is_inert in secret-scan.sh) so bare identifiers, env refs and dotted
|
||||
# code access are not reported as credentials. Omit the column to report every
|
||||
# regex hit.
|
||||
# Matching is case-insensitive, so `API_KEY` and `api_key` both count.
|
||||
#
|
||||
# Add a rule here, never inline in secret-scan.sh: this file is the single
|
||||
# auditable list of what the guard considers credential-shaped.
|
||||
openai-key \bsk-[A-Za-z0-9_-]{16,} OpenAI/LiteLLM-style "sk-" secret key (also hyphenated sk-proj- keys)
|
||||
openrouter-key \bsk-or-v1-[A-Za-z0-9_-]{8,} OpenRouter API key
|
||||
stripe-live-key \bsk_live_[A-Za-z0-9]{8,} Stripe live secret key
|
||||
proxmox-token PVEAPIToken=[^[:space:]"']+ Proxmox API token literal
|
||||
bearer-token bearer[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) literal Bearer token (http header or prose)
|
||||
auth-header authorization:[[:space:]]+["']?(«.{3,}»|[A-Za-z0-9_./+=-]{20,}) Authorization header carrying a raw literal value
|
||||
private-key -----BEGIN [A-Z ]*PRIVATE KEY----- PEM private key block
|
||||
cred-prose credentials?[[:space:]]*[:=][[:space:]]*[^[:space:]] prose credential line carrying a value
|
||||
secret-assign (api[_-]?key|apikey|passwd|password|secret|token)s?["']?[[:space:]]*[:=][[:space:]]*["']?(«.{3,}»|[A-Za-z0-9_./+=-]{8,}) credential assignment carrying a literal value value
|
||||
|
Can't render this file because it contains an unexpected character in line 5 and column 48.
|
Executable
+269
@@ -0,0 +1,269 @@
|
||||
#!/usr/bin/env bash
|
||||
# secret-scan.sh — commit-time secret guard. FAILS (exit 1) on a credential-shaped
|
||||
# string, so a build cannot go green with a credential committed to it.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/secret-scan.sh # scan the whole git-tracked tree (default)
|
||||
# scripts/secret-scan.sh --tree
|
||||
# scripts/secret-scan.sh --path DIR # scan an arbitrary directory (git not required)
|
||||
# scripts/secret-scan.sh --staged # scan added lines in the index (pre-commit)
|
||||
# scripts/secret-scan.sh --diff REF # scan added lines since REF (e.g. origin/master)
|
||||
# --quiet only print the verdict and findings, no per-mode banner
|
||||
#
|
||||
# Exit codes: 0 clean, 1 credential found, 2 usage/config error.
|
||||
#
|
||||
# Patterns live in scripts/secret-patterns.tsv
|
||||
# Exceptions live in scripts/secret-allowlist.tsv (every entry carries a reason;
|
||||
# a missing reason is a hard error, so the guard fails closed).
|
||||
#
|
||||
# Dependencies are deliberately bash + coreutils + grep + sed/awk + git. The
|
||||
# Gitea Actions runner executes job steps INSIDE the runner container, which
|
||||
# has no node and no python by default: keep this script free of both.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SELF_DIR=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||
ROOT=$(cd -- "$SELF_DIR/.." && pwd)
|
||||
PATTERNS_FILE="$SELF_DIR/secret-patterns.tsv"
|
||||
ALLOWLIST_FILE="$SELF_DIR/secret-allowlist.tsv"
|
||||
|
||||
# The guard's own definition files are not scannable content: the pattern list
|
||||
# necessarily contains the pattern text, and the allowlist necessarily contains
|
||||
# the allowed literals. Narrow, exact-path exclusion — not a wildcard.
|
||||
SELF_FILES=(
|
||||
"scripts/secret-scan.sh"
|
||||
"scripts/secret-patterns.tsv"
|
||||
"scripts/secret-allowlist.tsv"
|
||||
)
|
||||
|
||||
MODE="tree"
|
||||
PATH_DIR=""
|
||||
DIFF_REF=""
|
||||
QUIET=0
|
||||
|
||||
usage() {
|
||||
sed -n '2,20p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'
|
||||
exit 2
|
||||
}
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
--tree) MODE="tree" ;;
|
||||
--path) MODE="path"; PATH_DIR="${2:-}"; shift ;;
|
||||
--staged) MODE="staged" ;;
|
||||
--diff) MODE="diff"; DIFF_REF="${2:-}"; shift ;;
|
||||
--quiet) QUIET=1 ;;
|
||||
-h|--help) usage ;;
|
||||
*) echo "secret-scan: unknown argument '$1'" >&2; usage ;;
|
||||
esac
|
||||
shift
|
||||
done
|
||||
|
||||
[ -f "$PATTERNS_FILE" ] || { echo "secret-scan: missing $PATTERNS_FILE" >&2; exit 2; }
|
||||
[ -f "$ALLOWLIST_FILE" ] || { echo "secret-scan: missing $ALLOWLIST_FILE" >&2; exit 2; }
|
||||
if [ "$MODE" = "path" ] && [ -z "$PATH_DIR" ]; then
|
||||
echo "secret-scan: --path needs a directory" >&2; exit 2
|
||||
fi
|
||||
if [ "$MODE" = "diff" ] && [ -z "$DIFF_REF" ]; then
|
||||
echo "secret-scan: --diff needs a base ref" >&2; exit 2
|
||||
fi
|
||||
|
||||
# ── Load patterns ──────────────────────────────────────────────────────────
|
||||
RULE_IDS=()
|
||||
RULE_RES=()
|
||||
RULE_DESCS=()
|
||||
RULE_CHECKS=()
|
||||
COMBINED=""
|
||||
while IFS=$'\t' read -r id re desc check; do
|
||||
case "$id" in ''|'#'*) continue ;; esac
|
||||
[ -n "$re" ] || continue
|
||||
RULE_IDS+=("$id"); RULE_RES+=("$re"); RULE_DESCS+=("$desc"); RULE_CHECKS+=("${check:-}")
|
||||
if [ -z "$COMBINED" ]; then COMBINED="($re)"; else COMBINED="$COMBINED|($re)"; fi
|
||||
done < "$PATTERNS_FILE"
|
||||
if [ "${#RULE_IDS[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: no patterns loaded from $PATTERNS_FILE" >&2; exit 2
|
||||
fi
|
||||
|
||||
# ── Load allowlist (fails closed on a missing reason) ──────────────────────
|
||||
AL_RULES=()
|
||||
AL_GLOBS=()
|
||||
AL_LITS=()
|
||||
AL_REASONS=()
|
||||
AL_LINENO=0
|
||||
while IFS=$'\t' read -r rule glob lit reason; do
|
||||
AL_LINENO=$((AL_LINENO + 1))
|
||||
case "$rule" in ''|'#'*) continue ;; esac
|
||||
if [ -z "$glob" ] || [ -z "$lit" ] || [ -z "$reason" ]; then
|
||||
echo "secret-scan: ❌ $ALLOWLIST_FILE:$AL_LINENO — allowlist entry needs <rule> <path-glob> <literal> <reason>; reason-based exceptions only, refusing to run" >&2
|
||||
exit 2
|
||||
fi
|
||||
AL_RULES+=("$rule"); AL_GLOBS+=("$glob"); AL_LITS+=("$lit"); AL_REASONS+=("$reason")
|
||||
done < "$ALLOWLIST_FILE"
|
||||
|
||||
# nocasematch is toggled only around the regex test; path globs must stay
|
||||
# case-sensitive, so it is never left on.
|
||||
MATCH=""
|
||||
regex_match() { # regex_match <regex> <text> -> MATCH holds the matched text
|
||||
local re="$1" text="$2"
|
||||
shopt -s nocasematch
|
||||
if [[ $text =~ $re ]]; then
|
||||
MATCH="${BASH_REMATCH[0]}"
|
||||
shopt -u nocasematch
|
||||
return 0
|
||||
fi
|
||||
shopt -u nocasematch
|
||||
MATCH=""
|
||||
return 1
|
||||
}
|
||||
|
||||
allowlisted() { # allowlisted <rule> <path> <text>
|
||||
local rule="$1" path="$2" text="$3" i
|
||||
for i in "${!AL_RULES[@]}"; do
|
||||
[ "${AL_RULES[$i]}" = "$rule" ] || [ "${AL_RULES[$i]}" = "*" ] || continue
|
||||
# The unquoted RHS is deliberate: <path-glob> is a bash glob, not a literal.
|
||||
# shellcheck disable=SC2053
|
||||
[[ $path == ${AL_GLOBS[$i]} ]] || continue
|
||||
[[ $text == *"${AL_LITS[$i]}"* ]] || continue
|
||||
return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
mask_value() { # mask_value <text> <match> — never echo a credential to logs.
|
||||
# Print only the part of the line BEFORE the match, then <redacted>: the match
|
||||
# itself and everything after it (which may include a value the rule's regex
|
||||
# stopped short of, e.g. `credentials:` followed by a backticked password) is
|
||||
# never written to stdout.
|
||||
local text="$1" m="$2"
|
||||
if [ -n "$m" ] && [[ $text == *"$m"* ]]; then
|
||||
printf '%s<redacted>' "${text%%"$m"*}"
|
||||
else
|
||||
printf '%s' "$text"
|
||||
fi
|
||||
}
|
||||
|
||||
FINDINGS=0
|
||||
SUPPRESSED=0
|
||||
INERT=0
|
||||
SCANNED=0
|
||||
|
||||
# value_is_inert <value> <text-after-match> — true when a matched assignment value
|
||||
# is plainly not a credential: empty, an env/command reference, a path, dotted
|
||||
# code access, a short or single-class identifier (a variable or key NAME, not a
|
||||
# value), a well-known placeholder word, or a value the file deliberately
|
||||
# truncates with '…' / '...' (a redacted prefix is not a usable credential).
|
||||
# Deliberately does NOT know the words "synthetic" or "example": a fabricated
|
||||
# example must be an explicit allowlist entry.
|
||||
value_is_inert() {
|
||||
local v="$1" rest="$2"
|
||||
case "$rest" in '…'*|'...'*) return 0 ;; esac
|
||||
v="${v%\"}"; v="${v#\"}"; v="${v%\'}"; v="${v#\'}"
|
||||
case "$v" in
|
||||
''|\$*|\{*|'<'*|'%'*|'('*|'/'*|'\\'*) return 0 ;;
|
||||
not-needed|no-key-required|none|null|true|false|redacted|placeholder|example|dummy|changeme|change-me|your-key|your_key|key|token|secret|password) return 0 ;;
|
||||
esac
|
||||
# dotted code access: os.environ.get / process.env.ZULIP_API_KEY / cfg.a
|
||||
if [[ $v =~ ^[a-z_][a-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)+$ ]]; then return 0; fi
|
||||
# bare identifier (no punctuation beyond _): a NAME, not a value. A real
|
||||
# secret in this shape is long and mixes letters with digits.
|
||||
if [[ $v =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
|
||||
[ "${#v}" -lt 20 ] && return 0
|
||||
[[ $v =~ [0-9] ]] || return 0
|
||||
return 1
|
||||
fi
|
||||
return 1
|
||||
}
|
||||
|
||||
report_finding() { # report_finding <path> <line> <text>
|
||||
local path="$1" line="$2" text="$3" i val
|
||||
for i in "${!RULE_IDS[@]}"; do
|
||||
regex_match "${RULE_RES[$i]}" "$text" || continue
|
||||
SCANNED=$((SCANNED + 1))
|
||||
if [ "${RULE_CHECKS[$i]}" = "value" ]; then
|
||||
val="${MATCH#*[:=]}"
|
||||
val="${val# }"
|
||||
if value_is_inert "$val" "${text#*"$MATCH"}"; then
|
||||
INERT=$((INERT + 1))
|
||||
continue
|
||||
fi
|
||||
fi
|
||||
if allowlisted "${RULE_IDS[$i]}" "$path" "$text"; then
|
||||
SUPPRESSED=$((SUPPRESSED + 1))
|
||||
continue
|
||||
fi
|
||||
FINDINGS=$((FINDINGS + 1))
|
||||
printf ' ❌ %s:%s [%s] %s\n' "$path" "$line" "${RULE_IDS[$i]}" "${RULE_DESCS[$i]}"
|
||||
printf ' | %s\n' "$(mask_value "$text" "$MATCH")"
|
||||
done
|
||||
}
|
||||
|
||||
self_excluded() { # self_excluded <repo-relative-path>
|
||||
local p="$1" s
|
||||
for s in "${SELF_FILES[@]}"; do
|
||||
[ "$p" = "$s" ] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
# ── Collect candidate lines and scan them ─────────────────────────────────
|
||||
if [ "$MODE" = "tree" ] || [ "$MODE" = "path" ]; then
|
||||
if [ "$MODE" = "tree" ]; then
|
||||
BASE="$ROOT"
|
||||
git -C "$BASE" rev-parse --git-dir >/dev/null 2>&1 || { echo "secret-scan: --tree needs a git checkout (use --path DIR)" >&2; exit 2; }
|
||||
mapfile -d '' candidate < <(git -C "$BASE" ls-files -z 2>/dev/null)
|
||||
if [ "${#candidate[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: ❌ no tracked files — refusing to report clean" >&2; exit 2
|
||||
fi
|
||||
else
|
||||
BASE=$(cd -- "$PATH_DIR" 2>/dev/null && pwd) || { echo "secret-scan: --path '$PATH_DIR' is not a directory" >&2; exit 2; }
|
||||
mapfile -t candidate < <(cd -- "$BASE" && find . -type f -not -path './.git/*' | sed 's|^\./||')
|
||||
if [ "${#candidate[@]}" -eq 0 ]; then
|
||||
echo "secret-scan: ❌ no files under $BASE — refusing to report clean" >&2; exit 2
|
||||
fi
|
||||
fi
|
||||
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan ($MODE): ${#candidate[@]} files under $BASE ──"
|
||||
for rel in "${candidate[@]}"; do
|
||||
[ -f "$BASE/$rel" ] || continue
|
||||
self_excluded "$rel" && continue
|
||||
while IFS= read -r hit; do
|
||||
[ -n "$hit" ] || continue
|
||||
report_finding "$rel" "${hit%%:*}" "${hit#*:}"
|
||||
done < <(grep -nEIi -e "$COMBINED" "$BASE/$rel" 2>/dev/null || true)
|
||||
done
|
||||
else
|
||||
# --staged / --diff: only ADDED lines, with the post-change line number.
|
||||
if [ "$MODE" = "staged" ]; then
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines in the index ──"
|
||||
DIFF_TEXT=$(git -C "$ROOT" diff --cached --unified=0 --no-color -- . 2>/dev/null)
|
||||
else
|
||||
[ "$QUIET" -eq 1 ] || echo "── secret scan: added lines since $DIFF_REF ──"
|
||||
DIFF_TEXT=$(git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"...HEAD 2>/dev/null \
|
||||
|| git -C "$ROOT" diff --unified=0 --no-color "$DIFF_REF"..HEAD 2>/dev/null)
|
||||
fi
|
||||
if [ -z "$DIFF_TEXT" ]; then
|
||||
[ "$QUIET" -eq 1 ] || echo " (no added lines)"
|
||||
fi
|
||||
while IFS=$'\t' read -r rel line text; do
|
||||
[ -n "$rel" ] || continue
|
||||
self_excluded "$rel" && continue
|
||||
report_finding "$rel" "$line" "$text"
|
||||
done < <(printf '%s\n' "$DIFF_TEXT" | awk '
|
||||
/^\+\+\+ / { f=$2; sub(/^b\//,"",f); next }
|
||||
/^@@ / { if (match($0, /\+[0-9]+/)) ln=substr($0, RSTART+1, RLENGTH-1)+0; next }
|
||||
(/^\+/ && !/^\+\+\+/) { print f "\t" ln "\t" substr($0,2); ln++; next }
|
||||
')
|
||||
fi
|
||||
|
||||
# ── Verdict ────────────────────────────────────────────────────────────────
|
||||
if [ "$FINDINGS" -gt 0 ]; then
|
||||
echo ""
|
||||
echo "❌ SECRET SCAN FAILED — $FINDINGS credential-shaped string(s) in ${MODE} content."
|
||||
echo " Fix: remove the credential and read it from the vault/env."
|
||||
echo " Only a deliberate synthetic example may be added to scripts/secret-allowlist.tsv,"
|
||||
echo " one entry per file/rule/literal, with a reason. Never allowlist a live credential."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "✅ secret scan clean (${MODE}; ${SUPPRESSED} allowlisted exception(s), ${INERT} inert value(s) ignored)"
|
||||
exit 0
|
||||
Executable
+228
@@ -0,0 +1,228 @@
|
||||
#!/bin/bash
|
||||
# test_infra_monitoring.sh — Asserts that probe calls use the documented targets.
|
||||
#
|
||||
# Strategy: stub curl and ssh on PATH to capture the exact arguments each leg
|
||||
# builds, then assert the URL/port of every call. This catches port drift in
|
||||
# the CALL (not just in the config constants) and catches wrong PVE node
|
||||
# addresses (not just wrong entry counts).
|
||||
#
|
||||
# Run: bash scripts/test_infra_monitoring.sh
|
||||
# Exits 0 if all assertions pass, 1 otherwise.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
SCRIPT="${SCRIPT_DIR}/infra-monitoring.sh"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
assert() {
|
||||
local desc="$1" condition="$2"
|
||||
if eval "$condition"; then
|
||||
echo " ✅ $desc"
|
||||
PASS=$((PASS+1))
|
||||
else
|
||||
echo " 🔴 $desc"
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
}
|
||||
|
||||
echo "=== test_infra_monitoring.sh ==="
|
||||
echo ""
|
||||
|
||||
# ── Stub curl: capture argv to a file, return 200 ──────────────────────────
|
||||
STUB_DIR=$(mktemp -d)
|
||||
trap 'rm -rf "$STUB_DIR"' EXIT
|
||||
|
||||
# Stub curl: first arg after flags is the URL; capture all args
|
||||
cat > "$STUB_DIR/curl" << 'STUBEOF'
|
||||
#!/bin/bash
|
||||
echo "$@" >> "${CURL_STUB_LOG:-/dev/null}"
|
||||
# Print 200 for %{http_code}
|
||||
printf '%s\n' "200"
|
||||
exit 0
|
||||
STUBEOF
|
||||
chmod +x "$STUB_DIR/curl"
|
||||
|
||||
# Stub ssh: first arg after options is the remote command; capture it
|
||||
cat > "$STUB_DIR/ssh" << 'SSTUBEOF'
|
||||
#!/bin/bash
|
||||
echo "SSH $@" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||
# The last arg is the remote command — extract and log curl args
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == curl* ]]; then
|
||||
echo "$arg" >> "${SSH_STUB_LOG:-/dev/null}"
|
||||
fi
|
||||
done
|
||||
printf '%s\n' "200"
|
||||
exit 0
|
||||
SSTUBEOF
|
||||
chmod +x "$STUB_DIR/ssh"
|
||||
|
||||
# ── Run the monitor with stubs ────────────────────────────────────────────
|
||||
CURL_LOG="$STUB_DIR/curl_calls.log"
|
||||
SSH_LOG="$STUB_DIR/ssh_calls.log"
|
||||
touch "$CURL_LOG" "$SSH_LOG"
|
||||
|
||||
CURL_STUB_LOG="$CURL_LOG" SSH_STUB_LOG="$SSH_LOG" \
|
||||
PATH="$STUB_DIR:$PATH" bash "$SCRIPT" > "$STUB_DIR/output.txt" 2>&1
|
||||
|
||||
# ── 1. Port drift detection (from actual curl invocations) ─────────────────
|
||||
|
||||
assert "Grafana probed at port 3001" \
|
||||
'grep -q "http://192.168.68.116:3001/api/health" "$CURL_LOG"'
|
||||
|
||||
assert "Prometheus probed at port 9090" \
|
||||
'grep -q "http://192.168.68.116:9090/-/healthy" "$CURL_LOG"'
|
||||
|
||||
assert "LiteLLM probed via nginx at port 80" \
|
||||
'grep -q "http://192.168.68.116:80/litellm/health" "$CURL_LOG"'
|
||||
|
||||
assert "PVE API probed at port 8006" \
|
||||
'grep -q ":8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "GPU exporter probed at port 9400" \
|
||||
'grep -q ":9400/metrics" "$CURL_LOG"'
|
||||
|
||||
# ── 2. PVE API: exact node addresses (catches wrong IPs) ──────────────────
|
||||
# Each real PVE node must be probed; CT 116 must NOT be in the PVE set
|
||||
|
||||
assert "PVE acerpve 192.168.68.9 probed" \
|
||||
'grep -q "https://192.168.68.9:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE minipve 192.168.68.12 probed" \
|
||||
'grep -q "https://192.168.68.12:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE storepve 192.168.68.6 probed" \
|
||||
'grep -q "https://192.168.68.6:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE amdpve 192.168.68.15 probed" \
|
||||
'grep -q "https://192.168.68.15:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
assert "PVE ocupve 192.168.68.5 probed" \
|
||||
'grep -q "https://192.168.68.5:8006/api2/json/version" "$CURL_LOG"'
|
||||
|
||||
# CT 116 (.116) must NOT appear as a PVE API target
|
||||
assert "CT 116 (.116) NOT probed as PVE API node" \
|
||||
'! grep -q "https://192.168.68.116:8006" "$CURL_LOG"'
|
||||
|
||||
# ── 3. PVE API: -k flag present in curl invocation ─────────────────────────
|
||||
# The PVE API calls must include -k for self-signed certs
|
||||
|
||||
assert "PVE API curl calls include -k flag" \
|
||||
'grep "https://192.168.68.9:8006" "$CURL_LOG" | grep -q -- "-k"'
|
||||
|
||||
# ── 4. Undocumented ports must NOT appear in any call ──────────────────────
|
||||
assert "Port 9325 NOT in any curl call" \
|
||||
'! grep -q ":9325" "$CURL_LOG"'
|
||||
|
||||
assert "Port 9405 NOT in any curl call" \
|
||||
'! grep -q ":9405" "$CURL_LOG"'
|
||||
|
||||
# ── 5. Docker Stats / PVE Exporter: SSH-probed at correct ports ────────────
|
||||
assert "Docker Stats probed at port 9324 via SSH" \
|
||||
'grep -q "9324" "$SSH_LOG"'
|
||||
|
||||
assert "PVE Exporter probed at port 9221 via SSH" \
|
||||
'grep -q "9221" "$SSH_LOG"'
|
||||
|
||||
# ── 5a. Per-leg assertions (proves which leg owns which port) ──────────────
|
||||
# The script source must show Docker Stats using $DOCKER_STATS_PORT and
|
||||
# PVE Exporter using $PVE_EXPORTER_PORT in the correct leg sections
|
||||
assert "Docker Stats leg uses DOCKER_STATS_PORT constant" \
|
||||
'grep -A 3 "# 6. Docker Stats" "$SCRIPT" | grep -q "\$DOCKER_STATS_PORT"'
|
||||
|
||||
assert "PVE Exporter leg uses PVE_EXPORTER_PORT constant" \
|
||||
'grep -A 3 "# 7. PVE Exporter" "$SCRIPT" | grep -q "\$PVE_EXPORTER_PORT"'
|
||||
|
||||
# Verify the constants themselves are set to the correct values
|
||||
assert "DOCKER_STATS_PORT constant set to 9324" \
|
||||
'grep -q "^DOCKER_STATS_PORT=\"9324\"" "$SCRIPT"'
|
||||
|
||||
assert "PVE_EXPORTER_PORT constant set to 9221" \
|
||||
'grep -q "^PVE_EXPORTER_PORT=\"9221\"" "$SCRIPT"'
|
||||
|
||||
# ── 5b. Stale port 9323 (dockerd) NOT probed ──────────────────────────────
|
||||
assert "Port 9323 (dockerd) NOT in SSH log" \
|
||||
'! grep -q "9323" "$SSH_LOG"'
|
||||
|
||||
# ── 6. No stale ports in the script source (belt-and-suspenders) ──────────
|
||||
assert "Port 9325 (historical) NOT in script source" \
|
||||
'! grep -q "9325" "$SCRIPT"'
|
||||
|
||||
assert "Port 9405 (historical) NOT in script source" \
|
||||
'! grep -q "9405" "$SCRIPT"'
|
||||
|
||||
# ── 7. Failure-line content includes non-empty kind ────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||
|
||||
# Test 7a: Unexpected status (500) → kind should be unexpected:500
|
||||
cat > "$TMP_DIR/curl" << 'EOF'
|
||||
#!/bin/bash
|
||||
# Stub: return 500 for Grafana port, 200 otherwise
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == *":3001"* ]]; then
|
||||
echo "500"
|
||||
exit 0
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/curl"
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||
assert "Grafana failure line exists (unexpected status)" \
|
||||
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||
assert "Grafana failure kind is non-empty (unexpected status)" \
|
||||
'[[ -n "$KIND" ]]'
|
||||
|
||||
# Test 7b: TLS error (000 + exit 60) → kind should be tls
|
||||
cat > "$TMP_DIR/curl" << 'EOF'
|
||||
#!/bin/bash
|
||||
# Stub: return 000 and exit 60 for Grafana port (TLS error) on ALL invocations
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == *":3001"* ]]; then
|
||||
echo "000"
|
||||
exit 60
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/curl"
|
||||
# Also stub ssh to return 000 + exit 60 for the retry
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
for arg in "$@"; do
|
||||
if [[ "$arg" == curl* ]]; then
|
||||
echo "000"
|
||||
exit 60
|
||||
fi
|
||||
done
|
||||
echo "200"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
export PATH="$TMP_DIR:$PATH"
|
||||
OUT=$(bash "$SCRIPT" 2>&1)
|
||||
GRAFANA_FAIL=$(echo "$OUT" | grep "Grafana: probe-failed")
|
||||
assert "Grafana failure line exists (TLS error)" \
|
||||
'[[ -n "$GRAFANA_FAIL" ]]'
|
||||
KIND=$(echo "$GRAFANA_FAIL" | grep -oP '\(<[^>]+>\)' | tr -d '()<>')
|
||||
assert "Grafana failure kind is tls" \
|
||||
'[[ "$KIND" == "tls" ]]'
|
||||
|
||||
# ── Summary ─────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||
if [ $FAIL -gt 0 ]; then
|
||||
echo " 🔴 TESTS FAILED"
|
||||
exit 1
|
||||
else
|
||||
echo " ✅ ALL TESTS PASSED"
|
||||
exit 0
|
||||
fi
|
||||
Executable
+125
@@ -0,0 +1,125 @@
|
||||
#!/bin/bash
|
||||
# test_proxmox_monitor.sh — Stub-driven tests for PBS GC liveness leg
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT="$(cd "$(dirname "$0")" && pwd)/proxmox-monitor.sh"
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
# ── Helpers ────────────────────────────────────────────────────────────────
|
||||
assert() {
|
||||
local desc="$1" cond="$2"
|
||||
if eval "$cond" 2>/dev/null; then
|
||||
echo " ✅ $desc"
|
||||
PASS=$((PASS+1))
|
||||
else
|
||||
echo " 🔴 $desc"
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
}
|
||||
|
||||
# ── 1. Healthy: fresh GC, 0 B pending ─────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
FRESH_ENDTIME=$(( $(date -u +%s) - (1 * 3600) )) # 1 hour ago
|
||||
cat > "$TMP_DIR/ssh" << EOF
|
||||
#!/bin/bash
|
||||
# Stub: return valid JSON with fresh endtime
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":$FRESH_ENDTIME,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Healthy: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Healthy: shows healthy verdict" '[[ "$PBS_LINE" == *"healthy"* ]]'
|
||||
assert "Healthy: shows pending-bytes 0 B" '[[ "$PBS_LINE" == *"pending-bytes: 0 B"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 2. Stale: GC >48h old ─────────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
OLD_ENDTIME=$(( $(date -u +%s) - (49 * 3600) )) # 49 hours ago
|
||||
cat > "$TMP_DIR/ssh" << EOF
|
||||
#!/bin/bash
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":$OLD_ENDTIME,"pending-bytes":1048576}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Stale: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Stale: shows stale verdict" '[[ "$PBS_LINE" == *"stale"* ]]'
|
||||
assert "Stale: shows pending-bytes 1048576 B" '[[ "$PBS_LINE" == *"pending-bytes: 1048576 B"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 3. Probe-failed: empty output ─────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Probe-failed (empty): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Probe-failed (empty): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 4. Probe-failed: unparseable output ───────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo "proxmox-backup-manager: command not found"
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Probe-failed (unparseable): PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Probe-failed (unparseable): shows probe-failed" '[[ "$PBS_LINE" == *"probe-failed"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 5. Null endtime: never-run ────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo '[{"store":"storepve-datastore","last-run-endtime":null,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Null endtime: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Null endtime: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── 6. Datastore absent ───────────────────────────────────────────────────
|
||||
TMP_DIR=$(mktemp -d)
|
||||
cat > "$TMP_DIR/ssh" << 'EOF'
|
||||
#!/bin/bash
|
||||
echo '[{"store":"s3-archive","last-run-endtime":null,"pending-bytes":0}]'
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "$TMP_DIR/ssh"
|
||||
|
||||
OUT=$(PATH="$TMP_DIR:$PATH" bash "$SCRIPT" 2>&1)
|
||||
PBS_LINE=$(echo "$OUT" | grep "PBS GC:" || echo "")
|
||||
assert "Datastore absent: PBS line exists" '[[ -n "$PBS_LINE" ]]'
|
||||
assert "Datastore absent: shows never-run" '[[ "$PBS_LINE" == *"never-run"* ]]'
|
||||
rm -rf "$TMP_DIR"
|
||||
|
||||
# ── Summary ────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "Results: ${PASS} passed, ${FAIL} failed"
|
||||
if [ $FAIL -gt 0 ]; then
|
||||
echo " 🔴 TESTS FAILED"
|
||||
exit 1
|
||||
else
|
||||
echo " ✅ ALL TESTS PASSED"
|
||||
exit 0
|
||||
fi
|
||||
+71
-21
@@ -8,12 +8,21 @@
|
||||
set -euo pipefail
|
||||
|
||||
# Credentials sourced from environment variable ZULIP_API_KEY (set by vault-backed start script)
|
||||
# Never fall back to a literal key
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:?ZULIP_API_KEY not set — refusing to run with no credential}"
|
||||
# Never fall back to a literal key.
|
||||
# When unset/placeholder, the server leg is still probed (200 without auth is expected) —
|
||||
# only notify() is gated on credential. The pi/Tanko/kagentz
|
||||
# legs do not need the Zulip API key. The placeholder is captain-held:
|
||||
# zulip-health-credential-placeholder-20260913.
|
||||
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
||||
ZULIP_SITE="https://chat.sysloggh.net"
|
||||
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
||||
OWNER_ZULIP_ID="9"
|
||||
|
||||
# Track whether the Zulip API credential is usable
|
||||
ZULIP_CRED_OK=1
|
||||
if [ -z "$ZULIP_API_KEY" ] || [[ "$ZULIP_API_KEY" == *"placeholder"* ]] || [[ "$ZULIP_API_KEY" == *"REDACTED"* ]]; then
|
||||
ZULIP_CRED_OK=0
|
||||
fi
|
||||
|
||||
LOG="/root/zulip-health-monitor.log"
|
||||
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
||||
@@ -24,23 +33,29 @@ notify() {
|
||||
local severity="$1" msg="$2"
|
||||
echo "[$severity] $msg"
|
||||
|
||||
# Zulip DM to owner
|
||||
local content="${severity} Zulip Monitor: ${msg}"
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||
> /dev/null 2>&1 \
|
||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
|
||||
# Zulip DM to owner (skip if no credential)
|
||||
if [ "$ZULIP_CRED_OK" -eq 1 ]; then
|
||||
local content="${severity} Zulip Monitor: ${msg}"
|
||||
local form
|
||||
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "${form}" > /dev/null 2>&1 || true
|
||||
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
||||
local stream_content="${severity} Zulip Monitor: ${msg}"
|
||||
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" \
|
||||
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
||||
> /dev/null 2>&1 \
|
||||
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)">> "$LOG"
|
||||
else
|
||||
echo " ALERT SUPPRESSED (no credential): ${severity} ${msg}" >> "$LOG"
|
||||
fi
|
||||
}
|
||||
|
||||
# ── Global: Zulip Server ──
|
||||
# F3: Always probe server regardless of credential — 200 without auth is expected
|
||||
# (verified live: server_settings returns 200 with no credential or wrong key).
|
||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||
https://chat.sysloggh.net/api/v1/server_settings \
|
||||
-u "${ZULIP_EMAIL}:${ZULIP_API_KEY}" 2>/dev/null) || SERVER_CODE="000"
|
||||
@@ -161,6 +176,11 @@ fi
|
||||
# never contact her former host.
|
||||
|
||||
# ── Platform C: Agent Zero (kagentz) ──
|
||||
# C1: A2A liveness (no credential needed) — probes the container's internal :80/a2a/
|
||||
# C2: A2A response verification (needs LITELLM_KEY) — probes POST /a2a with auth
|
||||
# C3: Public access path (no credential needed) — probes https://kagentz.sysloggh.net/
|
||||
|
||||
# C1: A2A liveness (container-internal probe)
|
||||
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
|
||||
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||
@@ -169,23 +189,53 @@ AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||
if [ "$AZ_A2A_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz A2A server DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||
echo " kagentz C1: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||
else
|
||||
case "$AZ_A2A_CODE" in
|
||||
200|401)
|
||||
echo " kagentz: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||
echo " kagentz C1: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||
echo " kagentz C1: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# C3: Public access path (the captain's point of view)
|
||||
# Probes the public URL that NetBird proxies to the container. 200/302/401 = alive,
|
||||
# 502 = proxy's "upstream refused" page (incident), connection failed = incident.
|
||||
# Never restarts anything — the contract forbids restarting the platform.
|
||||
KAGENTZ_PUBLIC_CODE=$(curl -s -o /dev/null --connect-timeout 10 --max-time 15 \
|
||||
-w '%{http_code}' https://kagentz.sysloggh.net/ 2>/dev/null) || KAGENTZ_PUBLIC_CODE="000"
|
||||
KAGENTZ_PUBLIC_CODE=$(printf '%s' "$KAGENTZ_PUBLIC_CODE" | tr -d '[:space:]')
|
||||
[ -n "$KAGENTZ_PUBLIC_CODE" ] || KAGENTZ_PUBLIC_CODE="000"
|
||||
|
||||
if [ "$KAGENTZ_PUBLIC_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz public URL DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: ❌ public URL down (HTTP 000)" >> "$LOG"
|
||||
elif [ "$KAGENTZ_PUBLIC_CODE" = "502" ]; then
|
||||
notify "🔴" "kagentz public URL 502 (upstream refused)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: ❌ public URL 502 (upstream refused)" >> "$LOG"
|
||||
else
|
||||
case "$KAGENTZ_PUBLIC_CODE" in
|
||||
200|302|401)
|
||||
echo " kagentz C3: ✅ public URL alive (HTTP $KAGENTZ_PUBLIC_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz public URL answered HTTP $KAGENTZ_PUBLIC_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz C3: 🟡 public URL unexpected http=$KAGENTZ_PUBLIC_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# ── Summary ──
|
||||
# The run verdict is non-optimistic: when ISSUES > 0, the run is an INCIDENT.
|
||||
# The lane must quote this Result line verbatim in its status report.
|
||||
if [ "$ISSUES" -eq 0 ]; then
|
||||
echo " Result: ✅ All healthy" >> "$LOG"
|
||||
echo " Result: ✅ 0 issues (all healthy)" >> "$LOG"
|
||||
else
|
||||
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
||||
echo " Result: 🔴 INCIDENT — $ISSUES issue(s) found" >> "$LOG"
|
||||
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
||||
fi
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ BASE = """
|
||||
model:
|
||||
api_key: ""
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
max_tokens: 4096
|
||||
default: syslog-auto
|
||||
provider: harness
|
||||
@@ -52,7 +52,7 @@ delegation:
|
||||
custom_providers:
|
||||
- name: harness
|
||||
key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
"""
|
||||
|
||||
|
||||
@@ -189,3 +189,73 @@ def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
|
||||
assert code == 1, out
|
||||
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_canonical_internal_path_passes(tmp_path):
|
||||
"""Rule 5 must accept the canonical internal base from hermes-key-enforcement.prose.md."""
|
||||
code, out = _run(
|
||||
tmp_path,
|
||||
"gpu-vision",
|
||||
)
|
||||
# Override the base_url in the config
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "canonical-internal.yaml", cfg_text)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
# Verify the correct message is shown
|
||||
assert "model.base_url is canonical" in out
|
||||
|
||||
|
||||
def test_wrong_base_url_fails(tmp_path):
|
||||
"""Rule 5 must reject paths outside the allowed list."""
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
" base_url: http://192.168.68.116/litellm/v1/responses",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "wrong-base.yaml", cfg_text)
|
||||
assert code == 1, out
|
||||
assert "RESULT: FAIL" in out
|
||||
assert "model.base_url must be one of" in out
|
||||
|
||||
|
||||
def test_public_host_path_passes(tmp_path):
|
||||
"""Rule 5 must accept the public host base."""
|
||||
cfg_text = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
" base_url: https://litellm.sysloggh.net/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "public-host.yaml", cfg_text)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_old_rule5_check_would_fail_canonical(tmp_path):
|
||||
"""
|
||||
Proof that the OLD Rule 5 check would fail the canonical internal path.
|
||||
This proves the bug existed before the fix.
|
||||
"""
|
||||
# OLD check expected /v1, so the canonical /litellm/v1 would have failed
|
||||
canonical_cfg = BASE.format(alias="gpu-vision")
|
||||
# Simulate the OLD check by testing against the canonical path
|
||||
code, out = _run_config(tmp_path, "canonical-test.yaml", canonical_cfg)
|
||||
# NEW check: canonical /litellm/v1 SHOULD pass
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
# OLD check expected /v1, so the internal /v1 would have passed
|
||||
# NEW check: internal /v1 is non-canonical but working (WARN not FAIL)
|
||||
old_cfg = BASE.format(alias="gpu-vision").replace(
|
||||
" base_url: http://192.168.68.116/litellm/v1",
|
||||
" base_url: http://192.168.68.116/v1",
|
||||
)
|
||||
code, out = _run_config(tmp_path, "old-check-test.yaml", old_cfg)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
# NEW check should pass
|
||||
new_cfg = BASE.format(alias="gpu-vision")
|
||||
code, out = _run_config(tmp_path, "new-check-test.yaml", new_cfg)
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
@@ -98,8 +98,8 @@ exit 0
|
||||
"""
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
|
||||
# record every call (including notify) payloads.
|
||||
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||
# kagentz C3 public URL, and record every call (including notify) payloads.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
@@ -109,6 +109,8 @@ case "$*" in
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
*kagentz.sysloggh.net*)
|
||||
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
@@ -121,7 +123,8 @@ def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a_code="401", az_a2a_exit=0):
|
||||
az_a2a_code="401", az_a2a_exit=0,
|
||||
kagentz_public_code="302"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
@@ -154,6 +157,7 @@ def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
@@ -169,8 +173,9 @@ def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
||||
assert "Server: ✅ HTTP 200" in log
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "Tanko: ✅ service=active http=200" in log
|
||||
assert "kagentz: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "Result: ✅ All healthy" in log
|
||||
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||
|
||||
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||
assert proc.stdout == ""
|
||||
@@ -207,8 +212,8 @@ def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
||||
# The rest of the monitor still ran alongside the failing Tanko leg.
|
||||
log = log_path.read_text()
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "kagentz: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "kagentz C1: ✅ A2A alive" in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
|
||||
|
||||
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||
@@ -216,9 +221,9 @@ def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||
assert "kagentz: ✅ A2A alive" not in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "kagentz C1: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||
assert "kagentz C1: ✅ A2A alive" not in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
assert "kagentz A2A server answered HTTP 500" in proc.stdout
|
||||
|
||||
|
||||
@@ -230,10 +235,10 @@ def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz: ❌ A2A down (HTTP 000)" in log
|
||||
assert "kagentz: ✅ A2A alive" not in log
|
||||
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||
assert "kagentz C1: ✅ A2A alive" not in log
|
||||
assert "unexpected" not in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "Result: 🔴 INCIDENT — 1 issue(s) found" in log
|
||||
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
|
||||
|
||||
|
||||
|
||||
Executable
+153
@@ -0,0 +1,153 @@
|
||||
#!/usr/bin/env bash
|
||||
# test_secret_scan.sh — self-test for the commit-time secret guard.
|
||||
#
|
||||
# Run: bash tests/test_secret_scan.sh
|
||||
# Exit: 0 all cases passed, 1 a case failed.
|
||||
#
|
||||
# WHY THIS FILE EXISTS: a scanner that is never observed to fail is not a guard.
|
||||
# Every fixture below is fabricated and pattern-shaped; the test writes it to a
|
||||
# temp tree (a path no allowlist entry covers) and asserts the guard FAILS. The
|
||||
# same fixtures are deliberately listed in scripts/secret-allowlist.tsv, so the
|
||||
# repo-wide tree scan stays quiet while a planted copy still bites — that is the
|
||||
# difference between an explicit, reasoned exception and a guard trained to
|
||||
# ignore a word.
|
||||
#
|
||||
# Only bash + coreutils + grep. No python/node: the Gitea runner executes job
|
||||
# steps inside the runner container, which has neither.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
HERE=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||
ROOT=$(cd -- "$HERE/.." && pwd)
|
||||
SCAN="$ROOT/scripts/secret-scan.sh"
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
LAST_OUT=""
|
||||
|
||||
ok() { PASS=$((PASS + 1)); echo " ✅ $1"; }
|
||||
bad() { FAIL=$((FAIL + 1)); echo " ❌ $1"; }
|
||||
|
||||
expect_exit() { # expect_exit <want-code> <label> <cmd...>
|
||||
local want="$1" label="$2"; shift 2
|
||||
local rc
|
||||
LAST_OUT=$("$@" 2>&1); rc=$?
|
||||
if [ "$rc" -eq "$want" ]; then ok "$label (exit $rc)"; else
|
||||
bad "$label (wanted exit $want, got $rc)"
|
||||
printf '%s\n' "$LAST_OUT" | sed 's/^/ /' | head -8
|
||||
fi
|
||||
}
|
||||
|
||||
expect_contains() { # expect_contains <label> <needle>
|
||||
if printf '%s' "$LAST_OUT" | grep -qF -- "$2"; then ok "$1"; else
|
||||
bad "$1 (output did not mention: $2)"
|
||||
fi
|
||||
}
|
||||
|
||||
TMPROOT=$(mktemp -d)
|
||||
trap 'rm -rf "$TMPROOT"' EXIT
|
||||
|
||||
echo "── secret-scan self-test ──"
|
||||
|
||||
# ── 1. Guard syntax ───────────────────────────────────────────────────────
|
||||
expect_exit 0 "scanner parses with bash -n" bash -n "$SCAN"
|
||||
|
||||
# ── 2. Guard FAILS on planted, pattern-matching fixtures ──────────────────
|
||||
mkdir -p "$TMPROOT/planted"
|
||||
cat > "$TMPROOT/planted/ops.env" <<'EOF'
|
||||
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
|
||||
EOF
|
||||
expect_exit 1 "planted sk-or-v1 key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted sk-or-v1 key names the openrouter-key rule" "[openrouter-key]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/curl.sh" <<'EOF'
|
||||
curl -s -H "Authorization: Bearer aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaabbbbbbbb" http://example.invalid/
|
||||
EOF
|
||||
expect_exit 1 "planted literal Bearer token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted Bearer token names the bearer-token rule" "[bearer-token]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/pve.sh" <<'EOF'
|
||||
AUTH="Authorization: PVEAPIToken=root@pam!monitor=11111111-2222-3333-4444-555555555555"
|
||||
EOF
|
||||
expect_exit 1 "planted Proxmox token fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted Proxmox token names the proxmox-token rule" "[proxmox-token]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/deploy-key.pem" <<'EOF'
|
||||
-----BEGIN OPENSSH PRIVATE KEY-----
|
||||
b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW
|
||||
-----END OPENSSH PRIVATE KEY-----
|
||||
EOF
|
||||
expect_exit 1 "planted PEM private key fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted PEM key names the private-key rule" "[private-key]"
|
||||
|
||||
# Prose is scanned exactly like code — the original exposures were in .md files.
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/handover.md" <<'EOF'
|
||||
- Admin credentials: `admin` / `correct-horse-battery-staple`
|
||||
EOF
|
||||
expect_exit 1 "planted prose credential line fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted prose line names the cred-prose rule" "[cred-prose]"
|
||||
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/config.env" <<'EOF'
|
||||
DB_PASSWORD=correct-horse-battery-staple
|
||||
EOF
|
||||
expect_exit 1 "planted password assignment fails the guard" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
expect_contains "planted password assignment names the secret-assign rule" "[secret-assign]"
|
||||
|
||||
# ── 3. Guard stays QUIET on inert values and on the real tree ─────────────
|
||||
mkdir -p "$TMPROOT/inert"
|
||||
cat > "$TMPROOT/inert/config.yaml" <<'EOF'
|
||||
api_key: not-needed
|
||||
bearer_token=monitor_key
|
||||
api_key: $LITELLM_API_KEY
|
||||
EOF
|
||||
expect_exit 0 "env refs, sentinels and variable names are not credentials" bash "$SCAN" --path "$TMPROOT/inert" --quiet
|
||||
|
||||
expect_exit 0 "current repo tree passes the guard" bash "$SCAN" --tree
|
||||
expect_contains "tree run reports the allowlisted exceptions it applied" "allowlisted exception(s)"
|
||||
|
||||
# ── 4. Allowlist entries are path-explicit, not word-based ────────────────
|
||||
# This exact line is allowlisted in infrastructure-control.prose.md; the same
|
||||
# text at an unlisted path must still fail, proving the exception is per-file
|
||||
# and reviewed, not a blanket "ignore the word vault".
|
||||
rm -f "$TMPROOT/planted/"*
|
||||
cat > "$TMPROOT/planted/unlisted.md" <<'EOF'
|
||||
- Admin credentials: `«vault: infrastructure/production STIRLING_ADMIN_PASSWORD»`
|
||||
EOF
|
||||
expect_exit 1 "allowlisted text at an unlisted path still fails" bash "$SCAN" --path "$TMPROOT/planted" --quiet
|
||||
|
||||
# ── 5. Commit-time mode: the guard blocks a STAGED credential ─────────────
|
||||
# A throwaway git repo with its own copy of the scanner, so this exercises the
|
||||
# real pre-commit path (--staged) without touching this repo's index.
|
||||
mkdir -p "$TMPROOT/repo/scripts"
|
||||
cp "$SCAN" "$TMPROOT/repo/scripts/secret-scan.sh"
|
||||
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/repo/scripts/secret-patterns.tsv"
|
||||
cp "$ROOT/scripts/secret-allowlist.tsv" "$TMPROOT/repo/scripts/secret-allowlist.tsv"
|
||||
git -C "$TMPROOT/repo" init -q
|
||||
git -C "$TMPROOT/repo" -c user.email=t@example.invalid -c user.name=test commit -q --allow-empty -m base
|
||||
cat > "$TMPROOT/repo/planted.env" <<'EOF'
|
||||
OPENROUTER_API_KEY=sk-or-v1-00000000000000000000000000000000000000000000000000000000deadbeef
|
||||
EOF
|
||||
git -C "$TMPROOT/repo" add planted.env
|
||||
expect_exit 1 "staged credential fails at commit time (--staged)" bash "$TMPROOT/repo/scripts/secret-scan.sh" --staged --quiet
|
||||
expect_contains "staged credential names the openrouter-key rule" "[openrouter-key]"
|
||||
|
||||
# ── 6. Fail closed: an allowlist entry without a reason is a hard error ───
|
||||
mkdir -p "$TMPROOT/scanner" "$TMPROOT/clean"
|
||||
cp "$SCAN" "$TMPROOT/scanner/secret-scan.sh"
|
||||
cp "$ROOT/scripts/secret-patterns.tsv" "$TMPROOT/scanner/secret-patterns.tsv"
|
||||
printf '*\t*.md\twhatever\n' > "$TMPROOT/scanner/secret-allowlist.tsv"
|
||||
echo "placeholder" > "$TMPROOT/clean/ok.md"
|
||||
expect_exit 2 "allowlist entry with no reason fails closed" bash "$TMPROOT/scanner/secret-scan.sh" --path "$TMPROOT/clean" --quiet
|
||||
|
||||
# ── Verdict ───────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
if [ "$FAIL" -gt 0 ]; then
|
||||
echo "❌ secret-scan self-test FAILED — $PASS passed, $FAIL failed"
|
||||
exit 1
|
||||
fi
|
||||
echo "✅ secret-scan self-test passed ($PASS cases)"
|
||||
@@ -0,0 +1,209 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Behavioural tests for zulip-monitor.sh kagentz C1/C3 legs and the Result verdict.
|
||||
|
||||
WHY THIS FILE EXISTS: the 2026-09-19 kagentz A2A outage was correctly detected
|
||||
by the monitor (C1 returned 000, the run logged an issue) but the lane's own
|
||||
summarization in ops.status wrote "OK" with the note "A2A server DOWN — expected
|
||||
(no credentials configured)". The optimistic verdict came from the lane, not the
|
||||
script. The fix adds a C3 public-access-path leg and makes the Result line say
|
||||
"INCIDENT" when issues are found, so the lane can quote it verbatim.
|
||||
|
||||
CONTRACT UNDER TEST:
|
||||
* C1 (A2A liveness, no credential): 000 → INCIDENT.
|
||||
* C3 (public access path, no credential): 502 → INCIDENT, 000 → INCIDENT,
|
||||
200/302/401 → alive.
|
||||
* Result verdict: when ISSUES > 0, the log's Result line says "INCIDENT",
|
||||
not just "issues found".
|
||||
* Healthy control: C1 401 + C3 302 → 0 issues, "all healthy".
|
||||
|
||||
HOW: behavioural execution using the sandbox pattern already in this repo
|
||||
(tests/test_mumuni_monitor_removal.py). The sandbox copies the shipped monitor
|
||||
verbatim, rewrites only its LOG constant, and runs it with stub ssh/curl on
|
||||
PATH. Each test asserts from the run's own log/verdict, not from file text.
|
||||
|
||||
Usage: python3 -m pytest tests/test_zulip_kagentz_legs.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import pathlib
|
||||
import stat
|
||||
import subprocess
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
|
||||
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
|
||||
|
||||
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
|
||||
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
|
||||
|
||||
|
||||
# ── Stub ssh: answers Tanko and Agent Zero probes by env vars ──────────
|
||||
|
||||
SSH_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub ssh: record the target host, then answer by host + remote command.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
|
||||
host=""
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
*@192.168.*) host="${a##*@}" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
|
||||
cmd="${*: -1}"
|
||||
case "$host" in
|
||||
192.168.68.15)
|
||||
case "$cmd" in
|
||||
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
|
||||
*curl*) printf '%s' "$TANKO_HTTP" ;;
|
||||
esac ;;
|
||||
192.168.68.14)
|
||||
case "$cmd" in
|
||||
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "${AZ_A2A_EXIT:-0}" ;;
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
# ── Stub curl: serves Zulip server, Abiba health, and C3 public URL ───
|
||||
|
||||
CURL_STUB = r"""#!/usr/bin/env bash
|
||||
# Stub curl: serve the Abiba health fixture, the Zulip server 200, and the
|
||||
# C3 public URL probe (https://kagentz.sysloggh.net/). Record every call.
|
||||
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
|
||||
case "$*" in
|
||||
*:9200/health*)
|
||||
case " $* " in
|
||||
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
|
||||
*) printf '%s' "$PI_BODY" ;; # body probe
|
||||
esac ;;
|
||||
*server_settings*)
|
||||
printf '%s' "$SERVER_HTTP" ;;
|
||||
*kagentz.sysloggh.net*)
|
||||
printf '%s' "$KAGENTZ_PUBLIC_CODE" ;;
|
||||
esac
|
||||
exit 0
|
||||
"""
|
||||
|
||||
|
||||
def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
path.write_text(body)
|
||||
path.chmod(path.stat().st_mode
|
||||
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a_code="401", az_a2a_exit=0,
|
||||
kagentz_public_code="302"):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
Everything else — legs, labels, notify logic — is the shipped script.
|
||||
"""
|
||||
sandbox = tmp_path / "sandbox"
|
||||
bindir = sandbox / "bin"
|
||||
record = sandbox / "record"
|
||||
bindir.mkdir(parents=True)
|
||||
record.mkdir()
|
||||
|
||||
_write_exec(bindir / "ssh", SSH_STUB)
|
||||
_write_exec(bindir / "curl", CURL_STUB)
|
||||
|
||||
source = ZULIP_MONITOR.read_text()
|
||||
log_line = 'LOG="/root/zulip-health-monitor.log"'
|
||||
assert log_line in source, "LOG constant moved — update the sandbox harness"
|
||||
log_path = sandbox / "zulip-health-monitor.log"
|
||||
script = sandbox / "zulip-monitor.sh"
|
||||
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
|
||||
|
||||
env = dict(os.environ)
|
||||
env.update({
|
||||
"PATH": f"{bindir}:{env['PATH']}",
|
||||
"RECORD_DIR": str(record),
|
||||
"TANKO_SVC": tanko_svc,
|
||||
"TANKO_HTTP": tanko_http,
|
||||
"AZ_A2A_CODE": az_a2a_code,
|
||||
"AZ_A2A_EXIT": str(az_a2a_exit),
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
"KAGENTZ_PUBLIC_CODE": kagentz_public_code,
|
||||
})
|
||||
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
|
||||
capture_output=True, text=True)
|
||||
return proc, record, log_path
|
||||
|
||||
|
||||
# ── Required behavioural cases ─────────────────────────────────────────
|
||||
|
||||
def test_c3_502_is_incident(tmp_path):
|
||||
"""C3 public leg returns 502 → the run's verdict is an INCIDENT,
|
||||
the C3 line names the 502, and the run is not summarised as healthy."""
|
||||
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="502")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C3 line names the 502.
|
||||
assert "kagentz C3: ❌ public URL 502 (upstream refused)" in log
|
||||
# The verdict is an INCIDENT, not healthy.
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
# The run is not summarised as healthy.
|
||||
assert "✅ 0 issues" not in log
|
||||
|
||||
|
||||
def test_c3_000_is_incident(tmp_path):
|
||||
"""C3 public leg returns 000 → INCIDENT."""
|
||||
proc, record, log_path = _run_monitor(tmp_path, kagentz_public_code="000")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C3 line reports the connection failure.
|
||||
assert "kagentz C3: ❌ public URL down (HTTP 000)" in log
|
||||
# The verdict is an INCIDENT.
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
assert "✅ 0 issues" not in log
|
||||
|
||||
|
||||
def test_healthy_control_c1_401_c3_302(tmp_path):
|
||||
"""Healthy control: C1 401 plus C3 302 → 0 issues and a healthy verdict,
|
||||
proving the new leg cannot cry wolf."""
|
||||
proc, record, log_path = _run_monitor(tmp_path,
|
||||
az_a2a_code="401",
|
||||
kagentz_public_code="302")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# Both legs report alive.
|
||||
assert "kagentz C1: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "kagentz C3: ✅ public URL alive (HTTP 302)" in log
|
||||
# Zero issues, healthy verdict.
|
||||
assert "Result: ✅ 0 issues (all healthy)" in log
|
||||
# No INCIDENT.
|
||||
assert "INCIDENT" not in log
|
||||
# No notify fired for kagentz.
|
||||
assert "kagentz public URL" not in proc.stdout
|
||||
assert "kagentz A2A server" not in proc.stdout
|
||||
|
||||
|
||||
def test_c1_000_is_incident(tmp_path):
|
||||
"""C1 returns 000 → INCIDENT, keeping the leg that actually caught this
|
||||
outage covered behaviourally."""
|
||||
proc, record, log_path = _run_monitor(tmp_path,
|
||||
az_a2a_code="000",
|
||||
az_a2a_exit=7,
|
||||
kagentz_public_code="302")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
# The C1 line reports the A2A down.
|
||||
assert "kagentz C1: ❌ A2A down (HTTP 000)" in log
|
||||
# The verdict is an INCIDENT (even though C3 is healthy).
|
||||
assert "Result: 🔴 INCIDENT" in log
|
||||
assert "all healthy" not in log
|
||||
# The notify fired for the A2A down.
|
||||
assert "kagentz A2A server DOWN" in proc.stdout
|
||||
+42
-19
@@ -1,9 +1,9 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-health
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||
version: 3.3.0
|
||||
version: 3.4.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
report_only_agents:
|
||||
@@ -30,7 +30,7 @@ session start.
|
||||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||||
- **PM2** on localhost for pi process management
|
||||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||||
- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200`
|
||||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||||
- **Relay access** via RA-H OS MCP for alert delivery
|
||||
|
||||
@@ -129,10 +129,10 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
|
||||
|
||||
### Liveness rule (scoped)
|
||||
|
||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN.
|
||||
|
||||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above).
|
||||
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
||||
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
||||
|
||||
@@ -469,9 +469,10 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
|
||||
> monitor issued a restart for something that could not start, posting a false
|
||||
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
|
||||
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
|
||||
> liveness only, and a probe must never restart a platform.
|
||||
> liveness/response and public-path access only, and a probe must never restart
|
||||
> a platform.
|
||||
|
||||
**C1: A2A Server Health**
|
||||
**C1: A2A Server Health (no credential needed)**
|
||||
|
||||
```bash
|
||||
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
|
||||
@@ -479,9 +480,9 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
|
||||
```
|
||||
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down (INCIDENT). Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**C2: A2A Response Verification**
|
||||
**C2: A2A Response Verification (requires LITELLM_KEY)**
|
||||
|
||||
```bash
|
||||
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
|
||||
@@ -491,15 +492,30 @@ ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:
|
||||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||
```
|
||||
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
|
||||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set (this is a credential issue, NOT a server-down incident).
|
||||
|
||||
**C3: Public Access Path (no credential needed)**
|
||||
|
||||
```bash
|
||||
# Probes the public URL that NetBird proxies to the agent-zero container.
|
||||
# This is the captain's point of view: if the captain can't reach it, it's down.
|
||||
# 200/302/401 = alive, 502 = proxy's "upstream refused" page (INCIDENT),
|
||||
# connection failed (000) = INCIDENT. Never restarts anything.
|
||||
curl -s -o /dev/null --connect-timeout 10 --max-time 15 -w '%{http_code}' https://kagentz.sysloggh.net/
|
||||
```
|
||||
|
||||
Expected: `200` (Agent Zero login page), `302` (redirect), or `401` (auth-gated) = alive. `502` = NetBird proxy's "upstream refused" page (the container's port 80 is not listening) = INCIDENT. `000` (connection failed) = INCIDENT. Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**Platform C Actions**
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
|
||||
| C1 A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||
| C1 A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||
| C2 A2A returns 401 | Check `LITELLM_KEY` is set (credential issue, NOT server-down) |
|
||||
| C3 public URL returns `502` (upstream refused) | Alert only — the container's port 80 is not listening; investigate the agent-zero container |
|
||||
| C3 public URL returns `000` (connection failed) | Alert only — the public path is down; investigate the NetBird proxy or the container |
|
||||
| C3 public URL returns a status other than `200`/`302`/`401` | Log/report as a warning — reported, never healed on |
|
||||
|
||||
### Step 5: Global Checks
|
||||
|
||||
@@ -513,12 +529,19 @@ If any bot processes >50 bot-originated messages in 15min → warning.
|
||||
|
||||
### Step 6: Compile and Report
|
||||
|
||||
1. Compile all platform checks and severity
|
||||
2. Determine `overall_severity` from worst per-agent severity
|
||||
3. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||
4. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||
5. If any agent critical or >2 degraded: send relay message to user
|
||||
6. Update `last_check` timestamp in `### Maintains` snapshot
|
||||
1. Run `scripts/zulip-monitor.sh` and take its final `Result:` line as the
|
||||
authoritative run verdict. The verdict line is either
|
||||
`Result: ✅ 0 issues (all healthy)` or
|
||||
`Result: 🔴 INCIDENT — N issue(s) found`.
|
||||
2. Quote that `Result:` line verbatim in the status report. When it says
|
||||
`INCIDENT`, the run MUST be reported as an incident — never summarised as
|
||||
OK/healthy and never annotated as "expected".
|
||||
3. Compile all platform checks and severity
|
||||
4. Determine `overall_severity` from worst per-agent severity
|
||||
5. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||||
6. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||||
7. If any agent critical or >2 degraded: send relay message to user
|
||||
8. Update `last_check` timestamp in `### Maintains` snapshot
|
||||
|
||||
### Restart Debounce
|
||||
|
||||
|
||||
Reference in New Issue
Block a user