The failure line now preserves both attempts' failure kinds instead of
hardcoding 'timeout after retry, 45s'. If both attempts fail, the report
shows: 'probe-failed: <model> <first kind> then <retry kind> (2 attempts)'.
This fixes the self-contradictory output when the first attempt timed out
but the retry failed with connection refused, and prevents the duration
from appearing twice when both attempts were timeouts.
Example outputs:
- timeout then timeout: 'probe-failed: gpu-dense timeout after 30s then timeout after 45s (2 attempts)'
- timeout then refused: 'probe-failed: gpu-dense timeout after 30s then connection refused (2 attempts)'
- refused then refused: 'probe-failed: gpu-dense connection refused then connection refused (2 attempts)'
Single-host models (gpu-dense, gpu-vision, strix-moe) now retry once at
45s on initial 30s timeout failure before declaring probe-failed. This
prevents a single transient timeout (cold prefill ~13s or concurrent
generation hold) from failing the entire health digest.
Evidence: 2026-09-19 ~06:55Z digest failed gpu-dense at 30s; 06:56Z
direct probe 200 in 1.04s.
The failed-probe-fails-the-run property is preserved: if both attempts
fail, the script still exits non-zero with the target and duration named.
Closes: daily-health-digest false negative on single transient timeout
The leg comments were wrong:
- Line 207: Docker Stats showed :9323 (dockerd port) but should be :9324
- Line 215: PVE Exporter showed :9324 (docker-stats port) but should be :9221
These were the exact pairing this PR exists to correct.
Read back the changed lines to verify:
scripts/infra-monitoring.sh:207 shows Docker Stats (CT 116 :9324, 127.0.0.1 via SSH)
scripts/infra-monitoring.sh:215 shows PVE Exporter (CT 116 :9221, 127.0.0.1 via SSH)
Branch: fix/infra-monitoring-probe-ports-20260919
F1: Fixed leg comments to match the actual ports
- Line 207: Docker Stats now shows :9324 (was :9323)
- Line 215: PVE Exporter now shows :9221 (was :9324)
These were the exact pairing this PR exists to correct.
F2: Added per-leg assertions that prove which leg owns which port
The new assertions verify:
1. Docker Stats leg uses $DOCKER_STATS_PORT constant
2. PVE Exporter leg uses $PVE_EXPORTER_PORT constant
3. DOCKER_STATS_PORT constant is set to 9324
4. PVE_EXPORTER_PORT constant is set to 9221
Proof the new assertions bite:
Under the both-constants-swapped mutation (DOCKER_STATS_PORT=9221,
PVE_EXPORTER_PORT=9324), the suite fails with 25 passed / 2 failed
(failing exactly the two constant-value assertions). This proves the
per-leg assertions pin which leg owns which port, not just that both
ports appear somewhere in the SSH log.
Branch: fix/infra-monitoring-probe-ports-20260919
The Liveness Check section now documents ALL SIX verdict shapes exactly
as emitted by proxmox-monitor.sh:
1. ✅ PBS GC: healthy (last run Nh ago, pending-bytes: N B)
2. 🔴 PBS GC: stale (last run Nh ago, pending-bytes: N B)
3. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)
4. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)
5. 🔴 PBS GC: never-run (storepve-datastore not found in GC list)
6. 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)
Fixed the quoted healthy example (line ~114) to include the pending-bytes
suffix the code now appends. Previously the prose only documented the
probe-failed shape, missing the PR's own headline cases (never-run).
Proof: grep -n 'never-run' proxmox-monitor.prose.md now returns two lines
(lines 106 and 108), documenting both never-run variants.
Branch: fix/pbs-gc-liveness-signal-20260919
Previously probed wrong ports:
- Docker Stats was at 9323 (dockerd metrics) but should be 9324
(harness-docker-stats, docker_container_* metrics)
- PVE Exporter was at 9324 (harness-docker-stats) but should be 9221
(harness-pve-exporter, 5 pve_* metrics)
Both exporters bind to 127.0.0.1 on CT 116 and must be probed via SSH.
Updated infrastructure-monitoring.prose.md to document the correct ports.
Added test assertions verifying the exact ports are probed.
Branch: fix/infra-monitoring-probe-ports-20260919
(a) Probe-failure detection: now treats empty OR unparseable JSON as
probe-failed, not never-run. This prevents 'command not found'
outputs from being rendered as service verdicts.
(b) pending-bytes: now extracted from JSON and reported in stale verdict.
(c) Null endtime: use .get() with explicit None check, not 0 fallback.
null values now correctly trigger never-run verdict instead of
arithmetic crash (set -u).
(d) Tests: Added 14-assertion stub-driven suite covering: healthy,
stale (>48h), probe-failed (empty and unparseable), null endtime,
datastore absent. Each test stubs ssh/curl to verify exact
behavior against the pre-fix head.
Branch: fix/pbs-gc-liveness-signal-20260919
Document the PBS GC schedule (00:00 UTC, not 20:00 UTC as PR #116 said),
what actually runs (pbs-gc.sh -> pct exec 107 -- proxmox-backup-manager),
datastore location (CT 107's /mnt/pbs-backup on /tank/pbs-backup, NOT
/media/easystore2), and the new liveness check (48h threshold, reports
age in hours, explicit healthy line).
Branch: fix/pbs-gc-liveness-signal-20260919
Add monitoring leg that checks storepve-datastore GC health:
- Reads GC state from CT 107 via pct exec
- FAILS if last-run-endtime is older than 48h
- Reports age in hours and pending-bytes status
- Uses JSON parsing for reliable data extraction
Test: All 5 legs OK, Exit 0.
Branch: fix/pbs-gc-liveness-signal-20260919
PR #118 finding F1 (low): The deployed LiteLLM on CT 116 uses
allowed_mcp_servers (193 occurrences in installed package), not
bare allowed_mcp. One-word doc fix.
Add the audit-hermes-config.py Rule 15 wording fix from PR #117:
- Violation message now reads 'URL is incorrect: <url> (expected: <expected>)'
- Detection logic unchanged
- Matches URL and not-in-known-list branches remain byte-identical
This consolidates relay #779 into a single PR (#118).
PR #117 follow-up (verify PASS-WITH-FINDINGS):
1. infrastructure-update.prose.md:
- Update access table: agent keys now have per-key MCP grants (2026-09-18)
- Strike-through old limitation: per-key grants now work
- Mark Migration Path as COMPLETED 2026-09-18
2. hermes-config-template.prose.md:
- Remove hedge ('may have been upgraded')
- State fact: per-key MCP grants verified 2026-09-18
This resolves the contradiction where one file asserted per-key
MCP access and the other denied it.
F1: Make kind classification real — append (<kind>) to every failure
line, assign kind=tls on curl exit 35/60, fix :112 where
kind=refused was set on successful retry. Update prose shape.
F3: Move credential placeholder skip from server probe to notify()
only — server is always probed (200 without auth verified live).
F4: notify() logs ALERT SUPPRESSED when credential unusable so
alerts from other legs are not silently dropped.
F7: Restore trailing newline in infra-monitoring.sh.
F5 (DO NOT CHANGE): Verified directly — ssh root@192.168.68.6
'grep -n keep-daily /etc/pve/jobs.cfg' returns five
prune-backups keep-daily=35 lines. Prose is CORRECT.
Test: 18 passed, 0 failed (bash scripts/test_infra_monitoring.sh)
When ZULIP_API_KEY is unset or contains 'placeholder'/'REDACTED', skip
the global Zulip server leg with a ⏭ marker instead of failing the whole
script. The pi/Tanko/kagentz legs do not need the Zulip API key and keep
their verdicts.
Tracked as: zulip-health-credential-placeholder-20260913 (captain-held)
This removes the repeated 'Action required' noise every cycle while
keeping the credential enforcement loud and visible.
Rewrote test to run the monitor with stubbed curl/ssh on PATH that
capture the exact argv of each probe call. The test now asserts the
URL+port of every leg actually requested, not source text or config
constants.
A2: PVE node assertions now check the exact URL in the curl log
(https://192.168.68.9:8006/... must appear), so a wrong IP
(e.g. .99) fails the test.
A3: Grafana/Prometheus/LiteLLM assertions check the URL the call
actually builds, so a hardcoded wrong port in the CALL (while the
config variable stays correct) fails the test.
Mutation evidence:
A2: sed s/192.168.68.9/192.168.68.99/ in PVE_NODES -> suite FAILS
A3: sed s/"$GRAFANA_PORT"/"9999"/ in probe call -> suite FAILS
Results: 18 passed, 0 failed (baseline); 17/18 on each mutation
A1: Test -k assertion now checks use_k:+-k syntax (actual bash pattern)
A2: PVE_NODES assertions now count expected nodes and verify exact array size
A3: Test now asserts liveness behavior (PVE_API_LIVENESS=1) not source text
B1: disk-gc GC schedule corrected: cron runs pbs-gc.sh (not proxmox-backup-manager),
schedule is 20:00 LOCAL (00:00 UTC, not 20:00 UTC), host timezone America/New_York
B2: PROBE SHAPE now documents actual output shape including TLS flag notes
C1: TLS kind is now printed in PVE API failure output
C2: SSH retry logic clarified - retry is in probe_http function (not unreachable)
HOST-RED on media volumes says 'capacity decision — owner to decide' not
'immediate owner attention'. Media volumes are report-only at all levels; the
urgency language was misleading. PBS datastore and host-root get the immediate
attention wording.
Closes the welcome-back proposal: 'how the disk-gc check should classify a
media volume so HOST-RED stops meaning nothing.'
The 20:00 UTC PBS GC cron (proxmox-backup-manager datastore prune)
applies only to /tank/pbs-backup (pbs-datastore). It does NOT touch
media volumes (/media/*) which are report-only at all threat levels.
This clarification prevents the recurring confusion where a 96% media
volume triggers a GC expectation, when the GC schedule never applies
to it.
Closes the 2026-09-17 correction: 'the GC schedule is now 20:00 UTC,
protects the backup datastore, NOT the nearly-full media volume.'
The 2026-09-17 false-verdict incident (third recurrence) showed that prose
policy is not a control: the agent probed :9325/:9405 (nonexistent ports),
CT 116 for PVE API (should be real PVE nodes), and rendered TLS failures as
connection-refused. This moves the canonical probe set into
scripts/infra-monitoring.sh (executed verbatim by the contract) and adds
scripts/test_infra_monitoring.sh which asserts every probed port matches the
documented value.
- scripts/infra-monitoring.sh: one script per contract pattern; all targets,
ports, paths, and expected-status rules in code; -k for PVE self-signed
certs; non-zero exit naming every failed target; no OK summary on failure
- scripts/test_infra_monitoring.sh: 20 assertions covering port drift,
monitoring-host-as-PVE-node, and missing -k flag
- infrastructure-monitoring.prose.md: check-health section now references the
script as executable owner; paste its raw output verbatim
Proof: all 13 legs pass (exit 0); deliberately broken Grafana port (9325)
produces 'probe-failed: 192.168.68.116:9325 (expected 200)' and exit 1.
Implement Rule 15 automated enforcement for MCP servers:
- Validate MCP server URLs against known endpoints (ra-h-os, litellm)
- Check for authentication headers on MCP server configs
- Warn if header values look like env-vars instead of literal keys
- Warn if no auth header is present
This ensures the MCP URL/header invariants from the prose contract
are enforced at the earliest shared boundary (before config
application).
Addressed all 4 ask-user findings from the review:
f1: Qualified the MCP access verification claim - noted that it may
contradict infrastructure-update.prose.md and that LiteLLM version may
have been upgraded since that contract was written.
f2: Added key rotation note documenting that MCP headers use literal keys
and do NOT auto-rotate with the vault. Added TODO to consider adding
MCP header regeneration to the Key Update Procedure.
f3: Added MCP server checks to audit-hermes-config.py (Rule 15):
- Validate MCP server URLs against known endpoints
- Check for authentication headers
- Warn if header values look like env-vars instead of literal keys
f4: Updated Rule 15 verification instruction to include MCP initialize
handshake test, not just /v1/models check.
f5: Added NetBird dependency note documenting that 502 errors on MCP
requests may indicate NetBird outage, not auth failure.
- Documented MCP endpoint verification (2026-08-07): tested with real key,
confirmed initialize handshake works and virtual keys have MCP access
- Added note about Accept header requirement (handled by MCP client library)
- Clarified that the Accept header is NOT part of the config template
- Added litellm MCP server entry to mcp_servers section with correct URL
(https://litellm.sysloggh.net/mcp) and header format
- Updated Rule 15 to be more specific about endpoint validation and
header requirements (REAL keys, not env-var references)
- Added MCP Server Configuration section with implementation details
- Documented the 2026-08-07 Tanko incident where ra-h-os was pointing
to litellm endpoint with env header causing 401 floods
- Updated frontmatter to reflect the changes
Fixes: #keyless-mcp-incident-20260807
Refs: Rule 15 (MCP Endpoint and Header Validation)
Scripts that had hardcoded credentials:
- scripts/zulip-monitor.sh:12 (was ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8")
- scripts/daily-infra-report.py:25 (was ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8")
Both now read from environment variable ZULIP_API_KEY (set by vault-backed start script)
with loud failure if not present.
Other credentials in scripts/:
- capture-dsh-token.sh: uses TOKEN variable with fallback (not a secret)
- pm2-self-heal.sh: reads TELEGRAM_BOT_TOKEN from /root/.pi/agent/extensions/telegram/.env (acceptable)
- prose-ai-review.sh: uses GITEA_TOKEN from .env file with LITELLM_KEY fallback (not secrets)
No other hardcoded credentials found.
Proof of behavior:
With ZULIP_API_KEY set:
bash scripts/zulip-monitor.sh → Server: HTTP 200 (authenticated)
python3 scripts/daily-infra-report.py --json → Collecting infrastructure data...
Without ZULIP_API_KEY set:
bash scripts/zulip-monitor.sh → "ZULIP_API_KEY not set — refusing to run with no credential"
python3 scripts/daily-infra-report.py --json → "ZULIP_API_KEY not set — refusing to run with no credential"
Cred source: environment variable ZULIP_API_KEY (set by vault-backed start script)
No key rotation (that is a separate decision).
Fix 1: GPU leg degradation is not just SSH probe failure — it also covers
gpu-no-port and gpu-ghost conditions. Rewrite to match check_gpu_ports reality.
Fix 2: Add skipped and partial exemplars for all four legs (LiteLLM keys,
GPU ports, CTs, Vault secrets) so the template covers the rule rather than
only the happy path.
Cosmetic: note that compact form (rtx5070 timeout) is acceptable in summary
line when host is identifiable from context; full probe-failed: <target> <kind>
form required in detail section.
Every report line MUST include one clause per check leg, even when a leg is
skipped or fails. Missing leg must never look the same as healthy leg.
Required legs:
- LiteLLM keys: N/M (names) status
- GPU ports: N/M (rtx3090, rtx5070, strixhalo) status — or SKIPPED (reason)
- CTs: N/M running (names)
- Vault secrets: status
GPU leg is never skipped by configuration; only SSH probe failure causes
degraded status.
The key-count line in litellm-health-check now reports self-describing
output: '18 total (10 on page 1)' instead of bare '18' or '4'. Uses
total_count from the paginated API response and names what was
counted. Previous bare numbers could not reconcile changes between
runs; now a reader sees both the total and the page 1 sample.
1. Host-side fix (tanko 192.168.68.122): Moved 5 config.yaml.bak-* files
from /root/.hermes/ to /root/hermes-config-backups/ so the scanner
pattern no longer matches. Dead credential (sk-b7d99... DEEPSEEK key
from July, 401 against gateway) is preserved in history without
cluttering the scanned tree.
2. Contract text: Added ACCEPTABLE PATTERN section to
hermes-key-enforcement.prose.md clarifying that agent keys live in
.env/.env.vault with 600 perms (koonimo's shape), while a plaintext
key in config.yaml or any config backup is a violation. Fix procedure:
move the backup file out of the scanned tree, don't delete.
FAIL 1: Cron cadence was */10 * * * * (every 10 min) but the real crontab
on CT 100 is 35 2,6,10,14,18,22 * * * (every 4 hours at :35). Fixed in
frontmatter, body, and Continuity section. Added 4-hour rationale note.
FAIL 2: Added gateway log health to the list of checks (frontmatter +
Strategies section). Added note that script may perform additional
diagnostics beyond the seven contract checks.
The consolidated agent health check contract (wraps scripts/agent-health-check.py v4).
Created during earlier work but never committed — was a stray untracked file in
the execution clone, making the home look dirty to the fleet update path.