The 20:00 UTC PBS GC cron (proxmox-backup-manager datastore prune)
applies only to /tank/pbs-backup (pbs-datastore). It does NOT touch
media volumes (/media/*) which are report-only at all threat levels.
This clarification prevents the recurring confusion where a 96% media
volume triggers a GC expectation, when the GC schedule never applies
to it.
Closes the 2026-09-17 correction: 'the GC schedule is now 20:00 UTC,
protects the backup datastore, NOT the nearly-full media volume.'
The 2026-09-17 false-verdict incident (third recurrence) showed that prose
policy is not a control: the agent probed :9325/:9405 (nonexistent ports),
CT 116 for PVE API (should be real PVE nodes), and rendered TLS failures as
connection-refused. This moves the canonical probe set into
scripts/infra-monitoring.sh (executed verbatim by the contract) and adds
scripts/test_infra_monitoring.sh which asserts every probed port matches the
documented value.
- scripts/infra-monitoring.sh: one script per contract pattern; all targets,
ports, paths, and expected-status rules in code; -k for PVE self-signed
certs; non-zero exit naming every failed target; no OK summary on failure
- scripts/test_infra_monitoring.sh: 20 assertions covering port drift,
monitoring-host-as-PVE-node, and missing -k flag
- infrastructure-monitoring.prose.md: check-health section now references the
script as executable owner; paste its raw output verbatim
Proof: all 13 legs pass (exit 0); deliberately broken Grafana port (9325)
produces 'probe-failed: 192.168.68.116:9325 (expected 200)' and exit 1.
Implement Rule 15 automated enforcement for MCP servers:
- Validate MCP server URLs against known endpoints (ra-h-os, litellm)
- Check for authentication headers on MCP server configs
- Warn if header values look like env-vars instead of literal keys
- Warn if no auth header is present
This ensures the MCP URL/header invariants from the prose contract
are enforced at the earliest shared boundary (before config
application).
Addressed all 4 ask-user findings from the review:
f1: Qualified the MCP access verification claim - noted that it may
contradict infrastructure-update.prose.md and that LiteLLM version may
have been upgraded since that contract was written.
f2: Added key rotation note documenting that MCP headers use literal keys
and do NOT auto-rotate with the vault. Added TODO to consider adding
MCP header regeneration to the Key Update Procedure.
f3: Added MCP server checks to audit-hermes-config.py (Rule 15):
- Validate MCP server URLs against known endpoints
- Check for authentication headers
- Warn if header values look like env-vars instead of literal keys
f4: Updated Rule 15 verification instruction to include MCP initialize
handshake test, not just /v1/models check.
f5: Added NetBird dependency note documenting that 502 errors on MCP
requests may indicate NetBird outage, not auth failure.
- Documented MCP endpoint verification (2026-08-07): tested with real key,
confirmed initialize handshake works and virtual keys have MCP access
- Added note about Accept header requirement (handled by MCP client library)
- Clarified that the Accept header is NOT part of the config template
- Added litellm MCP server entry to mcp_servers section with correct URL
(https://litellm.sysloggh.net/mcp) and header format
- Updated Rule 15 to be more specific about endpoint validation and
header requirements (REAL keys, not env-var references)
- Added MCP Server Configuration section with implementation details
- Documented the 2026-08-07 Tanko incident where ra-h-os was pointing
to litellm endpoint with env header causing 401 floods
- Updated frontmatter to reflect the changes
Fixes: #keyless-mcp-incident-20260807
Refs: Rule 15 (MCP Endpoint and Header Validation)
Scripts that had hardcoded credentials:
- scripts/zulip-monitor.sh:12 (was ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8")
- scripts/daily-infra-report.py:25 (was ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8")
Both now read from environment variable ZULIP_API_KEY (set by vault-backed start script)
with loud failure if not present.
Other credentials in scripts/:
- capture-dsh-token.sh: uses TOKEN variable with fallback (not a secret)
- pm2-self-heal.sh: reads TELEGRAM_BOT_TOKEN from /root/.pi/agent/extensions/telegram/.env (acceptable)
- prose-ai-review.sh: uses GITEA_TOKEN from .env file with LITELLM_KEY fallback (not secrets)
No other hardcoded credentials found.
Proof of behavior:
With ZULIP_API_KEY set:
bash scripts/zulip-monitor.sh → Server: HTTP 200 (authenticated)
python3 scripts/daily-infra-report.py --json → Collecting infrastructure data...
Without ZULIP_API_KEY set:
bash scripts/zulip-monitor.sh → "ZULIP_API_KEY not set — refusing to run with no credential"
python3 scripts/daily-infra-report.py --json → "ZULIP_API_KEY not set — refusing to run with no credential"
Cred source: environment variable ZULIP_API_KEY (set by vault-backed start script)
No key rotation (that is a separate decision).
Fix 1: GPU leg degradation is not just SSH probe failure — it also covers
gpu-no-port and gpu-ghost conditions. Rewrite to match check_gpu_ports reality.
Fix 2: Add skipped and partial exemplars for all four legs (LiteLLM keys,
GPU ports, CTs, Vault secrets) so the template covers the rule rather than
only the happy path.
Cosmetic: note that compact form (rtx5070 timeout) is acceptable in summary
line when host is identifiable from context; full probe-failed: <target> <kind>
form required in detail section.
Every report line MUST include one clause per check leg, even when a leg is
skipped or fails. Missing leg must never look the same as healthy leg.
Required legs:
- LiteLLM keys: N/M (names) status
- GPU ports: N/M (rtx3090, rtx5070, strixhalo) status — or SKIPPED (reason)
- CTs: N/M running (names)
- Vault secrets: status
GPU leg is never skipped by configuration; only SSH probe failure causes
degraded status.
The key-count line in litellm-health-check now reports self-describing
output: '18 total (10 on page 1)' instead of bare '18' or '4'. Uses
total_count from the paginated API response and names what was
counted. Previous bare numbers could not reconcile changes between
runs; now a reader sees both the total and the page 1 sample.
1. Host-side fix (tanko 192.168.68.122): Moved 5 config.yaml.bak-* files
from /root/.hermes/ to /root/hermes-config-backups/ so the scanner
pattern no longer matches. Dead credential (sk-b7d99... DEEPSEEK key
from July, 401 against gateway) is preserved in history without
cluttering the scanned tree.
2. Contract text: Added ACCEPTABLE PATTERN section to
hermes-key-enforcement.prose.md clarifying that agent keys live in
.env/.env.vault with 600 perms (koonimo's shape), while a plaintext
key in config.yaml or any config backup is a violation. Fix procedure:
move the backup file out of the scanned tree, don't delete.
FAIL 1: Cron cadence was */10 * * * * (every 10 min) but the real crontab
on CT 100 is 35 2,6,10,14,18,22 * * * (every 4 hours at :35). Fixed in
frontmatter, body, and Continuity section. Added 4-hour rationale note.
FAIL 2: Added gateway log health to the list of checks (frontmatter +
Strategies section). Added note that script may perform additional
diagnostics beyond the seven contract checks.
The consolidated agent health check contract (wraps scripts/agent-health-check.py v4).
Created during earlier work but never committed — was a stray untracked file in
the execution clone, making the home look dirty to the fleet update path.
The state file is runtime state (rewritten every scan), so tracking it in git
means:
- every executor's clone becomes permanently dirty after one run
- a scan in one clone produces a merge conflict with a scan in another
- the committed baseline can be stale in a way nobody notices
Added to .gitignore and removed from the index. Contract updated to say
'the state file lives at <abs path> and is gitignored runtime state - the
scanner creates it on first run'.
The contract said the state file was written after every scan, but the script
had no state-file logic at all. This PR adds:
1. Host filesystem scanning (probe_host_filesystems) - probes df on all PVE nodes
2. State file I/O (read_state_file/write_state_file) - absolute path from script location
3. Band classification (classify_band) - HOST-WARN/AMBER/RED thresholds
4. Transition detection (detect_transitions) - alerts on escalation/recovery
5. CLI flags (--hosts-only, --guests-only) to control which parts run
The contract now specifies the state file path resolves from the script's own
location (not CWD-relative), so two different execution contexts cannot write
to two different places.
A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER,
AMBER->RED) and ONCE when it drops back down (recovery notice). While it stays
in the same band, it is reported in the scan output only — no DM, no channel
alert. This stops the same 96% easystore2 from re-DMing the owner on every 6h
scan.
State lives in a small JSON state file (state/host-disk-bands.json), keyed by
host/volume -> last-seen band. The scanner reads the prior band, compares to the
current band, and DMs only on a transition; the state file is written after
every scan. Chosen over a periodic digest because the scan already runs every
6h and a transition is genuinely new, actionable state.
First-run behavior: when the state file does not yet exist, the current band of
every volume is recorded as baseline WITHOUT alerting — a first run would
otherwise DM every already-elevated volume at once.
Report-only restriction and volume-naming output kept exactly as-is.
Add separate threat bands for HOST filesystems (distinct from guest bands):
- HOST-WARN at 85%: name volume + % + absolute free space in scan output
- HOST-AMBER at 90%: flag for owner attention, Zulip DM
- HOST-RED at 95%: flag for immediate owner attention, Zulip DM + channel alert
Volume naming rule: every host line MUST name the volume and what lives on it.
Action classes by volume type:
- host-root: near full = real risk (backup staging, thin-pool metadata)
- media (/media/*): near full = capacity decision for owner, never auto-delete
- pbs-datastore (tank): near full = breaks Proxmox Backup Server
Report-only restriction: no automatic deletion of media or datastore content ever.
Justification (measured 2026-09-15): storepve /media/easystore2 at 96% was
reported but never banded or acted on. Two incidents this weekend showed the
host filesystem is the thing that breaks, not the guest's.
Added HOST-WARN/AMBER/RED alert templates.
Added report-only execution rule for host filesystems.
pm2-self-heal.prose.md:
- Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2
- gpu-watchdog is decommissioned and folded into gpu-monitor.service
- gitea-runner is KEPT; abiba-zulip is KEPT (online for days)
- spoton-service was deleted; live PM2 set is 4 processes
- Preserve historical context for crash-loop guard
litellm-health.prose.md:
- Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100)
- Note .4:9100 is DEAD target (no route, down for weeks)
- Clarify this does not read as 6 healthy nodes
The pr-pipeline workflow filtered both push and pull_request on
paths (**.prose.md, scripts/**.sh, **.yaml, **.yml). A PR whose diff
touched none of those paths — e.g. PR #77, deliverables/-only —
produced no Gitea Actions run at all, so validate/lint/ai-review and
the merge gate were silently skipped.
Remove the paths filter from both triggers and record in the file
that the trigger is intentionally unfiltered. No job, step, needs,
if or command is changed.
Add warning that bare PVE hostnames (acerpve, amdpve, etc.) resolve to VPS
via *.dns.sysloggh.net wildcard, not to actual nodes. List IP addresses:
- acerpve 192.168.68.9
- amdpve 192.168.68.15
- storepve 192.168.68.6
- minipve 192.168.68.12
- ocupve 192.168.68.5
Update acerpve example in backup preflight to include address (192.168.68.9).
Fix dmsetup comment to show full field order:
=start =length =thin-pool =transaction-id
=metadata_used/metadata_total =data_used/data_total
remaining fields are flags
Make it clear lvs command is the primary source for percentages, dmsetup is only for error-state check.
Incidents now include addresses: acerpve (192.168.68.9) and amdpve (192.168.68.15).
Fix two errors in the PREFLIGHT section (measured on acerpve 2026-09-15):
1. lvs -o ... pve/data (not pve-data-tpool) - this is the PRIMARY check that yields percentages directly
- Quote the acerpve example: data 29.95% 1.22% <816.21g
2. dmsetup status pve-data-tpool - document fields correctly:
- = transaction ID (99), NOT data_percent
- = metadata used/total blocks
- = data used/total sectors
- Show how to derive percentages if needed
3. Keep the error-state check (grep -q 'Error|Fail') - this is how the incident presented
Everything else stays: 1777 tmpdir requirement with EACCES symptom, incidents as rationale,
GPU-host fact, --output-format json rule, honest note that metadata/snapshot pressure is unproven.
Add documented preflight checks for VM/CT backups on LVM thin-pool hosts:
- dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70%
- Exit 1 if pool shows Error/Fail state (takes down entire VG including host root)
- tmpdir must be mode 1777 (world-traversable) for vzdump archive step
- --output-format json for tasks started from truncating shells
- Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14)
- Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control
- Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup
- get_response_body() returns first 200 chars of response body (single line)
- On 401/403 model probe: report code + body + key_alias
- Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116)
- Failed connections stay probe-failed, 200 stays plain 200
- Do not turn other statuses into credential faults
Signed-off-by: Abiba
- Replace Infisical retrieval path with proven docker exec + .env note
- State explicitly that master key is NOT in Infisical project=infrastructure
- Keep the live-key check and never-trust-a-literal instruction
- All other corrections from PR #95 preserved
Signed-off-by: Abiba
- probe_http now returns (code, failure_kind) tuple
- Model probes report 'probe-failed: <model> <kind> (Ns timeout)' on 000
- Do not assert a service verdict from a failed probe
- 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill)
- 60s timeout for syslog-auto pool alias with retry on 000
Signed-off-by: Abiba
- Replace broken --plain flag (prints nothing on CLI 0.43.110) with awk parsing
- Note that --plain is broken so nobody fixes it back
- Replace unproven nginx path with verified direct endpoint http://127.0.0.1:4000/key/list
Signed-off-by: Abiba
- hermes-key-enforcement.prose.md:
- State that expiry must be set EXPLICITLY at creation with duration
- Record that config default is NOT honoured by LiteLLM 1.99.1
- Describe daily audit as AUDIT-ONLY (reports non-expiring and soon-to-expire)
- State that renewal is NOT implemented
- Document exclusions: abiba-pi and all crewmate keys stay WITHOUT expiry
- koby is report-only
- litellm-api-keys.prose.md:
- Replace literal master key with retrieval path (docker exec + infisical)
- State that literal values must never be trusted again (key rotates)
- Add live-key check (200 from /key/list)
Signed-off-by: Abiba
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output
Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind
Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running