Files
prose-contracts/daily-health-digest.prose.md
T
root cd00cd0475
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
feat: add the missing daily-health-digest contract and register it
Backlog row daily-health-digest-contract-missing-20260925. The digest has been
dispatched on a schedule with NO contract file at all - no *daily*.prose.md,
absent from contract-registry.yaml, the only reference anywhere being the CT100
cron line. That absence is why the choice of execution copy was silently the
operator's, and how a stale clone could run the check unnoticed.

New daily-health-digest.prose.md states:
* the PINNED execution path /root/abiba-workspace/projects/prose-contracts/
  scripts/daily-infra-report.py and the pinned clone - the cron's FM_HOME clone,
  the only stable non-ephemeral copy; the treehouse clone is a per-agent working
  copy and must NOT be pinned;
* the output shape (all 18 top-level --json keys) and what a healthy run is;
* the exit-code semantics AS THEY ACTUALLY BEHAVE, verified case by case:
  missing PVE_TOKEN or an unreachable probe exits 1 and raises an alert, while
  a missing EMAIL credential is a deliberate DEGRADED leg that still exits 0 and
  still produces the report. PROBE_FAILURES and DEGRADED_LEGS are separate lists
  on purpose and must not be merged;
* the email-delivery dependency, that EMAIL_PASSWORD must be a Google app
  password, that it is failing with 534 5.7.9 as of 2026-09-25, and that a
  delivery failure is a credential dependency rather than a code defect;
* what counts as a failure versus degraded.

Registered in contract-registry.yaml (contracts entry plus index.by_category
.monitoring and index.by_domain.infrastructure). Verified: YAML parses, 31
contracts, exactly one daily-health-digest entry.
2026-09-25 11:08:28 +00:00

7.2 KiB

kind, name, description, version
kind name description version
function daily-health-digest Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails it as an HTML report. Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to the ops lane, which executes the producer below. Until 2026-09-25 this ran with NO contract file at all, which is why the choice of execution copy was silently the operator's rather than the contract's. EXECUTION IS PINNED. The producer must be run from the clone named under "Execution pinning" — not from an agent working copy. Exit-code semantics (as they actually behave, verified 2026-09-25): * missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert * missing EMAIL credential -> deliberate DEGRADED leg, exit 0, report still produced * email send failure -> exit 1 (a delivery fault, not a code defect) 1.0.0

Purpose

Give one daily, machine-collected picture of the estate so drift and outages are seen the day they happen rather than when something breaks. It is a report, not a repair: it changes nothing.

Execution pinning

Pinned execution path:

/root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py

Pinned clone: /root/abiba-workspace/projects/prose-contracts

That is the cron's FM_HOME clone and the only stable, non-ephemeral copy. The treehouse clone (/root/.treehouse/agent-workspace-*/…/projects/prose-contracts) is a per-agent working copy and must NOT be pinned or executed from — it drifts onto feature branches, which is exactly how a stale producer reported a stale picture and nobody noticed.

See docs/contract-execution-pinning.md. Schedule and alerting live in /etc/cron.d/contract-runner on CT 100:

30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1

Invocation (credentials come from the vault; never inline them):

cd /root/abiba-workspace/projects/prose-contracts
infisical run --env=prod -- python3 scripts/daily-infra-report.py

Output shape

Modes:

invocation effect
(none) collect, build the HTML dashboard, email it
--test-email same but with a 🧪 TEST — subject prefix
--json print the collected data as JSON to stdout and send no email

--json emits a single object with these top-level keys (observed on a live run 2026-09-25):

key type meaning
nodes object (5) per-node cpu/ram/disk/uptime/status
node_count, nodes_online int Proxmox node totals
pve_probe_status, resources_probe_status ok|unreachable probe outcome
total_vms, running_vms, stopped_vms, vms_by_node — guest inventory
storage, nfs array datastore and mount usage
litellm object inference checks
zulip_ext object Zulip queue/serving state
agents object per-agent health
docker_vm, docker_syslog, docker_netbird, endpoints object/array Docker hosts and probed endpoints

What a healthy run looks like

$ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
  "nodes_online": 5,  "node_count": 5,  "pve_probe_status": "ok",
  "resources_probe_status": "ok",  "running_vms": 22,  "total_vms": 22
EXIT=0

and in mail mode:

   Sending email...
   ✅ All legs fully credentialed
📋 Summary:
   Proxmox: 5/5 nodes online
   VMs/CTs: 22/22 running

Healthy means: every probe reports ok, nodes_online == node_count, and the email leg reports a successful send.

Exit-code semantics — as they actually behave

Verified on 2026-09-25 by running each case deliberately.

condition exit alert notes
all probes reachable, email sent 0 — healthy
missing PVE_TOKEN 1 yes PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down), and cluster resources unreachable
Proxmox probe unreachable 1 yes same path as above; pve_probe_status: unreachable
missing EMAIL_PASSWORD 0 no deliberate DEGRADED leg (credential-missing: EMAIL_PASSWORD); the report is still produced
email send fails 1 yes e.g. Gmail 534 5.7.9 Application-specific password required
degraded legs present (non-email) 0 no logged under ⚠️ Degraded legs

The distinction is deliberate and must not be flattened:

  • A missing PVE token or an unreachable probe is a real failure — the report would otherwise claim zero nodes and still look successful. That was the 2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1.
  • A missing email credential is survivable — the report is still produced and is still useful. It is a DEGRADED leg and exits 0 by design.

PROBE_FAILURES and DEGRADED_LEGS are separate lists for exactly this reason. Do not merge them.

Email-delivery dependency

Delivery is a credential dependency, not a code path. The producer authenticates to smtp.gmail.com:587 as jtabiri@gmail.com with EMAIL_PASSWORD from the vault and sends to jerome@sysloggh.com.

  • Since that Google account has two-step verification, EMAIL_PASSWORD must be a Google app password, not the account password.
  • As of 2026-09-25 delivery is failing with 534 5.7.9 Application-specific password required; the fix is for the captain to generate a fresh app password and place it in Infisical (infrastructure/production) as EMAIL_PASSWORD.
  • A delivery failure is not a code defect. Investigation of a failed send should start at the credential, not the script. Chasing it as a code bug wastes the effort; verify the credential path first with --test-email.
  • Tracked separately as daily-digest-mail-transport-20260921.

What counts as a failure

A run FAILS (exit 1) when the report cannot be trusted or delivered:

  • any probe is unreachable, so a section would silently be empty;
  • PVE_TOKEN is missing;
  • the email send fails.

A run is DEGRADED (exit 0, report still produced) when a non-load-bearing credential is absent, currently only EMAIL_PASSWORD.

Failure behaviour

  • Non-zero exit with the alert text above; on the scheduled path the dispatch is a firstmate message, so the ops lane sees it and reports it.
  • On a probe failure the report must not be treated as evidence about the estate — a 0/0 Proxmox section means "could not look", not "nothing there". That reading is why the 2026-09-25 defect went unnoticed.

Verification

# data path, no email
cd /root/abiba-workspace/projects/prose-contracts
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \
  | grep -E 'pve_probe_status|node_count|nodes_online'

# delivery path
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email

Regression tests: tests/test_daily_infra_report.py (7 tests). Four of them fail against the pre-fix script, which is what makes them bite.

Maintains

  • daily-infra-dashboard: { status: "degraded", reason: "email credential", last_check: timestamp }
  • pve-probe: { status: "ok|unreachable", last_check: timestamp }