--- kind: function name: daily-health-digest description: > Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails it as an HTML report. Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to the ops lane, which executes the producer below. Until 2026-09-25 this ran with NO contract file at all, which is why the choice of execution copy was silently the operator's rather than the contract's. EXECUTION IS PINNED. The producer must be run from the clone named under "Execution pinning" — not from an agent working copy. Exit-code semantics (as they actually behave, verified 2026-09-25): * missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert * missing EMAIL credential -> deliberate DEGRADED leg, exit 0, report still produced * email send failure -> exit 1 (a delivery fault, not a code defect) version: 1.0.0 --- ## Purpose Give one daily, machine-collected picture of the estate so drift and outages are seen the day they happen rather than when something breaks. It is a *report*, not a repair: it changes nothing. ## Execution pinning **Pinned execution path:** ``` /root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py ``` **Pinned clone:** `/root/abiba-workspace/projects/prose-contracts` That is the cron's `FM_HOME` clone and the only stable, non-ephemeral copy. The treehouse clone (`/root/.treehouse/agent-workspace-*/…/projects/prose-contracts`) is a **per-agent working copy and must NOT be pinned or executed from** — it drifts onto feature branches, which is exactly how a stale producer reported a stale picture and nobody noticed. See `docs/contract-execution-pinning.md`. Schedule and alerting live in `/etc/cron.d/contract-runner` on CT 100: ``` 30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1 ``` Invocation (credentials come from the vault; never inline them): ```bash cd /root/abiba-workspace/projects/prose-contracts infisical run --env=prod -- python3 scripts/daily-infra-report.py ``` ## Output shape Modes: | invocation | effect | | --- | --- | | *(none)* | collect, build the HTML dashboard, email it | | `--test-email` | same but with a `🧪 TEST —` subject prefix | | `--json` | print the collected data as JSON to stdout and **send no email** | `--json` emits a single object with these top-level keys (observed on a live run 2026-09-25): | key | type | meaning | | --- | --- | --- | | `nodes` | object (5) | per-node cpu/ram/disk/uptime/status | | `node_count`, `nodes_online` | int | Proxmox node totals | | `pve_probe_status`, `resources_probe_status` | `ok`\|`unreachable` | probe outcome | | `total_vms`, `running_vms`, `stopped_vms`, `vms_by_node` | — | guest inventory | | `storage`, `nfs` | array | datastore and mount usage | | `litellm` | object | inference checks | | `zulip_ext` | object | Zulip queue/serving state | | `agents` | object | per-agent health | | `docker_vm`, `docker_syslog`, `docker_netbird`, `endpoints` | object/array | Docker hosts and probed endpoints | ## What a healthy run looks like ``` $ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json "nodes_online": 5, "node_count": 5, "pve_probe_status": "ok", "resources_probe_status": "ok", "running_vms": 22, "total_vms": 22 EXIT=0 ``` and in mail mode: ``` Sending email... ✅ All legs fully credentialed 📋 Summary: Proxmox: 5/5 nodes online VMs/CTs: 22/22 running ``` Healthy means: every probe reports `ok`, `nodes_online == node_count`, and the email leg reports a successful send. ## Exit-code semantics — as they actually behave Verified on 2026-09-25 by running each case deliberately. | condition | exit | alert | notes | | --- | --- | --- | --- | | all probes reachable, email sent | 0 | — | healthy | | **missing `PVE_TOKEN`** | **1** | yes | `PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down)`, and `cluster resources unreachable` | | **Proxmox probe unreachable** | **1** | yes | same path as above; `pve_probe_status: unreachable` | | **missing `EMAIL_PASSWORD`** | **0** | no | deliberate **DEGRADED** leg (`credential-missing: EMAIL_PASSWORD`); the report is still produced | | **email send fails** | **1** | yes | e.g. Gmail `534 5.7.9 Application-specific password required` | | degraded legs present (non-email) | 0 | no | logged under `⚠️ Degraded legs` | The distinction is deliberate and must not be flattened: * A **missing PVE token or an unreachable probe is a real failure** — the report would otherwise claim zero nodes and still look successful. That was the 2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1. * A **missing email credential is survivable** — the report is still produced and is still useful. It is a `DEGRADED` leg and exits 0 by design. `PROBE_FAILURES` and `DEGRADED_LEGS` are separate lists for exactly this reason. Do not merge them. ## Email-delivery dependency Delivery is a **credential dependency, not a code path**. The producer authenticates to `smtp.gmail.com:587` as `jtabiri@gmail.com` with `EMAIL_PASSWORD` from the vault and sends to `jerome@sysloggh.com`. * Since that Google account has two-step verification, `EMAIL_PASSWORD` must be a Google **app password**, not the account password. * As of 2026-09-25 delivery is **failing** with `534 5.7.9 Application-specific password required`; the fix is for the captain to generate a fresh app password and place it in Infisical (`infrastructure/production`) as `EMAIL_PASSWORD`. * **A delivery failure is not a code defect.** Investigation of a failed send should start at the credential, not the script. Chasing it as a code bug wastes the effort; verify the credential path first with `--test-email`. * Tracked separately as `daily-digest-mail-transport-20260921`. ## What counts as a failure A run FAILS (exit 1) when the report cannot be trusted or delivered: * any probe is unreachable, so a section would silently be empty; * `PVE_TOKEN` is missing; * the email send fails. A run is DEGRADED (exit 0, report still produced) when a non-load-bearing credential is absent, currently only `EMAIL_PASSWORD`. ## Failure behaviour * Non-zero exit with the alert text above; on the scheduled path the dispatch is a firstmate message, so the ops lane sees it and reports it. * On a probe failure the report must **not** be treated as evidence about the estate — a `0/0` Proxmox section means "could not look", not "nothing there". That reading is why the 2026-09-25 defect went unnoticed. ## Verification ```bash # data path, no email cd /root/abiba-workspace/projects/prose-contracts infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \ | grep -E 'pve_probe_status|node_count|nodes_online' # delivery path infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email ``` Regression tests: `tests/test_daily_infra_report.py` (7 tests). Four of them fail against the pre-fix script, which is what makes them bite. ## Maintains - daily-infra-dashboard: { status: "degraded", reason: "email credential", last_check: timestamp } - pve-probe: { status: "ok|unreachable", last_check: timestamp }