Backlog row daily-health-digest-contract-missing-20260925. The digest has been dispatched on a schedule with NO contract file at all - no *daily*.prose.md, absent from contract-registry.yaml, the only reference anywhere being the CT100 cron line. That absence is why the choice of execution copy was silently the operator's, and how a stale clone could run the check unnoticed. New daily-health-digest.prose.md states: * the PINNED execution path /root/abiba-workspace/projects/prose-contracts/ scripts/daily-infra-report.py and the pinned clone - the cron's FM_HOME clone, the only stable non-ephemeral copy; the treehouse clone is a per-agent working copy and must NOT be pinned; * the output shape (all 18 top-level --json keys) and what a healthy run is; * the exit-code semantics AS THEY ACTUALLY BEHAVE, verified case by case: missing PVE_TOKEN or an unreachable probe exits 1 and raises an alert, while a missing EMAIL credential is a deliberate DEGRADED leg that still exits 0 and still produces the report. PROBE_FAILURES and DEGRADED_LEGS are separate lists on purpose and must not be merged; * the email-delivery dependency, that EMAIL_PASSWORD must be a Google app password, that it is failing with 534 5.7.9 as of 2026-09-25, and that a delivery failure is a credential dependency rather than a code defect; * what counts as a failure versus degraded. Registered in contract-registry.yaml (contracts entry plus index.by_category .monitoring and index.by_domain.infrastructure). Verified: YAML parses, 31 contracts, exactly one daily-health-digest entry.
7.2 KiB
kind, name, description, version
| kind | name | description | version |
|---|---|---|---|
| function | daily-health-digest | Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails it as an HTML report. Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to the ops lane, which executes the producer below. Until 2026-09-25 this ran with NO contract file at all, which is why the choice of execution copy was silently the operator's rather than the contract's. EXECUTION IS PINNED. The producer must be run from the clone named under "Execution pinning" — not from an agent working copy. Exit-code semantics (as they actually behave, verified 2026-09-25): * missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert * missing EMAIL credential -> deliberate DEGRADED leg, exit 0, report still produced * email send failure -> exit 1 (a delivery fault, not a code defect) | 1.0.0 |
Purpose
Give one daily, machine-collected picture of the estate so drift and outages are seen the day they happen rather than when something breaks. It is a report, not a repair: it changes nothing.
Execution pinning
Pinned execution path:
/root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py
Pinned clone: /root/abiba-workspace/projects/prose-contracts
That is the cron's FM_HOME clone and the only stable, non-ephemeral copy.
The treehouse clone (/root/.treehouse/agent-workspace-*/…/projects/prose-contracts)
is a per-agent working copy and must NOT be pinned or executed from — it
drifts onto feature branches, which is exactly how a stale producer reported a
stale picture and nobody noticed.
See docs/contract-execution-pinning.md. Schedule and alerting live in
/etc/cron.d/contract-runner on CT 100:
30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1
Invocation (credentials come from the vault; never inline them):
cd /root/abiba-workspace/projects/prose-contracts
infisical run --env=prod -- python3 scripts/daily-infra-report.py
Output shape
Modes:
| invocation | effect |
|---|---|
| (none) | collect, build the HTML dashboard, email it |
--test-email |
same but with a 🧪 TEST — subject prefix |
--json |
print the collected data as JSON to stdout and send no email |
--json emits a single object with these top-level keys (observed on a live
run 2026-09-25):
| key | type | meaning |
|---|---|---|
nodes |
object (5) | per-node cpu/ram/disk/uptime/status |
node_count, nodes_online |
int | Proxmox node totals |
pve_probe_status, resources_probe_status |
ok|unreachable |
probe outcome |
total_vms, running_vms, stopped_vms, vms_by_node |
— | guest inventory |
storage, nfs |
array | datastore and mount usage |
litellm |
object | inference checks |
zulip_ext |
object | Zulip queue/serving state |
agents |
object | per-agent health |
docker_vm, docker_syslog, docker_netbird, endpoints |
object/array | Docker hosts and probed endpoints |
What a healthy run looks like
$ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
"nodes_online": 5, "node_count": 5, "pve_probe_status": "ok",
"resources_probe_status": "ok", "running_vms": 22, "total_vms": 22
EXIT=0
and in mail mode:
Sending email...
✅ All legs fully credentialed
📋 Summary:
Proxmox: 5/5 nodes online
VMs/CTs: 22/22 running
Healthy means: every probe reports ok, nodes_online == node_count, and the
email leg reports a successful send.
Exit-code semantics — as they actually behave
Verified on 2026-09-25 by running each case deliberately.
| condition | exit | alert | notes |
|---|---|---|---|
| all probes reachable, email sent | 0 | — | healthy |
missing PVE_TOKEN |
1 | yes | PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down), and cluster resources unreachable |
| Proxmox probe unreachable | 1 | yes | same path as above; pve_probe_status: unreachable |
missing EMAIL_PASSWORD |
0 | no | deliberate DEGRADED leg (credential-missing: EMAIL_PASSWORD); the report is still produced |
| email send fails | 1 | yes | e.g. Gmail 534 5.7.9 Application-specific password required |
| degraded legs present (non-email) | 0 | no | logged under ⚠️ Degraded legs |
The distinction is deliberate and must not be flattened:
- A missing PVE token or an unreachable probe is a real failure — the report would otherwise claim zero nodes and still look successful. That was the 2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1.
- A missing email credential is survivable — the report is still produced
and is still useful. It is a
DEGRADEDleg and exits 0 by design.
PROBE_FAILURES and DEGRADED_LEGS are separate lists for exactly this
reason. Do not merge them.
Email-delivery dependency
Delivery is a credential dependency, not a code path. The producer
authenticates to smtp.gmail.com:587 as jtabiri@gmail.com with
EMAIL_PASSWORD from the vault and sends to jerome@sysloggh.com.
- Since that Google account has two-step verification,
EMAIL_PASSWORDmust be a Google app password, not the account password. - As of 2026-09-25 delivery is failing with
534 5.7.9 Application-specific password required; the fix is for the captain to generate a fresh app password and place it in Infisical (infrastructure/production) asEMAIL_PASSWORD. - A delivery failure is not a code defect. Investigation of a failed send
should start at the credential, not the script. Chasing it as a code bug
wastes the effort; verify the credential path first with
--test-email. - Tracked separately as
daily-digest-mail-transport-20260921.
What counts as a failure
A run FAILS (exit 1) when the report cannot be trusted or delivered:
- any probe is unreachable, so a section would silently be empty;
PVE_TOKENis missing;- the email send fails.
A run is DEGRADED (exit 0, report still produced) when a non-load-bearing
credential is absent, currently only EMAIL_PASSWORD.
Failure behaviour
- Non-zero exit with the alert text above; on the scheduled path the dispatch is a firstmate message, so the ops lane sees it and reports it.
- On a probe failure the report must not be treated as evidence about the
estate — a
0/0Proxmox section means "could not look", not "nothing there". That reading is why the 2026-09-25 defect went unnoticed.
Verification
# data path, no email
cd /root/abiba-workspace/projects/prose-contracts
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \
| grep -E 'pve_probe_status|node_count|nodes_online'
# delivery path
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
Regression tests: tests/test_daily_infra_report.py (7 tests). Four of them
fail against the pre-fix script, which is what makes them bite.
Maintains
- daily-infra-dashboard: { status: "degraded", reason: "email credential", last_check: timestamp }
- pve-probe: { status: "ok|unreachable", last_check: timestamp }