PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Backlog row daily-health-digest-contract-missing-20260925. The digest has been dispatched on a schedule with NO contract file at all - no *daily*.prose.md, absent from contract-registry.yaml, the only reference anywhere being the CT100 cron line. That absence is why the choice of execution copy was silently the operator's, and how a stale clone could run the check unnoticed. New daily-health-digest.prose.md states: * the PINNED execution path /root/abiba-workspace/projects/prose-contracts/ scripts/daily-infra-report.py and the pinned clone - the cron's FM_HOME clone, the only stable non-ephemeral copy; the treehouse clone is a per-agent working copy and must NOT be pinned; * the output shape (all 18 top-level --json keys) and what a healthy run is; * the exit-code semantics AS THEY ACTUALLY BEHAVE, verified case by case: missing PVE_TOKEN or an unreachable probe exits 1 and raises an alert, while a missing EMAIL credential is a deliberate DEGRADED leg that still exits 0 and still produces the report. PROBE_FAILURES and DEGRADED_LEGS are separate lists on purpose and must not be merged; * the email-delivery dependency, that EMAIL_PASSWORD must be a Google app password, that it is failing with 534 5.7.9 as of 2026-09-25, and that a delivery failure is a credential dependency rather than a code defect; * what counts as a failure versus degraded. Registered in contract-registry.yaml (contracts entry plus index.by_category .monitoring and index.by_domain.infrastructure). Verified: YAML parses, 31 contracts, exactly one daily-health-digest entry.
188 lines
7.2 KiB
Markdown
188 lines
7.2 KiB
Markdown
---
|
|
kind: function
|
|
name: daily-health-digest
|
|
description: >
|
|
Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox
|
|
nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails
|
|
it as an HTML report.
|
|
|
|
Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to
|
|
the ops lane, which executes the producer below. Until 2026-09-25 this ran
|
|
with NO contract file at all, which is why the choice of execution copy was
|
|
silently the operator's rather than the contract's.
|
|
|
|
EXECUTION IS PINNED. The producer must be run from the clone named under
|
|
"Execution pinning" — not from an agent working copy.
|
|
|
|
Exit-code semantics (as they actually behave, verified 2026-09-25):
|
|
* missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert
|
|
* missing EMAIL credential -> deliberate DEGRADED leg, exit 0, report still
|
|
produced
|
|
* email send failure -> exit 1 (a delivery fault, not a code defect)
|
|
|
|
version: 1.0.0
|
|
---
|
|
|
|
## Purpose
|
|
|
|
Give one daily, machine-collected picture of the estate so drift and outages
|
|
are seen the day they happen rather than when something breaks. It is a
|
|
*report*, not a repair: it changes nothing.
|
|
|
|
## Execution pinning
|
|
|
|
**Pinned execution path:**
|
|
|
|
```
|
|
/root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py
|
|
```
|
|
|
|
**Pinned clone:** `/root/abiba-workspace/projects/prose-contracts`
|
|
|
|
That is the cron's `FM_HOME` clone and the only stable, non-ephemeral copy.
|
|
The treehouse clone (`/root/.treehouse/agent-workspace-*/…/projects/prose-contracts`)
|
|
is a **per-agent working copy and must NOT be pinned or executed from** — it
|
|
drifts onto feature branches, which is exactly how a stale producer reported a
|
|
stale picture and nobody noticed.
|
|
|
|
See `docs/contract-execution-pinning.md`. Schedule and alerting live in
|
|
`/etc/cron.d/contract-runner` on CT 100:
|
|
|
|
```
|
|
30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1
|
|
```
|
|
|
|
Invocation (credentials come from the vault; never inline them):
|
|
|
|
```bash
|
|
cd /root/abiba-workspace/projects/prose-contracts
|
|
infisical run --env=prod -- python3 scripts/daily-infra-report.py
|
|
```
|
|
|
|
## Output shape
|
|
|
|
Modes:
|
|
|
|
| invocation | effect |
|
|
| --- | --- |
|
|
| *(none)* | collect, build the HTML dashboard, email it |
|
|
| `--test-email` | same but with a `🧪 TEST —` subject prefix |
|
|
| `--json` | print the collected data as JSON to stdout and **send no email** |
|
|
|
|
`--json` emits a single object with these top-level keys (observed on a live
|
|
run 2026-09-25):
|
|
|
|
| key | type | meaning |
|
|
| --- | --- | --- |
|
|
| `nodes` | object (5) | per-node cpu/ram/disk/uptime/status |
|
|
| `node_count`, `nodes_online` | int | Proxmox node totals |
|
|
| `pve_probe_status`, `resources_probe_status` | `ok`\|`unreachable` | probe outcome |
|
|
| `total_vms`, `running_vms`, `stopped_vms`, `vms_by_node` | — | guest inventory |
|
|
| `storage`, `nfs` | array | datastore and mount usage |
|
|
| `litellm` | object | inference checks |
|
|
| `zulip_ext` | object | Zulip queue/serving state |
|
|
| `agents` | object | per-agent health |
|
|
| `docker_vm`, `docker_syslog`, `docker_netbird`, `endpoints` | object/array | Docker hosts and probed endpoints |
|
|
|
|
## What a healthy run looks like
|
|
|
|
```
|
|
$ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
|
|
"nodes_online": 5, "node_count": 5, "pve_probe_status": "ok",
|
|
"resources_probe_status": "ok", "running_vms": 22, "total_vms": 22
|
|
EXIT=0
|
|
```
|
|
|
|
and in mail mode:
|
|
|
|
```
|
|
Sending email...
|
|
✅ All legs fully credentialed
|
|
📋 Summary:
|
|
Proxmox: 5/5 nodes online
|
|
VMs/CTs: 22/22 running
|
|
```
|
|
|
|
Healthy means: every probe reports `ok`, `nodes_online == node_count`, and the
|
|
email leg reports a successful send.
|
|
|
|
## Exit-code semantics — as they actually behave
|
|
|
|
Verified on 2026-09-25 by running each case deliberately.
|
|
|
|
| condition | exit | alert | notes |
|
|
| --- | --- | --- | --- |
|
|
| all probes reachable, email sent | 0 | — | healthy |
|
|
| **missing `PVE_TOKEN`** | **1** | yes | `PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down)`, and `cluster resources unreachable` |
|
|
| **Proxmox probe unreachable** | **1** | yes | same path as above; `pve_probe_status: unreachable` |
|
|
| **missing `EMAIL_PASSWORD`** | **0** | no | deliberate **DEGRADED** leg (`credential-missing: EMAIL_PASSWORD`); the report is still produced |
|
|
| **email send fails** | **1** | yes | e.g. Gmail `534 5.7.9 Application-specific password required` |
|
|
| degraded legs present (non-email) | 0 | no | logged under `⚠️ Degraded legs` |
|
|
|
|
The distinction is deliberate and must not be flattened:
|
|
|
|
* A **missing PVE token or an unreachable probe is a real failure** — the report
|
|
would otherwise claim zero nodes and still look successful. That was the
|
|
2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1.
|
|
* A **missing email credential is survivable** — the report is still produced
|
|
and is still useful. It is a `DEGRADED` leg and exits 0 by design.
|
|
|
|
`PROBE_FAILURES` and `DEGRADED_LEGS` are separate lists for exactly this
|
|
reason. Do not merge them.
|
|
|
|
## Email-delivery dependency
|
|
|
|
Delivery is a **credential dependency, not a code path**. The producer
|
|
authenticates to `smtp.gmail.com:587` as `jtabiri@gmail.com` with
|
|
`EMAIL_PASSWORD` from the vault and sends to `jerome@sysloggh.com`.
|
|
|
|
* Since that Google account has two-step verification, `EMAIL_PASSWORD` must be
|
|
a Google **app password**, not the account password.
|
|
* As of 2026-09-25 delivery is **failing** with
|
|
`534 5.7.9 Application-specific password required`; the fix is for the
|
|
captain to generate a fresh app password and place it in Infisical
|
|
(`infrastructure/production`) as `EMAIL_PASSWORD`.
|
|
* **A delivery failure is not a code defect.** Investigation of a failed send
|
|
should start at the credential, not the script. Chasing it as a code bug
|
|
wastes the effort; verify the credential path first with `--test-email`.
|
|
* Tracked separately as `daily-digest-mail-transport-20260921`.
|
|
|
|
## What counts as a failure
|
|
|
|
A run FAILS (exit 1) when the report cannot be trusted or delivered:
|
|
|
|
* any probe is unreachable, so a section would silently be empty;
|
|
* `PVE_TOKEN` is missing;
|
|
* the email send fails.
|
|
|
|
A run is DEGRADED (exit 0, report still produced) when a non-load-bearing
|
|
credential is absent, currently only `EMAIL_PASSWORD`.
|
|
|
|
## Failure behaviour
|
|
|
|
* Non-zero exit with the alert text above; on the scheduled path the dispatch is
|
|
a firstmate message, so the ops lane sees it and reports it.
|
|
* On a probe failure the report must **not** be treated as evidence about the
|
|
estate — a `0/0` Proxmox section means "could not look", not "nothing there".
|
|
That reading is why the 2026-09-25 defect went unnoticed.
|
|
|
|
## Verification
|
|
|
|
```bash
|
|
# data path, no email
|
|
cd /root/abiba-workspace/projects/prose-contracts
|
|
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \
|
|
| grep -E 'pve_probe_status|node_count|nodes_online'
|
|
|
|
# delivery path
|
|
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
|
|
```
|
|
|
|
Regression tests: `tests/test_daily_infra_report.py` (7 tests). Four of them
|
|
fail against the pre-fix script, which is what makes them bite.
|
|
|
|
## Maintains
|
|
|
|
- daily-infra-dashboard: { status: "degraded", reason: "email credential", last_check: timestamp }
|
|
- pve-probe: { status: "ok|unreachable", last_check: timestamp }
|