diff --git a/contract-registry.yaml b/contract-registry.yaml index 2e58f7b..eec0d9b 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -47,6 +47,7 @@ index: - infrastructure-monitoring - zulip-health - litellm-health + - daily-health-digest remediation: - litellm-self-heal - pm2-self-heal @@ -96,6 +97,7 @@ index: - infrastructure-maintenance - pm2-self-heal - disk-gc-threat-response + - daily-health-digest gpu: - gpu-monitor - gpu-fleet @@ -1922,6 +1924,62 @@ contracts: - abiba - mumuni - ops +- name: daily-health-digest + file: daily-health-digest.prose.md + kind: function + category: monitoring + sensitivity: normal + status: active + owner: abiba + version: 1.0.0 + trigger: + type: scheduled + cadence: 30 10 * * * + description: Daily at 10:30 UTC, dispatched on CT 100 as a firstmate message + to the ops lane, which executes the pinned producer + cron_job_id: null + execution: + agent: abiba + timeout: 300 + requires: + - infisical (vault credentials injected at run time) + protocol: + - 'Execute from the PINNED clone only: /root/abiba-workspace/projects/prose-contracts' + - cd /root/abiba-workspace/projects/prose-contracts + - infisical run --env=prod -- python3 scripts/daily-infra-report.py + - 'Never execute from a per-agent working copy (treehouse) - it drifts onto feature branches' + - 'On failure: do not treat a 0/0 Proxmox section as evidence about the estate - it means could not look' + verification: + postconditions: + - check: every Proxmox probe is reachable + verify: >- + infisical run --env=prod -- python3 scripts/daily-infra-report.py --json + | grep -c 'pve_probe_status' + expect: '1' + - check: all nodes reported online + verify: >- + infisical run --env=prod -- python3 scripts/daily-infra-report.py --json + | grep 'nodes_online' + expect: nodes_online == node_count + artifact: timestamped HTML dashboard emailed to jerome@sysloggh.com + verify_commands: + - infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email + - python3 -m pytest tests/test_daily_infra_report.py -q + email_dependency: + transport: smtp.gmail.com:587 + identity: jtabiri@gmail.com + secret: EMAIL_PASSWORD (must be a Google app password) + status: DEGRADED as of 2026-09-25 - 534 5.7.9 Application-specific password required + note: A delivery failure is a credential dependency, not a code defect. Tracked + as daily-digest-mail-transport-20260921. + exit_semantics: + '1': missing PVE_TOKEN, unreachable Proxmox probe, or failed email send - raises an alert + '0': healthy, or a deliberate DEGRADED leg where the email credential is absent + and the report is still produced + depends_on: [] + last_run: null + last_status: null + drift_alerts: [] koby_report_only: true koby_host: "CT 111 (tdunna)" diff --git a/daily-health-digest.prose.md b/daily-health-digest.prose.md new file mode 100644 index 0000000..9cf7fdd --- /dev/null +++ b/daily-health-digest.prose.md @@ -0,0 +1,187 @@ +--- +kind: function +name: daily-health-digest +description: > + Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox + nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails + it as an HTML report. + + Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to + the ops lane, which executes the producer below. Until 2026-09-25 this ran + with NO contract file at all, which is why the choice of execution copy was + silently the operator's rather than the contract's. + + EXECUTION IS PINNED. The producer must be run from the clone named under + "Execution pinning" — not from an agent working copy. + + Exit-code semantics (as they actually behave, verified 2026-09-25): + * missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert + * missing EMAIL credential -> deliberate DEGRADED leg, exit 0, report still + produced + * email send failure -> exit 1 (a delivery fault, not a code defect) + +version: 1.0.0 +--- + +## Purpose + +Give one daily, machine-collected picture of the estate so drift and outages +are seen the day they happen rather than when something breaks. It is a +*report*, not a repair: it changes nothing. + +## Execution pinning + +**Pinned execution path:** + +``` +/root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py +``` + +**Pinned clone:** `/root/abiba-workspace/projects/prose-contracts` + +That is the cron's `FM_HOME` clone and the only stable, non-ephemeral copy. +The treehouse clone (`/root/.treehouse/agent-workspace-*/…/projects/prose-contracts`) +is a **per-agent working copy and must NOT be pinned or executed from** — it +drifts onto feature branches, which is exactly how a stale producer reported a +stale picture and nobody noticed. + +See `docs/contract-execution-pinning.md`. Schedule and alerting live in +`/etc/cron.d/contract-runner` on CT 100: + +``` +30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1 +``` + +Invocation (credentials come from the vault; never inline them): + +```bash +cd /root/abiba-workspace/projects/prose-contracts +infisical run --env=prod -- python3 scripts/daily-infra-report.py +``` + +## Output shape + +Modes: + +| invocation | effect | +| --- | --- | +| *(none)* | collect, build the HTML dashboard, email it | +| `--test-email` | same but with a `🧪 TEST —` subject prefix | +| `--json` | print the collected data as JSON to stdout and **send no email** | + +`--json` emits a single object with these top-level keys (observed on a live +run 2026-09-25): + +| key | type | meaning | +| --- | --- | --- | +| `nodes` | object (5) | per-node cpu/ram/disk/uptime/status | +| `node_count`, `nodes_online` | int | Proxmox node totals | +| `pve_probe_status`, `resources_probe_status` | `ok`\|`unreachable` | probe outcome | +| `total_vms`, `running_vms`, `stopped_vms`, `vms_by_node` | — | guest inventory | +| `storage`, `nfs` | array | datastore and mount usage | +| `litellm` | object | inference checks | +| `zulip_ext` | object | Zulip queue/serving state | +| `agents` | object | per-agent health | +| `docker_vm`, `docker_syslog`, `docker_netbird`, `endpoints` | object/array | Docker hosts and probed endpoints | + +## What a healthy run looks like + +``` +$ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json + "nodes_online": 5, "node_count": 5, "pve_probe_status": "ok", + "resources_probe_status": "ok", "running_vms": 22, "total_vms": 22 +EXIT=0 +``` + +and in mail mode: + +``` + Sending email... + ✅ All legs fully credentialed +📋 Summary: + Proxmox: 5/5 nodes online + VMs/CTs: 22/22 running +``` + +Healthy means: every probe reports `ok`, `nodes_online == node_count`, and the +email leg reports a successful send. + +## Exit-code semantics — as they actually behave + +Verified on 2026-09-25 by running each case deliberately. + +| condition | exit | alert | notes | +| --- | --- | --- | --- | +| all probes reachable, email sent | 0 | — | healthy | +| **missing `PVE_TOKEN`** | **1** | yes | `PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down)`, and `cluster resources unreachable` | +| **Proxmox probe unreachable** | **1** | yes | same path as above; `pve_probe_status: unreachable` | +| **missing `EMAIL_PASSWORD`** | **0** | no | deliberate **DEGRADED** leg (`credential-missing: EMAIL_PASSWORD`); the report is still produced | +| **email send fails** | **1** | yes | e.g. Gmail `534 5.7.9 Application-specific password required` | +| degraded legs present (non-email) | 0 | no | logged under `⚠️ Degraded legs` | + +The distinction is deliberate and must not be flattened: + +* A **missing PVE token or an unreachable probe is a real failure** — the report + would otherwise claim zero nodes and still look successful. That was the + 2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1. +* A **missing email credential is survivable** — the report is still produced + and is still useful. It is a `DEGRADED` leg and exits 0 by design. + +`PROBE_FAILURES` and `DEGRADED_LEGS` are separate lists for exactly this +reason. Do not merge them. + +## Email-delivery dependency + +Delivery is a **credential dependency, not a code path**. The producer +authenticates to `smtp.gmail.com:587` as `jtabiri@gmail.com` with +`EMAIL_PASSWORD` from the vault and sends to `jerome@sysloggh.com`. + +* Since that Google account has two-step verification, `EMAIL_PASSWORD` must be + a Google **app password**, not the account password. +* As of 2026-09-25 delivery is **failing** with + `534 5.7.9 Application-specific password required`; the fix is for the + captain to generate a fresh app password and place it in Infisical + (`infrastructure/production`) as `EMAIL_PASSWORD`. +* **A delivery failure is not a code defect.** Investigation of a failed send + should start at the credential, not the script. Chasing it as a code bug + wastes the effort; verify the credential path first with `--test-email`. +* Tracked separately as `daily-digest-mail-transport-20260921`. + +## What counts as a failure + +A run FAILS (exit 1) when the report cannot be trusted or delivered: + +* any probe is unreachable, so a section would silently be empty; +* `PVE_TOKEN` is missing; +* the email send fails. + +A run is DEGRADED (exit 0, report still produced) when a non-load-bearing +credential is absent, currently only `EMAIL_PASSWORD`. + +## Failure behaviour + +* Non-zero exit with the alert text above; on the scheduled path the dispatch is + a firstmate message, so the ops lane sees it and reports it. +* On a probe failure the report must **not** be treated as evidence about the + estate — a `0/0` Proxmox section means "could not look", not "nothing there". + That reading is why the 2026-09-25 defect went unnoticed. + +## Verification + +```bash +# data path, no email +cd /root/abiba-workspace/projects/prose-contracts +infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \ + | grep -E 'pve_probe_status|node_count|nodes_online' + +# delivery path +infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email +``` + +Regression tests: `tests/test_daily_infra_report.py` (7 tests). Four of them +fail against the pre-fix script, which is what makes them bite. + +## Maintains + +- daily-infra-dashboard: { status: "degraded", reason: "email credential", last_check: timestamp } +- pve-probe: { status: "ok|unreachable", last_check: timestamp }