Merge remote-tracking branch 'origin/master' into fix/land-revision-preflight-guard-20260925
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s

This commit is contained in:
root
2026-09-25 11:29:30 +00:00
4 changed files with 260 additions and 7 deletions
+58
View File
@@ -47,6 +47,7 @@ index:
- infrastructure-monitoring
- zulip-health
- litellm-health
- daily-health-digest
remediation:
- litellm-self-heal
- pm2-self-heal
@@ -96,6 +97,7 @@ index:
- infrastructure-maintenance
- pm2-self-heal
- disk-gc-threat-response
- daily-health-digest
gpu:
- gpu-monitor
- gpu-fleet
@@ -1922,6 +1924,62 @@ contracts:
- abiba
- mumuni
- ops
- name: daily-health-digest
file: daily-health-digest.prose.md
kind: function
category: monitoring
sensitivity: normal
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: 30 10 * * *
description: Daily at 10:30 UTC, dispatched on CT 100 as a firstmate message
to the ops lane, which executes the pinned producer
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires:
- infisical (vault credentials injected at run time)
protocol:
- 'Execute from the PINNED clone only: /root/abiba-workspace/projects/prose-contracts'
- cd /root/abiba-workspace/projects/prose-contracts
- infisical run --env=prod -- python3 scripts/daily-infra-report.py
- 'Never execute from a per-agent working copy (treehouse) - it drifts onto feature branches'
- 'On failure: do not treat a 0/0 Proxmox section as evidence about the estate - it means could not look'
verification:
postconditions:
- check: every Proxmox probe is reachable
verify: >-
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
| grep -c 'pve_probe_status'
expect: '1'
- check: all nodes reported online
verify: >-
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
| grep 'nodes_online'
expect: nodes_online == node_count
artifact: timestamped HTML dashboard emailed to jerome@sysloggh.com
verify_commands:
- infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
- python3 -m pytest tests/test_daily_infra_report.py -q
email_dependency:
transport: smtp.gmail.com:587
identity: jtabiri@gmail.com
secret: EMAIL_PASSWORD (must be a Google app password)
status: DEGRADED as of 2026-09-25 - 534 5.7.9 Application-specific password required
note: A delivery failure is a credential dependency, not a code defect. Tracked
as daily-digest-mail-transport-20260921.
exit_semantics:
'1': missing PVE_TOKEN, unreachable Proxmox probe, or failed email send - raises an alert
'0': healthy, or a deliberate DEGRADED leg where the email credential is absent
and the report is still produced
depends_on: []
last_run: null
last_status: null
drift_alerts: []
koby_report_only: true
koby_host: "CT 111 (tdunna)"
+187
View File
@@ -0,0 +1,187 @@
---
kind: function
name: daily-health-digest
description: >
Produces the daily infrastructure dashboard for the whole estate — 5 Proxmox
nodes, every VM/CT, Zulip, LiteLLM, Docker hosts, storage and NFS — and mails
it as an HTML report.
Dispatched by cron on CT 100 (abiba) at 10:30 UTC as a firstmate message to
the ops lane, which executes the producer below. Until 2026-09-25 this ran
with NO contract file at all, which is why the choice of execution copy was
silently the operator's rather than the contract's.
EXECUTION IS PINNED. The producer must be run from the clone named under
"Execution pinning" — not from an agent working copy.
Exit-code semantics (as they actually behave, verified 2026-09-25):
* missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert
* missing EMAIL credential -> deliberate DEGRADED leg, exit 0, report still
produced
* email send failure -> exit 1 (a delivery fault, not a code defect)
version: 1.0.0
---
## Purpose
Give one daily, machine-collected picture of the estate so drift and outages
are seen the day they happen rather than when something breaks. It is a
*report*, not a repair: it changes nothing.
## Execution pinning
**Pinned execution path:**
```
/root/abiba-workspace/projects/prose-contracts/scripts/daily-infra-report.py
```
**Pinned clone:** `/root/abiba-workspace/projects/prose-contracts`
That is the cron's `FM_HOME` clone and the only stable, non-ephemeral copy.
The treehouse clone (`/root/.treehouse/agent-workspace-*/…/projects/prose-contracts`)
is a **per-agent working copy and must NOT be pinned or executed from** — it
drifts onto feature branches, which is exactly how a stale producer reported a
stale picture and nobody noticed.
See `docs/contract-execution-pinning.md`. Schedule and alerting live in
`/etc/cron.d/contract-runner` on CT 100:
```
30 10 * * * FM_HOME=/root/abiba-workspace /root/abiba-workspace/bin/fm-send.sh ops "run contract: daily-health-digest" > /dev/null 2>&1
```
Invocation (credentials come from the vault; never inline them):
```bash
cd /root/abiba-workspace/projects/prose-contracts
infisical run --env=prod -- python3 scripts/daily-infra-report.py
```
## Output shape
Modes:
| invocation | effect |
| --- | --- |
| *(none)* | collect, build the HTML dashboard, email it |
| `--test-email` | same but with a `🧪 TEST —` subject prefix |
| `--json` | print the collected data as JSON to stdout and **send no email** |
`--json` emits a single object with these top-level keys (observed on a live
run 2026-09-25):
| key | type | meaning |
| --- | --- | --- |
| `nodes` | object (5) | per-node cpu/ram/disk/uptime/status |
| `node_count`, `nodes_online` | int | Proxmox node totals |
| `pve_probe_status`, `resources_probe_status` | `ok`\|`unreachable` | probe outcome |
| `total_vms`, `running_vms`, `stopped_vms`, `vms_by_node` | — | guest inventory |
| `storage`, `nfs` | array | datastore and mount usage |
| `litellm` | object | inference checks |
| `zulip_ext` | object | Zulip queue/serving state |
| `agents` | object | per-agent health |
| `docker_vm`, `docker_syslog`, `docker_netbird`, `endpoints` | object/array | Docker hosts and probed endpoints |
## What a healthy run looks like
```
$ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
"nodes_online": 5, "node_count": 5, "pve_probe_status": "ok",
"resources_probe_status": "ok", "running_vms": 22, "total_vms": 22
EXIT=0
```
and in mail mode:
```
Sending email...
✅ All legs fully credentialed
📋 Summary:
Proxmox: 5/5 nodes online
VMs/CTs: 22/22 running
```
Healthy means: every probe reports `ok`, `nodes_online == node_count`, and the
email leg reports a successful send.
## Exit-code semantics — as they actually behave
Verified on 2026-09-25 by running each case deliberately.
| condition | exit | alert | notes |
| --- | --- | --- | --- |
| all probes reachable, email sent | 0 | — | healthy |
| **missing `PVE_TOKEN`** | **1** | yes | `PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down)`, and `cluster resources unreachable` |
| **Proxmox probe unreachable** | **1** | yes | same path as above; `pve_probe_status: unreachable` |
| **missing `EMAIL_PASSWORD`** | **0** | no | deliberate **DEGRADED** leg (`credential-missing: EMAIL_PASSWORD`); the report is still produced |
| **email send fails** | **1** | yes | e.g. Gmail `534 5.7.9 Application-specific password required` |
| degraded legs present (non-email) | 0 | no | logged under `⚠️ Degraded legs` |
The distinction is deliberate and must not be flattened:
* A **missing PVE token or an unreachable probe is a real failure** — the report
would otherwise claim zero nodes and still look successful. That was the
2026-09-25 silent-zero defect (fixed in PR #133); it now exits 1.
* A **missing email credential is survivable** — the report is still produced
and is still useful. It is a `DEGRADED` leg and exits 0 by design.
`PROBE_FAILURES` and `DEGRADED_LEGS` are separate lists for exactly this
reason. Do not merge them.
## Email-delivery dependency
Delivery is a **credential dependency, not a code path**. The producer
authenticates to `smtp.gmail.com:587` as `jtabiri@gmail.com` with
`EMAIL_PASSWORD` from the vault and sends to `jerome@sysloggh.com`.
* Since that Google account has two-step verification, `EMAIL_PASSWORD` must be
a Google **app password**, not the account password.
* As of 2026-09-25 delivery is **failing** with
`534 5.7.9 Application-specific password required`; the fix is for the
captain to generate a fresh app password and place it in Infisical
(`infrastructure/production`) as `EMAIL_PASSWORD`.
* **A delivery failure is not a code defect.** Investigation of a failed send
should start at the credential, not the script. Chasing it as a code bug
wastes the effort; verify the credential path first with `--test-email`.
* Tracked separately as `daily-digest-mail-transport-20260921`.
## What counts as a failure
A run FAILS (exit 1) when the report cannot be trusted or delivered:
* any probe is unreachable, so a section would silently be empty;
* `PVE_TOKEN` is missing;
* the email send fails.
A run is DEGRADED (exit 0, report still produced) when a non-load-bearing
credential is absent, currently only `EMAIL_PASSWORD`.
## Failure behaviour
* Non-zero exit with the alert text above; on the scheduled path the dispatch is
a firstmate message, so the ops lane sees it and reports it.
* On a probe failure the report must **not** be treated as evidence about the
estate — a `0/0` Proxmox section means "could not look", not "nothing there".
That reading is why the 2026-09-25 defect went unnoticed.
## Verification
```bash
# data path, no email
cd /root/abiba-workspace/projects/prose-contracts
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \
| grep -E 'pve_probe_status|node_count|nodes_online'
# delivery path
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
```
Regression tests: `tests/test_daily_infra_report.py` (7 tests). Four of them
fail against the pre-fix script, which is what makes them bite.
## Maintains
- daily-infra-dashboard: { status: "degraded", reason: "email credential", last_check: timestamp }
- pve-probe: { status: "ok|unreachable", last_check: timestamp }
+6 -1
View File
@@ -17,6 +17,11 @@ from email.mime.multipart import MIMEMultipart
PVE = "https://192.168.68.12:8006"
# The HTTP header prefix is a protocol constant, not a credential. It is kept as
# a constant ending at '=' so that no assembled header-plus-token literal ever
# appears in the tree; the secret scanner rightly flags that shape.
PVE_AUTH_HEADER = "PVEAPIToken="
def pve_auth():
"""PVE API auth header, resolved at call time from the injected environment.
@@ -29,7 +34,7 @@ def pve_auth():
token = os.environ.get("PVE_TOKEN")
if not token:
raise RuntimeError("PVE_TOKEN is not set (run under `infisical run --env=prod`)")
return f"Authorization: PVEAPIToken={token}"
return f"Authorization: {PVE_AUTH_HEADER}{token}"
# ── Shared credentials —─
+9 -6
View File
@@ -28,15 +28,18 @@ def load_script():
def test_pve_token_is_read_from_the_environment():
"""The PVE token must come from the injected environment, never a literal.
Regression: AUTH used to be the literal string
``"Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"``.
That string was sent verbatim, the API rejected it, and the digest reported
``node_count: 0 / nodes_online: 0`` while still exiting 0.
Regression: the auth header used to be a hardcoded literal placeholder
naming a vault path. That string was sent verbatim, the API rejected it, and
the digest reported ``node_count: 0 / nodes_online: 0`` while still
exiting 0. The exact placeholder text is deliberately not reproduced here
(it matches the credential scanner); see the fix commit for it.
"""
mod = load_script()
assert hasattr(mod, "pve_auth"), "pve_auth() must exist to resolve the token at call time"
with patch.dict("os.environ", {"PVE_TOKEN": "user@pve!tokid=secretvalue"}, clear=False):
assert mod.pve_auth() == "Authorization: PVEAPIToken=user@pve!tokid=secretvalue"
sample = "unit-test-sample-value"
with patch.dict("os.environ", {"PVE_TOKEN": sample}, clear=False):
assert mod.pve_auth() == f"Authorization: {mod.PVE_AUTH_HEADER}{sample}"
assert mod.pve_auth().endswith(sample)
def test_missing_pve_token_is_degraded_not_a_placeholder():