Files
prose-contracts/docs/contract-execution-pinning.md
root c5a6dcd42a
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
fix(revision-preflight): default to warn, name the refusal class, bound the fetch
Bounded correction round on PR #134 after a PASS-WITH-FINDINGS review whose
Finding 4 is High. The guard's purpose and its fail-closed fix stand; the
problem was that with 'enforce' as the default it gates EVERY scheduled
contract, and three legitimate states produce a refusal - a clone legitimately
ahead of origin/master mid-review, a detached HEAD, and an offline or failed
fetch - so any of them would turn the fleet's monitoring into withheld
verdicts. That risk outweighs the staleness the guard catches.

1. DEFAULT IS NOW 'warn'. 'enforce' remains available and documented. The
   criteria for flipping the default later are written into the doc as a
   decision with evidence - a sustained window (30 days / 200+ runs) with zero
   mismatch:* and zero cannot-verify:* refusals, no fetch blips, and a pinned
   clone demonstrably kept current - explicitly as its own change, not a silent
   flip.

2. 'COULD NOT CHECK' IS NOW DISTINGUISHABLE FROM 'THIS COPY IS WRONG'. Every
   non-zero exit prints a machine-readable REASON=<class> line:
     cannot-verify:fetch-failed | cannot-verify:ref-unresolvable   (exit 2)
     mismatch:path-absent | mismatch:content
     mismatch:detached-head | mismatch:clone-ahead                (exit 1)
   detached-head and clone-ahead are named separately because they are
   legitimate states, far less alarming than a hand-edited file. clone-ahead
   requires HEAD to be STRICTLY ahead; an uncommitted edit on a commit that IS
   the ref is a plain content mismatch (my own first cut got this wrong and the
   new test 7d caught it).

3. THE DEFAULT FETCH IS BOUNDED: --fetch-timeout, default 20s, 0 = unbounded,
   and a missing 'timeout' binary is itself a cannot-verify rather than an
   unbounded fetch inside a scheduled contract.

4. TEST COVERAGE ADDED for every new class: fetch failure, fetch timeout
   (asserted to return promptly under a 1s bound), unresolvable ref, detached
   HEAD, clone-ahead, genuine content mismatch, and the contract-run.sh default.
   The pre-fix draft fixture comparisons are kept: 31 passed, 0 failed.

5. MERGE-TIME SEQUENCE documented: fast-forward /opt/contract-runner, confirm
   clean, prove a contract runs and reports. Baseline recorded as of today -
   firstmate has already fast-forwarded it to 9faffe4 - with the note that an
   untracked file blocks a fast-forward even when byte-identical.

Live behaviour re-verified on the real runner path:
  default: REASON=mismatch:clone-ahead -> 'continuing because ...=warn' -> VERDICT: PASS, exit 0
  enforce: REASON=mismatch:clone-ahead -> 'VERDICT WITHHELD: mismatch:clone-ahead', exit 2

MANDATORY CHECKS (master went red once from a credential-SHAPED string, so
these are now run on every shippable branch):
  bash scripts/prose-lint.sh        -> LINT PASSED (18 warning(s))
  secret scan                       -> secret scan clean (tree; 34 allowlisted,
                                       24 inert value(s) ignored); No committed credentials
  shellcheck revision-preflight.sh  -> clean
  shellcheck test_revision_preflight.sh -> clean
  shellcheck contract-run.sh        -> SC2034 x1, SC2086 x2 - byte-identical on
                                       master, i.e. pre-existing, none introduced

tests/test_probe_drift.py::test_prose_lint_accepts_report_format_with_provenance
fails both before and after this branch (it runs prose-lint from a temp CWD and
cannot find its sibling secret-scan.sh). Pre-existing, unrelated, not fixed here.
2026-09-25 11:29:15 +00:00

170 lines
7.8 KiB
Markdown

# Contract execution pinning
Which copy of a contract script actually ran, and how that is proven.
## Why this exists
Three times on 2026-09-25 a contract reported a verdict from a copy that was
not the merged one:
1. The ops lane's own clone sat on the merged feature branch
`fix/search-stack-multi-engine-20260925` at `8b2eba4` with no `pve_auth`
fix, while it executed the daily digest from a different clone. Nothing in
the workflow noticed.
2. `scripts/search-stack-check.py` was deployed into the pinned runner clone
by hand rather than through git.
3. A stale local `origin/master` ref made an ancestry check report
"unlanded work" for a branch that had in fact merged — the same staleness
would have passed a stale script as current.
A contract verdict is only meaningful if it came from the merged copy. The
control is `scripts/revision-preflight.sh`.
## The rule
**Every contract pins exactly one clone for execution: the clone that
`scripts/contract-run.sh` itself lives in.**
`contract-run.sh` derives that from its own location (`SCRIPTS_DIR`) and checks
the script it is about to run against `origin/master` in the same clone. There
is no second path to configure, and no contract may be executed from a
hand-copied location.
| Contract | Script | Pinned clone |
| --- | --- | --- |
| `infrastructure-monitoring` | `scripts/infra-monitoring.sh` | the clone containing `contract-run.sh` |
| `proxmox-monitor` | `scripts/proxmox-monitor.sh` | same |
| `zulip-health` | `scripts/zulip-monitor.sh` | same |
| `agent-health-check` | `scripts/agent-health-check.py` | same |
| `litellm-health` | `scripts/litellm-health-check.py` | same |
| `disk-gc-threat-response` | `scripts/disk-gc-scan.py` | same |
| `pm2-self-heal` | `scripts/pm2-self-heal.sh` | same |
| `search-stack-visibility` | `scripts/search-stack-check.py` | same |
### The deployed runner
The scheduler on **CT 100 (abiba)** runs contracts from
**`/opt/contract-runner`** via `/etc/cron.d/contract-runner`. That clone is the
pinned execution copy for every scheduled contract, and it must be kept current
with `master` by fast-forward. Its `origin` is a local path to the upstream
working copy, not a network remote.
`daily-health-digest` is **not** in the table above because it has no contract
file and no mapping — it is dispatched by cron as
`fm-send.sh ops "run contract: daily-health-digest"` and was, until
2026-09-25, executed by hand from whichever clone the operator happened to be
in. Creating its contract file and pinning it to a clone is an open follow-up.
## How the check works
`scripts/revision-preflight.sh <script-path> <clone-path>`:
* resolves the **repo-relative** path of the executing script inside the clone;
* **fetches** the remote first, so a stale local ref cannot make a stale script
look current — bounded by `--fetch-timeout` (default 20s) so a hung remote
cannot block a scheduled contract;
* compares the script's sha256 against `<ref>:<repo-relative-path>`;
* **fails closed** — a path absent from the ref, an unresolvable ref, or a
failed fetch is a failure, never a warning.
### Exit codes and reason classes
The guard distinguishes **"I could not check"** from **"this copy is wrong"**,
and every non-zero exit prints a machine-readable `REASON=<class>` line before
the human text, because a warning nobody can classify is not actionable — and
the flip to `enforce` (below) depends on being able to read these apart.
| exit | `REASON=` | meaning |
| --- | --- | --- |
| 0 | — | verified match |
| 2 | `cannot-verify:fetch-failed` | remote unreachable, failed, or timed out |
| 2 | `cannot-verify:ref-unresolvable` | `<ref>` does not exist in the clone |
| 1 | `mismatch:path-absent` | the script does not exist in `<ref>` |
| 1 | `mismatch:content` | the script differs from `<ref>` |
| 1 | `mismatch:detached-head` | the clone is on a detached HEAD |
| 1 | `mismatch:clone-ahead` | local HEAD is strictly ahead of `<ref>` (mid-review) |
`detached-head` and `clone-ahead` are named separately on purpose: they are
*legitimate* states that merely fail to be "the merged copy", and they are far
less alarming than a hand-edited file. `clone-ahead` requires HEAD to be
**strictly** ahead — an uncommitted edit on a commit that *is* the ref is a
plain `content` mismatch.
## Modes in `contract-run.sh`
| `CONTRACT_REVISION_PREFLIGHT` | Behaviour |
| --- | --- |
| unset / **`warn` (default)** | log the refusal and its class, then still report |
| `enforce` | withhold the verdict, alert, exit `2` |
| `off` | skip the check entirely |
**The default is `warn`, deliberately.** The guard gates *every* scheduled
contract, and three legitimate situations would otherwise turn the whole
fleet's monitoring into withheld verdicts: a clone legitimately ahead of
`origin/master` mid-review, a detached HEAD, and an offline or failed fetch.
That is a bigger risk than the staleness the guard exists to catch. `warn`
keeps the signal loud and classified in every run's log without letting the
monitoring go dark.
### Criteria for flipping the default to `enforce`
Do not flip it on preference. Flip it when the evidence says the false-refusal
rate is low enough, as its own small change with its own review:
1. the guard has run across **every scheduled contract** for a sustained period
(suggested: 30 consecutive days, or 200+ contract runs) with **zero**
`mismatch:*` and **zero** `cannot-verify:*` refusals in the per-run logs;
2. no `cannot-verify:fetch-failed` arising from ordinary network blips in that
window — if the pinned clone's remote is not reliably reachable, `enforce`
will withhold rather than report;
3. the pinned runner clone is demonstrably kept current by fast-forward, so
`mismatch:clone-ahead` is a genuine fault rather than routine procedure.
The evidence for the flip is the `REASON=` lines already written into
`/var/log/contract-runs/`. Until then the default stays `warn`.
## Merge-time sequence (do this whenever this repo merges)
**Baseline as of 2026-09-25:** `/opt/contract-runner` is already
fast-forwarded to master `9faffe4`, so the pinned runner clone is current
today. This sequence exists to keep it that way.
After any merge to `master`:
```bash
# 1. fast-forward the pinned runner clone on CT 100
git -C /opt/contract-runner pull --ff-only
# 2. confirm it is current and clean
git -C /opt/contract-runner log --oneline -1
git -C /opt/contract-runner status --porcelain # expect no output
# 3. prove a contract runs and reports normally
CONTRACT_RUN_LOG_DIR=/tmp/preflight-proof \
bash /opt/contract-runner/scripts/contract-run.sh search-stack-visibility
echo "EXIT=$?" # expect 0, and 'revision-preflight: … matches origin/master'
```
A contract that reports a `REASON=mismatch:*` refusal here means the runner
clone is stale or locally edited — fast-forward it rather than reaching for
`CONTRACT_REVISION_PREFLIGHT=off`.
**Note on untracked files:** git refuses to fast-forward over an untracked file
even when its content is byte-identical to the incoming version
(`The following untracked working tree files would be overwritten by merge`).
A dirty clone will therefore block step 1. Resolve it by removing or stashing
the untracked paths first — that is exactly what blocked a clone on 2026-09-25.
## Operating notes
* Under the default `warn`, a stale pinned clone still produces verdicts but
every run logs the refusal and its class. Read those lines; do not ignore
them.
* Under `enforce`, a stale pinned clone **withholds**. That is the intended
failure. Recover by fast-forwarding:
`git -C /opt/contract-runner pull --ff-only`.
* When a contract legitimately changes, land it through the normal branch + PR
path and fast-forward the pinned clone. Do not copy files into it by hand.
* `--no-fetch` exists for offline inspection; it prints that freshness is
assumed rather than verified, and it is not used by `contract-run.sh`.