# Contract execution pinning Which copy of a contract script actually ran, and how that is proven. ## Why this exists Three times on 2026-09-25 a contract reported a verdict from a copy that was not the merged one: 1. The ops lane's own clone sat on the merged feature branch `fix/search-stack-multi-engine-20260925` at `8b2eba4` with no `pve_auth` fix, while it executed the daily digest from a different clone. Nothing in the workflow noticed. 2. `scripts/search-stack-check.py` was deployed into the pinned runner clone by hand rather than through git. 3. A stale local `origin/master` ref made an ancestry check report "unlanded work" for a branch that had in fact merged — the same staleness would have passed a stale script as current. A contract verdict is only meaningful if it came from the merged copy. The control is `scripts/revision-preflight.sh`. ## The rule **Every contract pins exactly one clone for execution: the clone that `scripts/contract-run.sh` itself lives in.** `contract-run.sh` derives that from its own location (`SCRIPTS_DIR`) and checks the script it is about to run against `origin/master` in the same clone. There is no second path to configure, and no contract may be executed from a hand-copied location. | Contract | Script | Pinned clone | | --- | --- | --- | | `infrastructure-monitoring` | `scripts/infra-monitoring.sh` | the clone containing `contract-run.sh` | | `proxmox-monitor` | `scripts/proxmox-monitor.sh` | same | | `zulip-health` | `scripts/zulip-monitor.sh` | same | | `agent-health-check` | `scripts/agent-health-check.py` | same | | `litellm-health` | `scripts/litellm-health-check.py` | same | | `disk-gc-threat-response` | `scripts/disk-gc-scan.py` | same | | `pm2-self-heal` | `scripts/pm2-self-heal.sh` | same | | `search-stack-visibility` | `scripts/search-stack-check.py` | same | ### The deployed runner The scheduler on **CT 100 (abiba)** runs contracts from **`/opt/contract-runner`** via `/etc/cron.d/contract-runner`. That clone is the pinned execution copy for every scheduled contract, and it must be kept current with `master` by fast-forward. Its `origin` is a local path to the upstream working copy, not a network remote. `daily-health-digest` is **not** in the table above because it has no contract file and no mapping — it is dispatched by cron as `fm-send.sh ops "run contract: daily-health-digest"` and was, until 2026-09-25, executed by hand from whichever clone the operator happened to be in. Creating its contract file and pinning it to a clone is an open follow-up. ## How the check works `scripts/revision-preflight.sh `: * resolves the **repo-relative** path of the executing script inside the clone; * **fetches** the remote first, so a stale local ref cannot make a stale script look current — bounded by `--fetch-timeout` (default 20s) so a hung remote cannot block a scheduled contract; * compares the script's sha256 against `:`; * **fails closed** — a path absent from the ref, an unresolvable ref, or a failed fetch is a failure, never a warning. ### Exit codes and reason classes The guard distinguishes **"I could not check"** from **"this copy is wrong"**, and every non-zero exit prints a machine-readable `REASON=` line before the human text, because a warning nobody can classify is not actionable — and the flip to `enforce` (below) depends on being able to read these apart. | exit | `REASON=` | meaning | | --- | --- | --- | | 0 | — | verified match | | 2 | `cannot-verify:fetch-failed` | remote unreachable, failed, or timed out | | 2 | `cannot-verify:ref-unresolvable` | `` does not exist in the clone | | 1 | `mismatch:path-absent` | the script does not exist in `` | | 1 | `mismatch:content` | the script differs from `` | | 1 | `mismatch:detached-head` | the clone is on a detached HEAD | | 1 | `mismatch:clone-ahead` | local HEAD is strictly ahead of `` (mid-review) | `detached-head` and `clone-ahead` are named separately on purpose: they are *legitimate* states that merely fail to be "the merged copy", and they are far less alarming than a hand-edited file. `clone-ahead` requires HEAD to be **strictly** ahead — an uncommitted edit on a commit that *is* the ref is a plain `content` mismatch. ## Modes in `contract-run.sh` | `CONTRACT_REVISION_PREFLIGHT` | Behaviour | | --- | --- | | unset / **`warn` (default)** | log the refusal and its class, then still report | | `enforce` | withhold the verdict, alert, exit `2` | | `off` | skip the check entirely | **The default is `warn`, deliberately.** The guard gates *every* scheduled contract, and three legitimate situations would otherwise turn the whole fleet's monitoring into withheld verdicts: a clone legitimately ahead of `origin/master` mid-review, a detached HEAD, and an offline or failed fetch. That is a bigger risk than the staleness the guard exists to catch. `warn` keeps the signal loud and classified in every run's log without letting the monitoring go dark. ### Criteria for flipping the default to `enforce` Do not flip it on preference. Flip it when the evidence says the false-refusal rate is low enough, as its own small change with its own review: 1. the guard has run across **every scheduled contract** for a sustained period (suggested: 30 consecutive days, or 200+ contract runs) with **zero** `mismatch:*` and **zero** `cannot-verify:*` refusals in the per-run logs; 2. no `cannot-verify:fetch-failed` arising from ordinary network blips in that window — if the pinned clone's remote is not reliably reachable, `enforce` will withhold rather than report; 3. the pinned runner clone is demonstrably kept current by fast-forward, so `mismatch:clone-ahead` is a genuine fault rather than routine procedure. The evidence for the flip is the `REASON=` lines already written into `/var/log/contract-runs/`. Until then the default stays `warn`. ## Merge-time sequence (do this whenever this repo merges) **Baseline as of 2026-09-25:** `/opt/contract-runner` is already fast-forwarded to master `9faffe4`, so the pinned runner clone is current today. This sequence exists to keep it that way. After any merge to `master`: ```bash # 1. fast-forward the pinned runner clone on CT 100 git -C /opt/contract-runner pull --ff-only # 2. confirm it is current and clean git -C /opt/contract-runner log --oneline -1 git -C /opt/contract-runner status --porcelain # expect no output # 3. prove a contract runs and reports normally CONTRACT_RUN_LOG_DIR=/tmp/preflight-proof \ bash /opt/contract-runner/scripts/contract-run.sh search-stack-visibility echo "EXIT=$?" # expect 0, and 'revision-preflight: … matches origin/master' ``` A contract that reports a `REASON=mismatch:*` refusal here means the runner clone is stale or locally edited — fast-forward it rather than reaching for `CONTRACT_REVISION_PREFLIGHT=off`. **Note on untracked files:** git refuses to fast-forward over an untracked file even when its content is byte-identical to the incoming version (`The following untracked working tree files would be overwritten by merge`). A dirty clone will therefore block step 1. Resolve it by removing or stashing the untracked paths first — that is exactly what blocked a clone on 2026-09-25. ## Operating notes * Under the default `warn`, a stale pinned clone still produces verdicts but every run logs the refusal and its class. Read those lines; do not ignore them. * Under `enforce`, a stale pinned clone **withholds**. That is the intended failure. Recover by fast-forwarding: `git -C /opt/contract-runner pull --ff-only`. * When a contract legitimately changes, land it through the normal branch + PR path and fast-forward the pinned clone. Do not copy files into it by hand. * `--no-fetch` exists for offline inspection; it prints that freshness is assumed rather than verified, and it is not used by `contract-run.sh`.