Files
prose-contracts/docs/contract-execution-pinning.md
T
root c5a6dcd42a
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
fix(revision-preflight): default to warn, name the refusal class, bound the fetch
Bounded correction round on PR #134 after a PASS-WITH-FINDINGS review whose
Finding 4 is High. The guard's purpose and its fail-closed fix stand; the
problem was that with 'enforce' as the default it gates EVERY scheduled
contract, and three legitimate states produce a refusal - a clone legitimately
ahead of origin/master mid-review, a detached HEAD, and an offline or failed
fetch - so any of them would turn the fleet's monitoring into withheld
verdicts. That risk outweighs the staleness the guard catches.

1. DEFAULT IS NOW 'warn'. 'enforce' remains available and documented. The
   criteria for flipping the default later are written into the doc as a
   decision with evidence - a sustained window (30 days / 200+ runs) with zero
   mismatch:* and zero cannot-verify:* refusals, no fetch blips, and a pinned
   clone demonstrably kept current - explicitly as its own change, not a silent
   flip.

2. 'COULD NOT CHECK' IS NOW DISTINGUISHABLE FROM 'THIS COPY IS WRONG'. Every
   non-zero exit prints a machine-readable REASON=<class> line:
     cannot-verify:fetch-failed | cannot-verify:ref-unresolvable   (exit 2)
     mismatch:path-absent | mismatch:content
     mismatch:detached-head | mismatch:clone-ahead                (exit 1)
   detached-head and clone-ahead are named separately because they are
   legitimate states, far less alarming than a hand-edited file. clone-ahead
   requires HEAD to be STRICTLY ahead; an uncommitted edit on a commit that IS
   the ref is a plain content mismatch (my own first cut got this wrong and the
   new test 7d caught it).

3. THE DEFAULT FETCH IS BOUNDED: --fetch-timeout, default 20s, 0 = unbounded,
   and a missing 'timeout' binary is itself a cannot-verify rather than an
   unbounded fetch inside a scheduled contract.

4. TEST COVERAGE ADDED for every new class: fetch failure, fetch timeout
   (asserted to return promptly under a 1s bound), unresolvable ref, detached
   HEAD, clone-ahead, genuine content mismatch, and the contract-run.sh default.
   The pre-fix draft fixture comparisons are kept: 31 passed, 0 failed.

5. MERGE-TIME SEQUENCE documented: fast-forward /opt/contract-runner, confirm
   clean, prove a contract runs and reports. Baseline recorded as of today -
   firstmate has already fast-forwarded it to 9faffe4 - with the note that an
   untracked file blocks a fast-forward even when byte-identical.

Live behaviour re-verified on the real runner path:
  default: REASON=mismatch:clone-ahead -> 'continuing because ...=warn' -> VERDICT: PASS, exit 0
  enforce: REASON=mismatch:clone-ahead -> 'VERDICT WITHHELD: mismatch:clone-ahead', exit 2

MANDATORY CHECKS (master went red once from a credential-SHAPED string, so
these are now run on every shippable branch):
  bash scripts/prose-lint.sh        -> LINT PASSED (18 warning(s))
  secret scan                       -> secret scan clean (tree; 34 allowlisted,
                                       24 inert value(s) ignored); No committed credentials
  shellcheck revision-preflight.sh  -> clean
  shellcheck test_revision_preflight.sh -> clean
  shellcheck contract-run.sh        -> SC2034 x1, SC2086 x2 - byte-identical on
                                       master, i.e. pre-existing, none introduced

tests/test_probe_drift.py::test_prose_lint_accepts_report_format_with_provenance
fails both before and after this branch (it runs prose-lint from a temp CWD and
cannot find its sibling secret-scan.sh). Pre-existing, unrelated, not fixed here.
2026-09-25 11:29:15 +00:00

7.8 KiB

Contract execution pinning

Which copy of a contract script actually ran, and how that is proven.

Why this exists

Three times on 2026-09-25 a contract reported a verdict from a copy that was not the merged one:

  1. The ops lane's own clone sat on the merged feature branch fix/search-stack-multi-engine-20260925 at 8b2eba4 with no pve_auth fix, while it executed the daily digest from a different clone. Nothing in the workflow noticed.
  2. scripts/search-stack-check.py was deployed into the pinned runner clone by hand rather than through git.
  3. A stale local origin/master ref made an ancestry check report "unlanded work" for a branch that had in fact merged — the same staleness would have passed a stale script as current.

A contract verdict is only meaningful if it came from the merged copy. The control is scripts/revision-preflight.sh.

The rule

Every contract pins exactly one clone for execution: the clone that scripts/contract-run.sh itself lives in.

contract-run.sh derives that from its own location (SCRIPTS_DIR) and checks the script it is about to run against origin/master in the same clone. There is no second path to configure, and no contract may be executed from a hand-copied location.

Contract Script Pinned clone
infrastructure-monitoring scripts/infra-monitoring.sh the clone containing contract-run.sh
proxmox-monitor scripts/proxmox-monitor.sh same
zulip-health scripts/zulip-monitor.sh same
agent-health-check scripts/agent-health-check.py same
litellm-health scripts/litellm-health-check.py same
disk-gc-threat-response scripts/disk-gc-scan.py same
pm2-self-heal scripts/pm2-self-heal.sh same
search-stack-visibility scripts/search-stack-check.py same

The deployed runner

The scheduler on CT 100 (abiba) runs contracts from /opt/contract-runner via /etc/cron.d/contract-runner. That clone is the pinned execution copy for every scheduled contract, and it must be kept current with master by fast-forward. Its origin is a local path to the upstream working copy, not a network remote.

daily-health-digest is not in the table above because it has no contract file and no mapping — it is dispatched by cron as fm-send.sh ops "run contract: daily-health-digest" and was, until 2026-09-25, executed by hand from whichever clone the operator happened to be in. Creating its contract file and pinning it to a clone is an open follow-up.

How the check works

scripts/revision-preflight.sh <script-path> <clone-path>:

  • resolves the repo-relative path of the executing script inside the clone;
  • fetches the remote first, so a stale local ref cannot make a stale script look current — bounded by --fetch-timeout (default 20s) so a hung remote cannot block a scheduled contract;
  • compares the script's sha256 against <ref>:<repo-relative-path>;
  • fails closed — a path absent from the ref, an unresolvable ref, or a failed fetch is a failure, never a warning.

Exit codes and reason classes

The guard distinguishes "I could not check" from "this copy is wrong", and every non-zero exit prints a machine-readable REASON=<class> line before the human text, because a warning nobody can classify is not actionable — and the flip to enforce (below) depends on being able to read these apart.

exit REASON= meaning
0 — verified match
2 cannot-verify:fetch-failed remote unreachable, failed, or timed out
2 cannot-verify:ref-unresolvable <ref> does not exist in the clone
1 mismatch:path-absent the script does not exist in <ref>
1 mismatch:content the script differs from <ref>
1 mismatch:detached-head the clone is on a detached HEAD
1 mismatch:clone-ahead local HEAD is strictly ahead of <ref> (mid-review)

detached-head and clone-ahead are named separately on purpose: they are legitimate states that merely fail to be "the merged copy", and they are far less alarming than a hand-edited file. clone-ahead requires HEAD to be strictly ahead — an uncommitted edit on a commit that is the ref is a plain content mismatch.

Modes in contract-run.sh

CONTRACT_REVISION_PREFLIGHT Behaviour
unset / warn (default) log the refusal and its class, then still report
enforce withhold the verdict, alert, exit 2
off skip the check entirely

The default is warn, deliberately. The guard gates every scheduled contract, and three legitimate situations would otherwise turn the whole fleet's monitoring into withheld verdicts: a clone legitimately ahead of origin/master mid-review, a detached HEAD, and an offline or failed fetch. That is a bigger risk than the staleness the guard exists to catch. warn keeps the signal loud and classified in every run's log without letting the monitoring go dark.

Criteria for flipping the default to enforce

Do not flip it on preference. Flip it when the evidence says the false-refusal rate is low enough, as its own small change with its own review:

  1. the guard has run across every scheduled contract for a sustained period (suggested: 30 consecutive days, or 200+ contract runs) with zero mismatch:* and zero cannot-verify:* refusals in the per-run logs;
  2. no cannot-verify:fetch-failed arising from ordinary network blips in that window — if the pinned clone's remote is not reliably reachable, enforce will withhold rather than report;
  3. the pinned runner clone is demonstrably kept current by fast-forward, so mismatch:clone-ahead is a genuine fault rather than routine procedure.

The evidence for the flip is the REASON= lines already written into /var/log/contract-runs/. Until then the default stays warn.

Merge-time sequence (do this whenever this repo merges)

Baseline as of 2026-09-25: /opt/contract-runner is already fast-forwarded to master 9faffe4, so the pinned runner clone is current today. This sequence exists to keep it that way.

After any merge to master:

# 1. fast-forward the pinned runner clone on CT 100
git -C /opt/contract-runner pull --ff-only

# 2. confirm it is current and clean
git -C /opt/contract-runner log --oneline -1
git -C /opt/contract-runner status --porcelain     # expect no output

# 3. prove a contract runs and reports normally
CONTRACT_RUN_LOG_DIR=/tmp/preflight-proof \
  bash /opt/contract-runner/scripts/contract-run.sh search-stack-visibility
echo "EXIT=$?"     # expect 0, and 'revision-preflight: … matches origin/master'

A contract that reports a REASON=mismatch:* refusal here means the runner clone is stale or locally edited — fast-forward it rather than reaching for CONTRACT_REVISION_PREFLIGHT=off.

Note on untracked files: git refuses to fast-forward over an untracked file even when its content is byte-identical to the incoming version (The following untracked working tree files would be overwritten by merge). A dirty clone will therefore block step 1. Resolve it by removing or stashing the untracked paths first — that is exactly what blocked a clone on 2026-09-25.

Operating notes

  • Under the default warn, a stale pinned clone still produces verdicts but every run logs the refusal and its class. Read those lines; do not ignore them.
  • Under enforce, a stale pinned clone withholds. That is the intended failure. Recover by fast-forwarding: git -C /opt/contract-runner pull --ff-only.
  • When a contract legitimately changes, land it through the normal branch + PR path and fast-forward the pinned clone. Do not copy files into it by hand.
  • --no-fetch exists for offline inspection; it prints that freshness is assumed rather than verified, and it is not used by contract-run.sh.