Bounded correction round on PR #134 after a PASS-WITH-FINDINGS review whose
Finding 4 is High. The guard's purpose and its fail-closed fix stand; the
problem was that with 'enforce' as the default it gates EVERY scheduled
contract, and three legitimate states produce a refusal - a clone legitimately
ahead of origin/master mid-review, a detached HEAD, and an offline or failed
fetch - so any of them would turn the fleet's monitoring into withheld
verdicts. That risk outweighs the staleness the guard catches.
1. DEFAULT IS NOW 'warn'. 'enforce' remains available and documented. The
criteria for flipping the default later are written into the doc as a
decision with evidence - a sustained window (30 days / 200+ runs) with zero
mismatch:* and zero cannot-verify:* refusals, no fetch blips, and a pinned
clone demonstrably kept current - explicitly as its own change, not a silent
flip.
2. 'COULD NOT CHECK' IS NOW DISTINGUISHABLE FROM 'THIS COPY IS WRONG'. Every
non-zero exit prints a machine-readable REASON=<class> line:
cannot-verify:fetch-failed | cannot-verify:ref-unresolvable (exit 2)
mismatch:path-absent | mismatch:content
mismatch:detached-head | mismatch:clone-ahead (exit 1)
detached-head and clone-ahead are named separately because they are
legitimate states, far less alarming than a hand-edited file. clone-ahead
requires HEAD to be STRICTLY ahead; an uncommitted edit on a commit that IS
the ref is a plain content mismatch (my own first cut got this wrong and the
new test 7d caught it).
3. THE DEFAULT FETCH IS BOUNDED: --fetch-timeout, default 20s, 0 = unbounded,
and a missing 'timeout' binary is itself a cannot-verify rather than an
unbounded fetch inside a scheduled contract.
4. TEST COVERAGE ADDED for every new class: fetch failure, fetch timeout
(asserted to return promptly under a 1s bound), unresolvable ref, detached
HEAD, clone-ahead, genuine content mismatch, and the contract-run.sh default.
The pre-fix draft fixture comparisons are kept: 31 passed, 0 failed.
5. MERGE-TIME SEQUENCE documented: fast-forward /opt/contract-runner, confirm
clean, prove a contract runs and reports. Baseline recorded as of today -
firstmate has already fast-forwarded it to 9faffe4 - with the note that an
untracked file blocks a fast-forward even when byte-identical.
Live behaviour re-verified on the real runner path:
default: REASON=mismatch:clone-ahead -> 'continuing because ...=warn' -> VERDICT: PASS, exit 0
enforce: REASON=mismatch:clone-ahead -> 'VERDICT WITHHELD: mismatch:clone-ahead', exit 2
MANDATORY CHECKS (master went red once from a credential-SHAPED string, so
these are now run on every shippable branch):
bash scripts/prose-lint.sh -> LINT PASSED (18 warning(s))
secret scan -> secret scan clean (tree; 34 allowlisted,
24 inert value(s) ignored); No committed credentials
shellcheck revision-preflight.sh -> clean
shellcheck test_revision_preflight.sh -> clean
shellcheck contract-run.sh -> SC2034 x1, SC2086 x2 - byte-identical on
master, i.e. pre-existing, none introduced
tests/test_probe_drift.py::test_prose_lint_accepts_report_format_with_provenance
fails both before and after this branch (it runs prose-lint from a temp CWD and
cannot find its sibling secret-scan.sh). Pre-existing, unrelated, not fixed here.
7.8 KiB
Contract execution pinning
Which copy of a contract script actually ran, and how that is proven.
Why this exists
Three times on 2026-09-25 a contract reported a verdict from a copy that was not the merged one:
- The ops lane's own clone sat on the merged feature branch
fix/search-stack-multi-engine-20260925at8b2eba4with nopve_authfix, while it executed the daily digest from a different clone. Nothing in the workflow noticed. scripts/search-stack-check.pywas deployed into the pinned runner clone by hand rather than through git.- A stale local
origin/masterref made an ancestry check report "unlanded work" for a branch that had in fact merged — the same staleness would have passed a stale script as current.
A contract verdict is only meaningful if it came from the merged copy. The
control is scripts/revision-preflight.sh.
The rule
Every contract pins exactly one clone for execution: the clone that
scripts/contract-run.sh itself lives in.
contract-run.sh derives that from its own location (SCRIPTS_DIR) and checks
the script it is about to run against origin/master in the same clone. There
is no second path to configure, and no contract may be executed from a
hand-copied location.
| Contract | Script | Pinned clone |
|---|---|---|
infrastructure-monitoring |
scripts/infra-monitoring.sh |
the clone containing contract-run.sh |
proxmox-monitor |
scripts/proxmox-monitor.sh |
same |
zulip-health |
scripts/zulip-monitor.sh |
same |
agent-health-check |
scripts/agent-health-check.py |
same |
litellm-health |
scripts/litellm-health-check.py |
same |
disk-gc-threat-response |
scripts/disk-gc-scan.py |
same |
pm2-self-heal |
scripts/pm2-self-heal.sh |
same |
search-stack-visibility |
scripts/search-stack-check.py |
same |
The deployed runner
The scheduler on CT 100 (abiba) runs contracts from
/opt/contract-runner via /etc/cron.d/contract-runner. That clone is the
pinned execution copy for every scheduled contract, and it must be kept current
with master by fast-forward. Its origin is a local path to the upstream
working copy, not a network remote.
daily-health-digest is not in the table above because it has no contract
file and no mapping — it is dispatched by cron as
fm-send.sh ops "run contract: daily-health-digest" and was, until
2026-09-25, executed by hand from whichever clone the operator happened to be
in. Creating its contract file and pinning it to a clone is an open follow-up.
How the check works
scripts/revision-preflight.sh <script-path> <clone-path>:
- resolves the repo-relative path of the executing script inside the clone;
- fetches the remote first, so a stale local ref cannot make a stale script
look current — bounded by
--fetch-timeout(default 20s) so a hung remote cannot block a scheduled contract; - compares the script's sha256 against
<ref>:<repo-relative-path>; - fails closed — a path absent from the ref, an unresolvable ref, or a failed fetch is a failure, never a warning.
Exit codes and reason classes
The guard distinguishes "I could not check" from "this copy is wrong",
and every non-zero exit prints a machine-readable REASON=<class> line before
the human text, because a warning nobody can classify is not actionable — and
the flip to enforce (below) depends on being able to read these apart.
| exit | REASON= |
meaning |
|---|---|---|
| 0 | — | verified match |
| 2 | cannot-verify:fetch-failed |
remote unreachable, failed, or timed out |
| 2 | cannot-verify:ref-unresolvable |
<ref> does not exist in the clone |
| 1 | mismatch:path-absent |
the script does not exist in <ref> |
| 1 | mismatch:content |
the script differs from <ref> |
| 1 | mismatch:detached-head |
the clone is on a detached HEAD |
| 1 | mismatch:clone-ahead |
local HEAD is strictly ahead of <ref> (mid-review) |
detached-head and clone-ahead are named separately on purpose: they are
legitimate states that merely fail to be "the merged copy", and they are far
less alarming than a hand-edited file. clone-ahead requires HEAD to be
strictly ahead — an uncommitted edit on a commit that is the ref is a
plain content mismatch.
Modes in contract-run.sh
CONTRACT_REVISION_PREFLIGHT |
Behaviour |
|---|---|
unset / warn (default) |
log the refusal and its class, then still report |
enforce |
withhold the verdict, alert, exit 2 |
off |
skip the check entirely |
The default is warn, deliberately. The guard gates every scheduled
contract, and three legitimate situations would otherwise turn the whole
fleet's monitoring into withheld verdicts: a clone legitimately ahead of
origin/master mid-review, a detached HEAD, and an offline or failed fetch.
That is a bigger risk than the staleness the guard exists to catch. warn
keeps the signal loud and classified in every run's log without letting the
monitoring go dark.
Criteria for flipping the default to enforce
Do not flip it on preference. Flip it when the evidence says the false-refusal rate is low enough, as its own small change with its own review:
- the guard has run across every scheduled contract for a sustained period
(suggested: 30 consecutive days, or 200+ contract runs) with zero
mismatch:*and zerocannot-verify:*refusals in the per-run logs; - no
cannot-verify:fetch-failedarising from ordinary network blips in that window — if the pinned clone's remote is not reliably reachable,enforcewill withhold rather than report; - the pinned runner clone is demonstrably kept current by fast-forward, so
mismatch:clone-aheadis a genuine fault rather than routine procedure.
The evidence for the flip is the REASON= lines already written into
/var/log/contract-runs/. Until then the default stays warn.
Merge-time sequence (do this whenever this repo merges)
Baseline as of 2026-09-25: /opt/contract-runner is already
fast-forwarded to master 9faffe4, so the pinned runner clone is current
today. This sequence exists to keep it that way.
After any merge to master:
# 1. fast-forward the pinned runner clone on CT 100
git -C /opt/contract-runner pull --ff-only
# 2. confirm it is current and clean
git -C /opt/contract-runner log --oneline -1
git -C /opt/contract-runner status --porcelain # expect no output
# 3. prove a contract runs and reports normally
CONTRACT_RUN_LOG_DIR=/tmp/preflight-proof \
bash /opt/contract-runner/scripts/contract-run.sh search-stack-visibility
echo "EXIT=$?" # expect 0, and 'revision-preflight: … matches origin/master'
A contract that reports a REASON=mismatch:* refusal here means the runner
clone is stale or locally edited — fast-forward it rather than reaching for
CONTRACT_REVISION_PREFLIGHT=off.
Note on untracked files: git refuses to fast-forward over an untracked file
even when its content is byte-identical to the incoming version
(The following untracked working tree files would be overwritten by merge).
A dirty clone will therefore block step 1. Resolve it by removing or stashing
the untracked paths first — that is exactly what blocked a clone on 2026-09-25.
Operating notes
- Under the default
warn, a stale pinned clone still produces verdicts but every run logs the refusal and its class. Read those lines; do not ignore them. - Under
enforce, a stale pinned clone withholds. That is the intended failure. Recover by fast-forwarding:git -C /opt/contract-runner pull --ff-only. - When a contract legitimately changes, land it through the normal branch + PR path and fast-forward the pinned clone. Do not copy files into it by hand.
--no-fetchexists for offline inspection; it prints that freshness is assumed rather than verified, and it is not used bycontract-run.sh.