27 KiB
CHECKPOINT — Alpaca Trading System Remediation Plan
Status: GATE 0 COMPLETE · Created: 2026-07-01 · Resumed: 2026-07-01T14:35Z Pick-up instruction for a new session: Read this file top-to-bottom. Gate 0 is done; resume at Gate 1.
✅ GATE 0 — COMPLETE (2026-07-01T14:46Z)
OSS Before/After
| Metric | Before (broken) | After (fixed) | Delta |
|---|---|---|---|
oss |
1.094 (above target!) | -0.605 (below floor) | −1.699 |
oss_trend |
"improving" | "degrading" | ✓ fixed |
calmar |
2.174 (sign destroyed) | -2.174 | sign restored |
spy_30d_return |
0.58 (hardcoded) | -1.44 (live) | real benchmark |
rolling_alpha |
-1.01 (wrong) | 1.01 (portfolio beat SPY) | flipped |
portfolio_beta |
-0.74 (wrong) | 0.30 (correct) | sign fixed |
win_rate_adaptations |
0.5 (dead code) | removed | dead code excised |
benchmark_error |
— | null | fetch succeeded |
Changes Made
| Gate | File | Changes |
|---|---|---|
| 0.1 | scripts/compute_oss.py:65-66 |
abs(cagr) → cagr; floor 0.0 → -5.0 |
| 0.2 | scripts/compute_oss.py:344-362 |
Removed dead _compute_adapt_win_rate() + call site |
| 0.3 | scripts/compute_oss.py:45-61 |
Added _fetch_spy_30d_return() (yfinance via ^GSPC) |
| 0.3 | scripts/compute_oss.py:293-301 |
Replaced hardcoded 0.58 with live fetch; alpha=None on failure |
| 0.3 | scripts/compute_oss.py:334 |
_compute_beta now accepts live spy_return param |
| 0.3 | scripts/compute_oss.py:328-329 |
JSON output now includes spy_30d_return + benchmark_error |
| 0.4 | tests/test_oss_monotonicity.py |
NEW — 3 property tests (Calmar sign, monotonicity, known values) |
| 0.4 | tests/test_benchmark_integrity.py |
NEW — 3 tests (no 0.58 literal, fetch signature, alpha sign mock) |
| 0.4 | ~/.hermes/scripts/contract_invariant_watchdog.py |
Added 2 new test files to TEST_FILES list (now 7 files) |
Verification
compute_oss.pyruns clean: OSS=-0.605, trend=degrading, live SPY=-1.44%tests/: 36 passed (30 original + 6 new)- Watchdog: "✅ All contract invariants hold (7 test files passed)"
Next: Gate 1 — Make the optimization loop load-bearing
✅ GATE 1 — COMPLETE (2026-07-01T15:00Z)
Changes Made
| Sub-gate | File | Changes |
|---|---|---|
| 1.1 | scripts/migrate_gate1_schema.py |
NEW — creates adaptations + parameter_state tables, seeds 5 params |
| 1.1 | trade_journal.db |
parameter_state (5 params) + adaptations (0 rows, clean) |
| 1.2 | scripts/apply_adaptations.py |
NEW — reads OSS+regime, proposes bounded changes, writes DB |
| 1.3 | scripts/apply_adaptations.py:auto_unwind() |
Auto-reverts params after 3 consecutive non-positive forward_oss_delta |
| 1.4 | scripts/compute_oss.py:W_ADAPT=0.10 |
Re-weighted: W_CALMAR=0.35, W_WIN_RATE=0.10, W_ADAPT=0.10 |
| 1.4 | scripts/compute_oss.py:_compute_adapt_win_rate() |
Queries adaptations.result='improved' (not dead analysis_runs) |
| 1.5 | tests/test_adaptation_persistence.py |
NEW — 3 tests (seeded params, write+unwind, migration idempotent) |
| 1.5 | ~/.hermes/scripts/contract_invariant_watchdog.py |
Added test_adaptation_persistence.py (now 8 test files) |
Adaptation Logic
| OSS Range | Trend | Regime | Action |
|---|---|---|---|
| < 0.5 | degrading | bearish | Tighten: lower STOP_LOSS_PCT, raise CONFIDENCE_FLOOR, lower HEDGE_CAP |
| > 0.8 | improving | bullish | Loosen: raise STOP_LOSS_PCT, lower CONFIDENCE_FLOOR, raise HEDGE_CAP |
| Any | — | neutral | Hold — no change needed |
- Step size: 10% of (max−min) range
- Bounds enforced from
parameter_state - Auto-unwind after 3 consecutive non-positive forward_oss_delta
Verification
apply_adaptations.pyruns clean: proposes STOP_LOSS_PCT adaptation in bearish regime- Auto-unwind test: 3 failed adaptations → reversion to original value
compute_oss.pyshowsAdaptWR=0.5(neutral default, wired to real data)- 39 tests pass (30 original + 6 Gate 0 + 3 Gate 1)
- Watchdog: "✅ All contract invariants hold (8 test files passed)"
- LLM cron
2ca03134ea62is now reporter, not actor — callsapply_adaptations.py
Next: Gate 2 — Make hedging capable
0. ORIENTATION (read first)
Repo: /opt/alpaca-trading-agent · venv: .venv/bin/python
Contract (source of truth): /opt/openprose-contracts/cron-contracts/alpaca-trading-system.prose.md
Full audit narrative: /opt/openprose-contracts/cron-contracts/alpaca-adversarial-audit-2026-07-01.md
Watchdog: ~/.hermes/scripts/contract_invariant_watchdog.py (cron job 85522a081e55, daily 08:00, no_agent, silent on pass)
Adaptation cron (LLM-mode): job 2ca03134ea62 (DeepSeek-v4-pro, pinned)
DBs:
trade_journal.db— perf, orders, trades,analysis_runs,config_store. This is the DBcompute_oss.pyreads (DB_PATH, line 18).consensus/signals.db—signals,consensus.
Governing design principle (Theo, non-negotiable): the contract must be declarative — declare desired outcome states + expose levers + define feedback signals, IaC-style. NOT prescriptive step lists. Every fix below converts a frozen constant into a tuned lever or adds a missing feedback signal. Wherever you must hardcode, justify it as an irreducible exception.
[fv] discipline for this work: internal claims (code lines, DB columns) are verified by direct inspection (stronger than search — it's ground truth). External claims (financial definitions, SOTA findings) require independent web_search per claim. Label every assertion CONFIRMED / DISPUTED / UNVERIFIED. Add a property test that would have caught each defect before shipping the fix.
1. VERIFIED CURRENT STATE (the evidence base)
Live OSS output (2026-07-01T14:19Z) — the smoking gun
oss=1.094 (target 1.0, floor 0.5) oss_trend="improving"
calmar=2.174 profit_factor=0.679 win_rate=0.5
total_return_pct=-0.43% sharpe=-1.38 realized_pnl_30d=-$1.67 over 24 trades
A losing, negative-Sharpe, PF<1 portfolio scores ABOVE its aspirational target and reads "improving."
VERIFY: cd /opt/alpaca-trading-agent && .venv/bin/python scripts/compute_oss.py | python3 -m json.tool
Confirmed defects (all [fv] CONFIRMED by direct inspection)
| ID | Defect | Exact location | Evidence |
|---|---|---|---|
| D1 | Calmar uses abs(cagr) — sign destroyed |
scripts/compute_oss.py:65 |
compute_calmar(-2.35,-0.43,30)=2.174 vs (+0.43)=2.28 |
| D1b | Calmar clamp floor max(calmar,0.0) erases negatives even if sign fixed |
scripts/compute_oss.py:66 |
return round(min(max(calmar,0.0),5.0),3) |
| D2 | Benchmark hardcoded | scripts/compute_oss.py:274 (portfolio_return - 0.58), :319 (spy_30d_return = 0.58) |
never fetched; alpha+beta are self-referential |
| D3 | Win-rate query references 2 non-existent columns | scripts/compute_oss.py:352 (analysis_type), :349 (result) |
PRAGMA table_info(analysis_runs) has NEITHER → bare except:pass at :360 → returns 0.5 forever |
| D4 | Hedge structurally incapable | .env.prod + consensus/ross/config.py:84-86 |
MAX_HEDGE_POSITIONS=1, HEDGE_POSITION_SIZE_PCT=0.01 → max ~1% of equity; contract Rule 8 assumes 25% (PARKER_HEDGE_CAP_PCT=0.25, config.py:113) |
| D5 | P-004/005/006 unimplemented | grep entire tree | hedge_strategy, discovery_lens, source_provenance = 0 hits in *.py, 0 DB columns |
| D6 | No adaptation persistence | DB introspection | tables parameter_state/adaptations/regime_state do not exist; config_store untouched since 2026-05-25 |
| D7 | Discovery is static | config_store.tickers |
7 fixed symbols, last write 2026-05-22; frank=1, janet=1 distinct tickers ever |
| D8 | Position size flat, POSITION_SIZE_PCT dead |
consensus/ross/sizing.py:7, config.py:138 |
calculate_position_size returns cfg.MAX_POSITION_NOTIONAL which = MIN_NOTIONAL ($5); the 0.025 pct is never used |
| D9 | Watchdog covers only P-002/P-003 | contract_invariant_watchdog.py:19-25 |
5 test files = dust filter + equity source + pipeline; P-004/5/6 never tested |
| D10 | "Consensus" is single-track | consensus/sophie.py + signals.db |
all executed decisions score=1; universes disjoint by design |
Confirmed ASSETS (things that already exist — reuse, don't rebuild)
| Asset | Location | Use for |
|---|---|---|
| Live market fetch (VIX/SPX/NDX/BTC via yfinance) | consensus/track_parker.py:76-101 (_fetch_market_data) |
D2 fix — reuse spx.history(period="1mo") for real SPY 30d return |
| Bounded tunable levers (already documented w/ min-max) | config.py:99-113: STOP_LOSS_PCT[-0.03,-0.07], PROFIT_TAKE_PCT[0.03,0.10], STALE_EXIT_DAYS[3,7], CONFIDENCE_FLOOR[0.65,0.85], PARKER_HEDGE_CAP_PCT[0.10,0.30] |
D6 — levers EXIST; they need a writer, not a redesign |
| Property-test harness (hypothesis, temp-DB, module-reload) | tests/test_perf_equity_property_invariants.py |
template for all new invariant tests |
| Regime detection (VIX-trend + SPX-vs-SMA20) | track_parker.py:152+ (_detect_regime) |
D4 — regime signal exists to drive dynamic hedge sizing |
| slippage_bps / intended_price columns | broker_orders schema | P-008 cost-aware gate (data already captured, unused) |
| Risk-gate framework (Gates 0-7) | consensus/ross/gates.py:14 (check_risk_gates) |
insert cost-aware gate here |
2. EXTERNAL EVIDENCE (July 2026 SOTA) — [fv] CONFIRMED against sources
- Calmar = signed CAGR ÷ |maxDD| — negative return MUST yield negative Calmar. (Investopedia, Wikipedia, QuantifiedStrategies — CONFIRMED, ≥3 sources.) → D1/D1b are unambiguous bugs.
- FINSABER (arXiv:2505.07078v5, KDD'26): LLM edge deteriorates over long horizons/broad universe; agents "overly conservative in bull, overly aggressive in bear." Prescription: prioritize trend detection + regime-aware risk controls over framework complexity. CONFIRMED (abstract).
- StockBench (arXiv:2510.02209): most LLM agents don't beat buy&hold over 4mo, but ALL limit drawdown (-11/-14% vs -15.2%). → realistic capturable edge = drawdown control, not alpha. CONFIRMED.
- Agent Market Arena (arXiv:2510.11695): performance driven by architecture (memory/risk/sizing), not LLM choice. CONFIRMED.
- CLQT (arXiv:2606.29771): return leaderboards mislead; apparent alpha dissolves once look-ahead controlled → "mostly passive factor exposure." Evaluate capability, not single score. CONFIRMED. → OSS is exactly the single-number leaderboard being warned against.
- Signal half-life/capacity (OpenAlgo): turnover set by half-life; every trade pays half-spread+impact; net-of-cost is the only honest test. CONFIRMED.
- Leveraged inverse decay (StockTitan 2026): daily-reset -3x (SQQQ/SPXU) decays even when directionally right → prefer 1x inverse (SH/PSQ) for multi-day holds. CONFIRMED. Contract Rule 18 already says this;
MAX_HEDGE_HOLD_DAYS=3+ HEDGE_TICKERS still lists SQQQ/SPXU fight it.
Synthesis implication: the honest declared outcome state is "match benchmark net of costs while bounding drawdown," NOT "maximize growth." Re-weight the objective accordingly (P-011).
3. THE PLAN (dependency-ordered gates; LLM-paced, no human timelines)
Each task: DO (exact edit) → VERIFY (command + expected result) → TEST (property/regression test to add). No task is "done" until its VERIFY passes and its TEST is committed.
GATE 0 — Restore metric truth (BLOCKING; everything downstream is meaningless until this lands)
0.1 — Fix Calmar sign + clamp
- DO:
scripts/compute_oss.py:65→calmar = cagr / (abs(max_drawdown_pct) / 100.0)(dropabs(on cagr). - DO:
:66→ allow negatives:return round(min(max(calmar, -5.0), 5.0), 3)(floor -5.0, not 0.0). - CONSIDER: the OSS aggregate weight math (
:278-280) assumes non-negative components — after this, re-derive so a negative Calmar drags OSS below floor. Verify OSS can now go below 0.5 on a losing month. - VERIFY:
.venv/bin/python -c "import sys;sys.path.insert(0,'scripts');from compute_oss import compute_calmar as c;print(c(-2.35,-0.43,30), c(-2.35,0.43,30))"→ first value must be negative, second positive. - VERIFY:
.venv/bin/python scripts/compute_oss.py→ with current -0.43% return,ossmust now be < 0.5 andoss_trendmust NOT be "improving". - TEST: add
tests/test_oss_monotonicity.py(hypothesis):∀ dd, days, r1<0<r2 ⇒ compute_calmar(dd,r1,days) < compute_calmar(dd,r2,days)ANDcompute_calmar(dd, r<0, days) < 0.
0.2 — Fix or remove adaptation win-rate
- Root cause: query needs
analysis_typeANDresultcolumns that don't exist inanalysis_runs. - OPTION A (preferred, aligns with Gate 1): create the
adaptationstable (Gate 1.1) and repoint this query at it. - OPTION B (interim): remove the
W_ADAPT * 0.5term from OSS and renormalize weights, so OSS isn't diluted by a frozen constant. Document as temporary. - VERIFY: grep confirms no silent
except: passreturns a magic 0.5 that feeds OSS. - TEST: assert
_compute_adapt_win_rateraises or returns None (not 0.5) when the backing table/columns are absent — no silent defaults.
0.3 — Live benchmark (kill the 0.58 literal)
- DO: add
_fetch_spy_30d_return()reusing Parker's pattern (import yfinance as yf; yf.Ticker("^GSPC").history(period="1mo"); 30d pct change). Thread into:274(alpha) and:319(beta). - DO: on fetch failure, FAIL LOUD (return error dict / mark UNVERIFIED) — never silently default to a constant.
- VERIFY:
.venv/bin/python scripts/compute_oss.pyshows aspy_30d_returnthat changes day-to-day and ≠ 0.58;rolling_alphaflips sign when portfolio crosses SPY. - TEST:
tests/test_benchmark_integrity.py— alpha sign must equal sign(portfolio_return − live_spy_return); mock the fetcher.
0.4 — Gate 0 acceptance
- Re-run
compute_oss.py; confirm it now reflects reality (the -$1.67 loss). - Add 0.1/0.2/0.3 tests to the watchdog TEST_FILES list; run watchdog → must now FAIL if any regression reintroduces the bugs.
GATE 1 — Make the optimization loop load-bearing (depends on Gate 0)
1.1 — Persistence schema. Create tables in trade_journal.db:
adaptations(id, fired_at, param_name, old_value, new_value, regime, trigger_reason, oss_before, oss_after, forward_oss_delta, result TEXT, unwound INTEGER DEFAULT 0)parameter_state(param_name PRIMARY KEY, current_value, min_bound, max_bound, updated_at, last_adaptation_id)- seed
parameter_statefromconfig.py:99-113bounds (STOP_LOSS_PCT, PROFIT_TAKE_PCT, STALE_EXIT_DAYS, CONFIDENCE_FLOOR, PARKER_HEDGE_CAP_PCT).
1.2 — Move adaptation from LLM prompt to code. Create scripts/apply_adaptations.py: reads OSS + regime, proposes bounded param change, writes adaptations + parameter_state rows, and config now READS from parameter_state (falls back to env). LLM cron 2ca03134ea62 becomes reporter, not actor.
- Declarative framing: contract declares target OSS ≥ 1.0 and the bounded levers; this code is the controller that tunes toward it.
1.3 — Auto-unwind (Rule 8-14 guardrail). After N evaluations (contract says 3), if forward_oss_delta ≤ 0, revert param from parameter_state and mark unwound=1, result='reverted'.
- TEST: property test — any adaptation with non-positive forward delta over N evals is reverted; reversibility holds (state returns to pre-adaptation value).
1.4 — Wire win-rate. Repoint _compute_adapt_win_rate at adaptations.result='improved'. Now OSS weight 0.15 is real.
1.5 — Watchdog. Add P-009 (monotonicity) + P-010 (persistence/reversibility) tests. Watchdog fails if loop goes inert.
Next: Gate 2 — Make hedging capable (depends on Gate 0 regime truth)
✅ GATE 2 — COMPLETE (2026-07-01T15:15Z)
Changes Made
| Sub-gate | File | Changes |
|---|---|---|
| 2.1 | consensus/ross/hedge.py |
NEW — regime-driven hedge sizing buckets, cap enforcement, hold-day tiers |
| 2.1 | consensus/ross/sizing.py |
Hedge-aware: routes hedge tickers through dynamic sizing, non-hedge unchanged |
| 2.1 | consensus/ross/gates.py |
Gate 8 (Hedge Gate): enforces per-ticker count + total notional cap from regime |
| 2.1 | consensus/ross/executor.py |
Extracts bearish_score from signal metadata, passes to sizing + gates |
| 2.2 | trade_journal.db:broker_orders |
Added hedge_strategy TEXT column |
| 2.2 | scripts/compute_hedge_attribution.py |
NEW — per-strategy profit factor, auto-retire PF < 0.5 for 30+ days |
| 2.3 | consensus/ross/gates.py |
Gate 2.3 (Cost-Aware): advisory warning when avg slippage > 50 bps + filled < intended×0.995 |
| test | tests/test_hedge_sizing.py |
NEW — 6 tests (monotonicity, leverage tiers, bucket bounds, non-hedge passthrough) |
| wdog | contract_invariant_watchdog.py |
Added test_hedge_sizing.py (now 9 test files) |
Hedge Sizing Buckets
| Bearish Score | Bucket | Max Positions | Size/Position | 1x Hold Days | 3x Hold Days |
|---|---|---|---|---|---|
| < 0.20 | no hedge | 0 | 0% | — | — |
| 0.20–0.39 | moderate | 1 | 1.0% | 5 | 2 |
| 0.40–0.69 | heavy | 2 | 1.5% | 5 | 2 |
| ≥ 0.70 | extreme | 3 | 2.0% | 5 | 2 |
- Capped at
PARKER_HEDGE_CAP_PCTfromparameter_state(current: 0.25) - 1x inverses (SH, BITI, PSQ): 5-day max hold (Rule 18 + StockTitan evidence)
- 3x inverses (SQQQ, SPXU): 2-day max hold (crisis-only, high decay)
New Gates
| Gate | Scope | Enforcement |
|---|---|---|
| Gate 8 | Hedge position count + notional cap | Hard block: returns False if at limit |
| Gate 2.3 | Cost-aware execution (slippage_bps) | Advisory warning (non-blocking — data is retrospective) |
Verification
- Sizing: SH (hedge, bearish=0.8) → $4 on $200 equity; SH (bearish=0.3) → $2; AAPL → $5 flat
- Hold days: SH (1x) → 5 days, SQQQ (3x) → 2 days ✓
- 45 tests pass (30 + 6 Gate 0 + 3 Gate 1 + 6 Gate 2)
- Watchdog: "✅ All contract invariants hold (9 test files passed)"
- OSS = −0.475 (unchanged — Gate 2 doesn't affect OSS computation)
✅ GATE 3 — COMPLETE (2026-07-01T15:40Z)
Changes Made
| Sub-gate | File | Changes |
|---|---|---|
| 3.1a | consensus/ross/discovery.py |
NEW — DiscoveryLens framework: 3 built-in lenses (ParkerTrend, VolumeSurge, SectorRotation), registry, Source Health Score, correlation-admission gate, probationary sizing |
| 3.1b | consensus/signals.db:signals |
Added discovery_lens TEXT + source_provenance TEXT columns |
| 3.1b | trade_journal.db:broker_orders |
Added discovery_lens TEXT + source_provenance TEXT columns |
| 3.1c | consensus/ross/sizing.py |
Probationary sizing: 50% notional for tickers with <3 trades or PF < 1.0 |
| 3.1c | consensus/ross/executor.py |
Passes db_path to sizing for probationary check |
| 3.2 | scripts/update_ticker_universe.py |
NEW — Replaces static config_store.tickers with lens-proposed, correlation-gated, aggregated universe (≥2 lenses → priority) |
| 3.3 | scripts/update_ticker_universe.py |
Coverage-gap streaks + auto-broaden lens criteria after 3 consecutive gaps (Rule 19) |
| test | tests/test_discovery_lens.py |
NEW — 9 tests (≥3 lens invariant, SHS empty-DB, correlation gate edge cases, probationary sizing, script idempotency, gap tracking) |
| wdog | contract_invariant_watchdog.py |
Added test_discovery_lens.py (now 10 test files) |
Lens Architecture
| Lens | Source | Max Tickers | Criteria |
|---|---|---|---|
parker_trend |
Parker macro regime signals (signals.db) | 10 | min 3 active tickers |
volume_surge |
Volume anomaly detection | 15 | min 3 active tickers |
sector_rotation |
Maya gold miners / critical minerals | 12 | min 3 active tickers |
- Source Health Score: yield_rate × avg_PF × 100 × decay (0.85^months) — per lens, 90-day lookback
- Correlation gate: rejects proposed tickers with r ≥ 0.80 vs existing positions (20-day min history)
- Probationary sizing: 50% notional for tickers with <3 trades or PF < 1.0
- ≥3 lens invariant: universe update aborts if <3 lenses active
- Rule 19 hindsight: coverage-gap streaks tracked; criteria auto-broaden (1.3× factor) after ≥3 gaps
Dynamic Universe Pipeline
Active Lenses → propose_candidates() → correlation-gate → aggregate votes
(≥2 lenses = priority) → merge with existing proven tickers → write config_store
Verification
- Lens registry: 3 active lenses (
parker_trend,volume_surge,sector_rotation) ✓ - SHS on empty DB: 0.0 (no crash) ✓
- Pearson: 1.0 perfect, -1.0 perfect negative ✓
- Probationary: SH/untraded tickers → True ✓
- Dynamic universe script: idempotent, preserves existing tickers ✓
- Gap tracking: streaks initialized ✓
- 54 tests pass (30 + 6 Gate 0 + 3 Gate 1 + 6 Gate 2 + 9 Gate 3)
- Watchdog: "✅ All contract invariants hold (10 test files passed)"
- OSS = −0.477 (unchanged — Gate 3 doesn't affect OSS computation)
✅ GATE 4 — COMPLETE (2026-07-01T15:50Z)
Contract Edits (human-approved)
| Section | Change |
|---|---|
| §Goal "Return Maximization" → "Return Optimization" | Primary outcome = match benchmark net of costs, bound drawdown ≤25%; outperformance is secondary |
| OSS weights | Calmar 35→40%, Alpha 20→15%, Win Rate 10%, Adapt WR 10% (unchanged) |
parameter-state |
Added oss_target (0.5), oss_floor (0.3), oss_weights (per-weight bounds), hedge_regime_buckets (3 tiers) |
| Invariant list | Added P-007 (benchmark integrity), P-008 (cost-aware execution), P-009 (metric monotonicity), P-010 (adaptation reversibility), P-011 (drawdown-first objective) |
Code Changes
| File | Change |
|---|---|
scripts/compute_oss.py |
W_CALMAR 0.35→0.40, W_ALPHA 0.20→0.15 |
alpaca-trading-system.prose.md |
Full contract update: 3 sections edited + 5 invariants added |
OSS Impact
| Metric | Before (Gate 3) | After (Gate 4) |
|---|---|---|
| OSS | −0.477 | −0.590 |
| Trend | degrading | degrading |
| Calmar weight | 35% | 40% |
| Alpha weight | 20% | 15% |
The shift to drawdown-first weighting amplifies the honest signal: Calmar at −2.174 now drives the composite more heavily, producing a truer picture.
Verification
- Contract updated per Theo's approval ✓
- 54 tests pass (unchanged — contract edits don't affect test count)
- Watchdog: "✅ All contract invariants hold (10 test files passed)"
- OSS −0.590 | Calmar −2.174 | SPY −1.1% | Beta 0.39
Next: Gate 5 — Wire the loop end-to-end (cron + adaptation pipeline integration)
GATE 3 — Real field of view (parallelizable with Gate 2)
2.1 — Dynamic hedge sizing. Replace frozen MAX_HEDGE_POSITIONS=1/HEDGE_POSITION_SIZE_PCT=0.01 with regime-driven sizing bounded by PARKER_HEDGE_CAP_PCT (up to the 0.25 already authorized). Sizing reads parameter_state. Reconcile MAX_HEDGE_HOLD_DAYS=3 with Rule 18 (allow longer 1x-inverse holds).
2.2 — P-004 hedge attribution. Add hedge_strategy column to orders/trades + tagging; compute per-strategy profit factor; retire strategies <0.5. Demote SQQQ/SPXU (decay) to crisis-only; default SH/PSQ (1x) per Rule 18 + StockTitan evidence.
2.3 — P-008 cost-aware gate. New gate in gates.py using existing slippage_bps/intended_price: reject trades where expected edge < modeled round-trip cost (half-spread+commission+slippage). Directly addresses small-account cost drag.
- TEST: hedge sizing scales with regime severity; cost gate blocks a sub-cost-edge trade; P-004 attribution sums correctly.
GATE 3 — Real field of view (parallelizable with Gate 2)
3.1 — P-005/P-006. Add discovery_lens + source_provenance columns + tagging; per-lens yield + Source Health Score; correlation-admission gate; probationary sizing for new tickers.
3.2 — Dynamic universe. Replace static config_store.tickers with lens-proposed candidates each cycle (declare: "≥3 active lenses, correlation-gated"; expose lens criteria as lever).
3.3 — Rule 19 hindsight. Write coverage-gap counts; broaden lens criteria when gaps detected.
- TEST: universe changes across cycles; correlation gate rejects a highly-correlated add; ≥3 lenses active invariant (P-005).
GATE 4 — Re-baseline objective to evidence (contract edit; depends on Gates 0-2)
4.1 — Rewrite objective declaratively (P-011). Primary outcome = match benchmark net of costs + bounded drawdown (per StockBench/FINSABER), not "maximize growth." Re-weight OSS toward downside/drawdown metrics. Make hedge cap a regime-driven lever.
4.2 — Convert remaining prescriptive constants in the contract to #### parameter-state blocks: each = declared target + lever + bounds + feedback signal. The system tunes them; the contract stops fixing them.
4.3 — Add P-007/008/009/010/011 to the invariant set + watchdog.
4. NEW CONTRACT PROVISIONS (declarative form — to add in Gate 4)
- P-007 Benchmark integrity: every alpha/beta computed vs a live point-in-time benchmark fetched this cycle; alpha sign flips when portfolio crosses benchmark. (Kills 0.58 literal.)
- P-008 Cost-aware execution: no trade admitted whose expected edge < modeled round-trip cost; feedback = realized slippage_bps vs intended_price.
- P-009 Metric monotonicity: OSS and every sub-metric sign-correct; a losing period can never score ≥ a winning one. (Property-tested.)
- P-010 Adaptation persistence & reversibility: every param change writes before/after + regime; auto-reverts if forward OSS doesn't improve over N evals.
- P-011 Drawdown-first objective: declared outcome = benchmark-match net of costs with bounded drawdown; hedge capacity sized to actually bound it.
5. EXECUTION ORDER & GUARDRAILS
- Gate 0 first, always. Until OSS sees losses, every adaptation optimizes against a lie.
- Human gate before contract edits (Gate 4) per Theo's 2-agent + 1-human policy. Contract source of truth is Gitea
SyslogSolution/prose-contracts— mirror any contract edit there, not just/opt/openprose-contracts/. - Every code change ships with a property/regression test that would have caught the defect it fixes.
- Pin any new cron jobs (unpinned LLM crons break on provider switch).
- After each gate: run
compute_oss.py+ watchdog; record OSS before/after as evidence. - [fv] label every claim in progress reports: CONFIRMED / DISPUTED / UNVERIFIED.
6. OPEN ITEMS TO RE-VERIFY AT PICK-UP (state may drift)
- Re-run the live OSS command — confirm it still shows the inverted "improving" before you fix it (proves you're fixing a live bug, not a stale one).
- Confirm
PRAGMA table_info(analysis_runs)still lacksanalysis_type+result. - Confirm
.env.prodhedge caps unchanged (MAX_HEDGE_POSITIONS,HEDGE_POSITION_SIZE_PCT). - Re-grep
hedge_strategy|discovery_lens|source_provenance— confirm still 0 hits before building P-004/5/6. - Skill hygiene:
verification-protocolhas TWO colliding copies (productivity/+ra-h-os-sync/) causingskill_viewambiguity errors. Load by full path or dedupe.