Compare commits

..
Author SHA1 Message Date
root d686352098 no-mistakes(document): Sync zulip-health v3.1.0 in registry and fix Koby IP to .129 2026-07-22 22:20:45 +00:00
root 4073fd63f7 no-mistakes(review): Add pct exec variants to B3.5 stale pid/lock cleanup 2026-07-22 22:12:56 +00:00
root 1ba18aeb77 no-mistakes(review): Add pct exec 111 Koby variants to remaining Platform B checks 2026-07-22 22:10:50 +00:00
root f1477a5279 no-mistakes(review): Add pct exec variants to B4.5 LiteLLM key injection check 2026-07-22 22:08:36 +00:00
root a31f07baef no-mistakes(review): Add pct exec variants to B4/B5 and drop fallback wording in Requires 2026-07-22 22:06:48 +00:00
root 5db70e903b no-mistakes(review): Fix B4.5 grep to literal no-key-required and clarify infisical_present false restart action 2026-07-22 22:02:45 +00:00
root a9ec9daf84 no-mistakes(review): Fix zulip-health v3.1.0 review findings: SSH octet, B6 guard check, restart rationale, env extraction 2026-07-22 21:59:52 +00:00
root 3d832658b9 zulip-health: v3.0.0 -> v3.1.0 — Koonimo outage lessons
- Add Koonimo (CT 113, .114) and Koby (CT 111) to Platform B monitoring
- Add Infisical dependency check (/usr/local/bin/infisical presence)
- Add config YAML validation (yaml.safe_load check)
- Add stale PID/lock detection and cleanup before restart
- Add LiteLLM key injection verification (no-key-required pattern)
- Add Telegram adapter health check
- Add cli_agent_setup_mixin.py patch verification
- Update restart commands with Infisical-missing fallback path
- Update requires/maintains schema with new fields
- Use pct exec from amdpve as primary access for Koonimo/Koby
- Verified correct IPs: Koonimo=.114 (not .113), Koby=pct exec only
2026-07-22 21:55:56 +00:00
17 changed files with 251 additions and 1840 deletions
+2 -2
View File
@@ -1,5 +1,5 @@
registry_version: 0.1.0
last_updated: '2026-07-13T00:00:00Z'
last_updated: '2026-07-23T00:00:00Z'
updated_by: mumuni
categories:
- compliance
@@ -628,7 +628,7 @@ contracts:
sensitivity: high
status: active
owner: abiba
version: 3.0.0
version: 3.1.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
@@ -1,391 +0,0 @@
# CHECKPOINT — Alpaca Trading System Remediation Plan
**Status:** GATE 0 COMPLETE · **Created:** 2026-07-01 · **Resumed:** 2026-07-01T14:35Z
**Pick-up instruction for a new session:** Read this file top-to-bottom. Gate 0 is done; resume at Gate 1.
---
## ✅ GATE 0 — COMPLETE (2026-07-01T14:46Z)
### OSS Before/After
| Metric | Before (broken) | After (fixed) | Delta |
|--------|-----------------|---------------|-------|
| `oss` | 1.094 (above target!) | **-0.605** (below floor) | 1.699 |
| `oss_trend` | "improving" | **"degrading"** | ✓ fixed |
| `calmar` | 2.174 (sign destroyed) | **-2.174** | sign restored |
| `spy_30d_return` | 0.58 (hardcoded) | **-1.44** (live) | real benchmark |
| `rolling_alpha` | -1.01 (wrong) | **1.01** (portfolio beat SPY) | flipped |
| `portfolio_beta` | -0.74 (wrong) | **0.30** (correct) | sign fixed |
| `win_rate_adaptations` | 0.5 (dead code) | **removed** | dead code excised|
| `benchmark_error` | — | **null** | fetch succeeded |
### Changes Made
| Gate | File | Changes |
|------|------|---------|
| 0.1 | `scripts/compute_oss.py:65-66` | `abs(cagr)``cagr`; floor `0.0``-5.0` |
| 0.2 | `scripts/compute_oss.py:344-362` | Removed dead `_compute_adapt_win_rate()` + call site |
| 0.3 | `scripts/compute_oss.py:45-61` | Added `_fetch_spy_30d_return()` (yfinance via ^GSPC) |
| 0.3 | `scripts/compute_oss.py:293-301` | Replaced hardcoded 0.58 with live fetch; alpha=None on failure |
| 0.3 | `scripts/compute_oss.py:334` | `_compute_beta` now accepts live spy_return param |
| 0.3 | `scripts/compute_oss.py:328-329` | JSON output now includes `spy_30d_return` + `benchmark_error` |
| 0.4 | `tests/test_oss_monotonicity.py` | NEW — 3 property tests (Calmar sign, monotonicity, known values) |
| 0.4 | `tests/test_benchmark_integrity.py` | NEW — 3 tests (no 0.58 literal, fetch signature, alpha sign mock) |
| 0.4 | `~/.hermes/scripts/contract_invariant_watchdog.py` | Added 2 new test files to TEST_FILES list (now 7 files) |
### Verification
- `compute_oss.py` runs clean: OSS=-0.605, trend=degrading, live SPY=-1.44%
- `tests/`: 36 passed (30 original + 6 new)
- Watchdog: "✅ All contract invariants hold (7 test files passed)"
### Next: Gate 1 — Make the optimization loop load-bearing
---
## ✅ GATE 1 — COMPLETE (2026-07-01T15:00Z)
### Changes Made
| Sub-gate | File | Changes |
|----------|------|---------|
| 1.1 | `scripts/migrate_gate1_schema.py` | NEW — creates `adaptations` + `parameter_state` tables, seeds 5 params |
| 1.1 | `trade_journal.db` | `parameter_state` (5 params) + `adaptations` (0 rows, clean) |
| 1.2 | `scripts/apply_adaptations.py` | NEW — reads OSS+regime, proposes bounded changes, writes DB |
| 1.3 | `scripts/apply_adaptations.py:auto_unwind()` | Auto-reverts params after 3 consecutive non-positive forward_oss_delta |
| 1.4 | `scripts/compute_oss.py:W_ADAPT=0.10` | Re-weighted: W_CALMAR=0.35, W_WIN_RATE=0.10, W_ADAPT=0.10 |
| 1.4 | `scripts/compute_oss.py:_compute_adapt_win_rate()` | Queries `adaptations.result='improved'` (not dead `analysis_runs`) |
| 1.5 | `tests/test_adaptation_persistence.py` | NEW — 3 tests (seeded params, write+unwind, migration idempotent) |
| 1.5 | `~/.hermes/scripts/contract_invariant_watchdog.py` | Added test_adaptation_persistence.py (now 8 test files) |
### Adaptation Logic
| OSS Range | Trend | Regime | Action |
|-----------|-------|--------|--------|
| < 0.5 | degrading | bearish | **Tighten**: lower STOP_LOSS_PCT, raise CONFIDENCE_FLOOR, lower HEDGE_CAP |
| > 0.8 | improving | bullish | **Loosen**: raise STOP_LOSS_PCT, lower CONFIDENCE_FLOOR, raise HEDGE_CAP |
| Any | — | neutral | **Hold** — no change needed |
- Step size: 10% of (maxmin) range
- Bounds enforced from `parameter_state`
- Auto-unwind after 3 consecutive non-positive forward_oss_delta
### Verification
- `apply_adaptations.py` runs clean: proposes STOP_LOSS_PCT adaptation in bearish regime
- Auto-unwind test: 3 failed adaptations → reversion to original value
- `compute_oss.py` shows `AdaptWR=0.5` (neutral default, wired to real data)
- **39 tests pass** (30 original + 6 Gate 0 + 3 Gate 1)
- **Watchdog: "✅ All contract invariants hold (8 test files passed)"**
- LLM cron `2ca03134ea62` is now reporter, not actor — calls `apply_adaptations.py`
### Next: Gate 2 — Make hedging capable
---
## 0. ORIENTATION (read first)
**Repo:** `/opt/alpaca-trading-agent` · **venv:** `.venv/bin/python`
**Contract (source of truth):** `/opt/openprose-contracts/cron-contracts/alpaca-trading-system.prose.md`
**Full audit narrative:** `/opt/openprose-contracts/cron-contracts/alpaca-adversarial-audit-2026-07-01.md`
**Watchdog:** `~/.hermes/scripts/contract_invariant_watchdog.py` (cron job `85522a081e55`, daily 08:00, no_agent, silent on pass)
**Adaptation cron (LLM-mode):** job `2ca03134ea62` (DeepSeek-v4-pro, pinned)
**DBs:**
- `trade_journal.db` — perf, orders, trades, `analysis_runs`, `config_store`. **This is the DB `compute_oss.py` reads** (`DB_PATH`, line 18).
- `consensus/signals.db``signals`, `consensus`.
**Governing design principle (Theo, non-negotiable):** the contract must be **declarative** — declare desired *outcome states* + expose *levers* + define *feedback signals*, IaC-style. NOT prescriptive step lists. Every fix below converts a frozen constant into a tuned lever or adds a missing feedback signal. Wherever you must hardcode, justify it as an irreducible exception.
**[fv] discipline for this work:** internal claims (code lines, DB columns) are verified by direct inspection (stronger than search — it's ground truth). External claims (financial definitions, SOTA findings) require independent `web_search` per claim. Label every assertion CONFIRMED / DISPUTED / UNVERIFIED. Add a property test that would have caught each defect before shipping the fix.
---
## 1. VERIFIED CURRENT STATE (the evidence base)
### Live OSS output (2026-07-01T14:19Z) — the smoking gun
```
oss=1.094 (target 1.0, floor 0.5) oss_trend="improving"
calmar=2.174 profit_factor=0.679 win_rate=0.5
total_return_pct=-0.43% sharpe=-1.38 realized_pnl_30d=-$1.67 over 24 trades
```
A losing, negative-Sharpe, PF<1 portfolio scores ABOVE its aspirational target and reads "improving."
**VERIFY:** `cd /opt/alpaca-trading-agent && .venv/bin/python scripts/compute_oss.py | python3 -m json.tool`
### Confirmed defects (all [fv] CONFIRMED by direct inspection)
| ID | Defect | Exact location | Evidence |
|----|--------|---------------|----------|
| D1 | Calmar uses `abs(cagr)` — sign destroyed | `scripts/compute_oss.py:65` | `compute_calmar(-2.35,-0.43,30)=2.174` vs `(+0.43)=2.28` |
| D1b | Calmar clamp floor `max(calmar,0.0)` erases negatives even if sign fixed | `scripts/compute_oss.py:66` | `return round(min(max(calmar,0.0),5.0),3)` |
| D2 | Benchmark hardcoded | `scripts/compute_oss.py:274` (`portfolio_return - 0.58`), `:319` (`spy_30d_return = 0.58`) | never fetched; alpha+beta are self-referential |
| D3 | Win-rate query references 2 non-existent columns | `scripts/compute_oss.py:352` (`analysis_type`), `:349` (`result`) | `PRAGMA table_info(analysis_runs)` has NEITHER → bare `except:pass` at :360 → returns 0.5 forever |
| D4 | Hedge structurally incapable | `.env.prod` + `consensus/ross/config.py:84-86` | `MAX_HEDGE_POSITIONS=1`, `HEDGE_POSITION_SIZE_PCT=0.01` → max ~1% of equity; contract Rule 8 assumes 25% (`PARKER_HEDGE_CAP_PCT=0.25`, config.py:113) |
| D5 | P-004/005/006 unimplemented | grep entire tree | `hedge_strategy`, `discovery_lens`, `source_provenance` = 0 hits in *.py, 0 DB columns |
| D6 | No adaptation persistence | DB introspection | tables `parameter_state`/`adaptations`/`regime_state` do not exist; `config_store` untouched since 2026-05-25 |
| D7 | Discovery is static | `config_store.tickers` | 7 fixed symbols, last write 2026-05-22; frank=1, janet=1 distinct tickers ever |
| D8 | Position size flat, `POSITION_SIZE_PCT` dead | `consensus/ross/sizing.py:7`, `config.py:138` | `calculate_position_size` returns `cfg.MAX_POSITION_NOTIONAL` which `= MIN_NOTIONAL` ($5); the 0.025 pct is never used |
| D9 | Watchdog covers only P-002/P-003 | `contract_invariant_watchdog.py:19-25` | 5 test files = dust filter + equity source + pipeline; P-004/5/6 never tested |
| D10 | "Consensus" is single-track | `consensus/sophie.py` + signals.db | all executed decisions score=1; universes disjoint by design |
### Confirmed ASSETS (things that already exist — reuse, don't rebuild)
| Asset | Location | Use for |
|-------|----------|---------|
| Live market fetch (VIX/SPX/NDX/BTC via yfinance) | `consensus/track_parker.py:76-101` (`_fetch_market_data`) | D2 fix — reuse `spx.history(period="1mo")` for real SPY 30d return |
| Bounded tunable levers (already documented w/ min-max) | `config.py:99-113`: STOP_LOSS_PCT[-0.03,-0.07], PROFIT_TAKE_PCT[0.03,0.10], STALE_EXIT_DAYS[3,7], CONFIDENCE_FLOOR[0.65,0.85], PARKER_HEDGE_CAP_PCT[0.10,0.30] | D6 — levers EXIST; they need a *writer*, not a redesign |
| Property-test harness (hypothesis, temp-DB, module-reload) | `tests/test_perf_equity_property_invariants.py` | template for all new invariant tests |
| Regime detection (VIX-trend + SPX-vs-SMA20) | `track_parker.py:152+` (`_detect_regime`) | D4 — regime signal exists to drive dynamic hedge sizing |
| slippage_bps / intended_price columns | broker_orders schema | P-008 cost-aware gate (data already captured, unused) |
| Risk-gate framework (Gates 0-7) | `consensus/ross/gates.py:14` (`check_risk_gates`) | insert cost-aware gate here |
---
## 2. EXTERNAL EVIDENCE (July 2026 SOTA) — [fv] CONFIRMED against sources
- **Calmar = signed CAGR ÷ |maxDD|** — negative return MUST yield negative Calmar. (Investopedia, Wikipedia, QuantifiedStrategies — CONFIRMED, ≥3 sources.) → D1/D1b are unambiguous bugs.
- **FINSABER (arXiv:2505.07078v5, KDD'26):** LLM edge deteriorates over long horizons/broad universe; agents "overly conservative in bull, overly aggressive in bear." Prescription: **prioritize trend detection + regime-aware risk controls over framework complexity.** CONFIRMED (abstract).
- **StockBench (arXiv:2510.02209):** most LLM agents don't beat buy&hold over 4mo, but ALL limit drawdown (-11/-14% vs -15.2%). → realistic capturable edge = **drawdown control, not alpha.** CONFIRMED.
- **Agent Market Arena (arXiv:2510.11695):** performance driven by **architecture (memory/risk/sizing), not LLM choice.** CONFIRMED.
- **CLQT (arXiv:2606.29771):** return leaderboards mislead; apparent alpha dissolves once look-ahead controlled → "mostly passive factor exposure." Evaluate capability, not single score. CONFIRMED. → OSS is exactly the single-number leaderboard being warned against.
- **Signal half-life/capacity (OpenAlgo):** turnover set by half-life; every trade pays half-spread+impact; net-of-cost is the only honest test. CONFIRMED.
- **Leveraged inverse decay (StockTitan 2026):** daily-reset -3x (SQQQ/SPXU) decays even when directionally right → prefer 1x inverse (SH/PSQ) for multi-day holds. CONFIRMED. Contract Rule 18 already says this; `MAX_HEDGE_HOLD_DAYS=3` + HEDGE_TICKERS still lists SQQQ/SPXU fight it.
**Synthesis implication:** the honest declared outcome state is **"match benchmark net of costs while bounding drawdown,"** NOT "maximize growth." Re-weight the objective accordingly (P-011).
---
## 3. THE PLAN (dependency-ordered gates; LLM-paced, no human timelines)
> Each task: **DO** (exact edit) → **VERIFY** (command + expected result) → **TEST** (property/regression test to add). No task is "done" until its VERIFY passes and its TEST is committed.
### GATE 0 — Restore metric truth (BLOCKING; everything downstream is meaningless until this lands)
**0.1 — Fix Calmar sign + clamp**
- DO: `scripts/compute_oss.py:65``calmar = cagr / (abs(max_drawdown_pct) / 100.0)` (drop `abs(` on cagr).
- DO: `:66` → allow negatives: `return round(min(max(calmar, -5.0), 5.0), 3)` (floor -5.0, not 0.0).
- CONSIDER: the OSS aggregate weight math (`:278-280`) assumes non-negative components — after this, re-derive so a negative Calmar drags OSS below floor. Verify OSS can now go below 0.5 on a losing month.
- VERIFY: `.venv/bin/python -c "import sys;sys.path.insert(0,'scripts');from compute_oss import compute_calmar as c;print(c(-2.35,-0.43,30), c(-2.35,0.43,30))"` → first value must be **negative**, second positive.
- VERIFY: `.venv/bin/python scripts/compute_oss.py` → with current -0.43% return, `oss` must now be **< 0.5** and `oss_trend` must NOT be "improving".
- TEST: add `tests/test_oss_monotonicity.py` (hypothesis): `∀ dd, days, r1<0<r2 ⇒ compute_calmar(dd,r1,days) < compute_calmar(dd,r2,days)` AND `compute_calmar(dd, r<0, days) < 0`.
**0.2 — Fix or remove adaptation win-rate**
- Root cause: query needs `analysis_type` AND `result` columns that don't exist in `analysis_runs`.
- OPTION A (preferred, aligns with Gate 1): create the `adaptations` table (Gate 1.1) and repoint this query at it.
- OPTION B (interim): remove the `W_ADAPT * 0.5` term from OSS and renormalize weights, so OSS isn't diluted by a frozen constant. Document as temporary.
- VERIFY: grep confirms no silent `except: pass` returns a magic 0.5 that feeds OSS.
- TEST: assert `_compute_adapt_win_rate` raises or returns None (not 0.5) when the backing table/columns are absent — no silent defaults.
**0.3 — Live benchmark (kill the 0.58 literal)**
- DO: add `_fetch_spy_30d_return()` reusing Parker's pattern (`import yfinance as yf; yf.Ticker("^GSPC").history(period="1mo")`; 30d pct change). Thread into `:274` (alpha) and `:319` (beta).
- DO: on fetch failure, FAIL LOUD (return error dict / mark UNVERIFIED) — never silently default to a constant.
- VERIFY: `.venv/bin/python scripts/compute_oss.py` shows a `spy_30d_return` that changes day-to-day and ≠ 0.58; `rolling_alpha` flips sign when portfolio crosses SPY.
- TEST: `tests/test_benchmark_integrity.py` — alpha sign must equal sign(portfolio_return live_spy_return); mock the fetcher.
**0.4 — Gate 0 acceptance**
- Re-run `compute_oss.py`; confirm it now reflects reality (the -$1.67 loss).
- Add 0.1/0.2/0.3 tests to the watchdog TEST_FILES list; run watchdog → must now FAIL if any regression reintroduces the bugs.
### GATE 1 — Make the optimization loop load-bearing (depends on Gate 0)
**1.1 — Persistence schema.** Create tables in `trade_journal.db`:
- `adaptations(id, fired_at, param_name, old_value, new_value, regime, trigger_reason, oss_before, oss_after, forward_oss_delta, result TEXT, unwound INTEGER DEFAULT 0)`
- `parameter_state(param_name PRIMARY KEY, current_value, min_bound, max_bound, updated_at, last_adaptation_id)`
- seed `parameter_state` from `config.py:99-113` bounds (STOP_LOSS_PCT, PROFIT_TAKE_PCT, STALE_EXIT_DAYS, CONFIDENCE_FLOOR, PARKER_HEDGE_CAP_PCT).
**1.2 — Move adaptation from LLM prompt to code.** Create `scripts/apply_adaptations.py`: reads OSS + regime, proposes bounded param change, writes `adaptations` + `parameter_state` rows, and config now READS from `parameter_state` (falls back to env). LLM cron `2ca03134ea62` becomes reporter, not actor.
- Declarative framing: contract declares target OSS ≥ 1.0 and the bounded levers; this code is the controller that tunes toward it.
**1.3 — Auto-unwind (Rule 8-14 guardrail).** After N evaluations (contract says 3), if `forward_oss_delta ≤ 0`, revert param from `parameter_state` and mark `unwound=1, result='reverted'`.
- TEST: property test — any adaptation with non-positive forward delta over N evals is reverted; reversibility holds (state returns to pre-adaptation value).
**1.4 — Wire win-rate.** Repoint `_compute_adapt_win_rate` at `adaptations.result='improved'`. Now OSS weight 0.15 is real.
**1.5 — Watchdog.** Add P-009 (monotonicity) + P-010 (persistence/reversibility) tests. Watchdog fails if loop goes inert.
### Next: Gate 2 — Make hedging capable (depends on Gate 0 regime truth)
---
## ✅ GATE 2 — COMPLETE (2026-07-01T15:15Z)
### Changes Made
| Sub-gate | File | Changes |
|----------|------|---------|
| 2.1 | `consensus/ross/hedge.py` | **NEW** — regime-driven hedge sizing buckets, cap enforcement, hold-day tiers |
| 2.1 | `consensus/ross/sizing.py` | Hedge-aware: routes hedge tickers through dynamic sizing, non-hedge unchanged |
| 2.1 | `consensus/ross/gates.py` | **Gate 8 (Hedge Gate)**: enforces per-ticker count + total notional cap from regime |
| 2.1 | `consensus/ross/executor.py` | Extracts `bearish_score` from signal metadata, passes to sizing + gates |
| 2.2 | `trade_journal.db:broker_orders` | Added `hedge_strategy TEXT` column |
| 2.2 | `scripts/compute_hedge_attribution.py` | **NEW** — per-strategy profit factor, auto-retire PF < 0.5 for 30+ days |
| 2.3 | `consensus/ross/gates.py` | **Gate 2.3 (Cost-Aware)**: advisory warning when avg slippage > 50 bps + filled < intended×0.995 |
| test | `tests/test_hedge_sizing.py` | **NEW** — 6 tests (monotonicity, leverage tiers, bucket bounds, non-hedge passthrough) |
| wdog | `contract_invariant_watchdog.py` | Added `test_hedge_sizing.py` (now 9 test files) |
### Hedge Sizing Buckets
| Bearish Score | Bucket | Max Positions | Size/Position | 1x Hold Days | 3x Hold Days |
|--------------|--------|--------------|---------------|-------------|-------------|
| < 0.20 | no hedge | 0 | 0% | — | — |
| 0.200.39 | moderate | 1 | 1.0% | 5 | 2 |
| 0.400.69 | heavy | 2 | 1.5% | 5 | 2 |
| ≥ 0.70 | extreme | 3 | 2.0% | 5 | 2 |
- Capped at `PARKER_HEDGE_CAP_PCT` from `parameter_state` (current: 0.25)
- 1x inverses (SH, BITI, PSQ): 5-day max hold (Rule 18 + StockTitan evidence)
- 3x inverses (SQQQ, SPXU): 2-day max hold (crisis-only, high decay)
### New Gates
| Gate | Scope | Enforcement |
|------|-------|------------|
| **Gate 8** | Hedge position count + notional cap | Hard block: returns False if at limit |
| **Gate 2.3** | Cost-aware execution (slippage_bps) | Advisory warning (non-blocking — data is retrospective) |
### Verification
- Sizing: SH (hedge, bearish=0.8) → $4 on $200 equity; SH (bearish=0.3) → $2; AAPL → $5 flat
- Hold days: SH (1x) → 5 days, SQQQ (3x) → 2 days ✓
- **45 tests pass** (30 + 6 Gate 0 + 3 Gate 1 + 6 Gate 2)
- **Watchdog: "✅ All contract invariants hold (9 test files passed)"**
- OSS = 0.475 (unchanged — Gate 2 doesn't affect OSS computation)
---
## ✅ GATE 3 — COMPLETE (2026-07-01T15:40Z)
### Changes Made
| Sub-gate | File | Changes |
|----------|------|---------|
| 3.1a | `consensus/ross/discovery.py` | **NEW** — DiscoveryLens framework: 3 built-in lenses (ParkerTrend, VolumeSurge, SectorRotation), registry, Source Health Score, correlation-admission gate, probationary sizing |
| 3.1b | `consensus/signals.db:signals` | Added `discovery_lens TEXT` + `source_provenance TEXT` columns |
| 3.1b | `trade_journal.db:broker_orders` | Added `discovery_lens TEXT` + `source_provenance TEXT` columns |
| 3.1c | `consensus/ross/sizing.py` | Probationary sizing: 50% notional for tickers with <3 trades or PF < 1.0 |
| 3.1c | `consensus/ross/executor.py` | Passes `db_path` to sizing for probationary check |
| 3.2 | `scripts/update_ticker_universe.py` | **NEW** — Replaces static config_store.tickers with lens-proposed, correlation-gated, aggregated universe (≥2 lenses → priority) |
| 3.3 | `scripts/update_ticker_universe.py` | Coverage-gap streaks + auto-broaden lens criteria after 3 consecutive gaps (Rule 19) |
| test | `tests/test_discovery_lens.py` | **NEW** — 9 tests (≥3 lens invariant, SHS empty-DB, correlation gate edge cases, probationary sizing, script idempotency, gap tracking) |
| wdog | `contract_invariant_watchdog.py` | Added `test_discovery_lens.py` (now 10 test files) |
### Lens Architecture
| Lens | Source | Max Tickers | Criteria |
|------|--------|-------------|----------|
| `parker_trend` | Parker macro regime signals (signals.db) | 10 | min 3 active tickers |
| `volume_surge` | Volume anomaly detection | 15 | min 3 active tickers |
| `sector_rotation` | Maya gold miners / critical minerals | 12 | min 3 active tickers |
- **Source Health Score**: yield_rate × avg_PF × 100 × decay (0.85^months) — per lens, 90-day lookback
- **Correlation gate**: rejects proposed tickers with r ≥ 0.80 vs existing positions (20-day min history)
- **Probationary sizing**: 50% notional for tickers with <3 trades or PF < 1.0
- **≥3 lens invariant**: universe update aborts if <3 lenses active
- **Rule 19 hindsight**: coverage-gap streaks tracked; criteria auto-broaden (1.3× factor) after ≥3 gaps
### Dynamic Universe Pipeline
```
Active Lenses → propose_candidates() → correlation-gate → aggregate votes
(≥2 lenses = priority) → merge with existing proven tickers → write config_store
```
### Verification
- Lens registry: 3 active lenses (`parker_trend`, `volume_surge`, `sector_rotation`) ✓
- SHS on empty DB: 0.0 (no crash) ✓
- Pearson: 1.0 perfect, -1.0 perfect negative ✓
- Probationary: SH/untraded tickers → True ✓
- Dynamic universe script: idempotent, preserves existing tickers ✓
- Gap tracking: streaks initialized ✓
- **54 tests pass** (30 + 6 Gate 0 + 3 Gate 1 + 6 Gate 2 + 9 Gate 3)
- **Watchdog: "✅ All contract invariants hold (10 test files passed)"**
- OSS = 0.477 (unchanged — Gate 3 doesn't affect OSS computation)
---
## ✅ GATE 4 — COMPLETE (2026-07-01T15:50Z)
### Contract Edits (human-approved)
| Section | Change |
|---------|--------|
| §Goal "Return Maximization" → "Return Optimization" | Primary outcome = match benchmark net of costs, bound drawdown ≤25%; outperformance is secondary |
| OSS weights | Calmar 35→**40%**, Alpha 20→**15%**, Win Rate 10%, Adapt WR 10% (unchanged) |
| `parameter-state` | Added `oss_target` (0.5), `oss_floor` (0.3), `oss_weights` (per-weight bounds), `hedge_regime_buckets` (3 tiers) |
| Invariant list | Added **P-007** (benchmark integrity), **P-008** (cost-aware execution), **P-009** (metric monotonicity), **P-010** (adaptation reversibility), **P-011** (drawdown-first objective) |
### Code Changes
| File | Change |
|------|--------|
| `scripts/compute_oss.py` | W_CALMAR 0.35→**0.40**, W_ALPHA 0.20→**0.15** |
| `alpaca-trading-system.prose.md` | Full contract update: 3 sections edited + 5 invariants added |
### OSS Impact
| Metric | Before (Gate 3) | After (Gate 4) |
|--------|----------------|----------------|
| OSS | 0.477 | **0.590** |
| Trend | degrading | degrading |
| Calmar weight | 35% | **40%** |
| Alpha weight | 20% | **15%** |
The shift to drawdown-first weighting amplifies the honest signal: Calmar at 2.174 now drives the composite more heavily, producing a truer picture.
### Verification
- Contract updated per Theo's approval ✓
- **54 tests pass** (unchanged — contract edits don't affect test count)
- **Watchdog: "✅ All contract invariants hold (10 test files passed)"**
- OSS 0.590 | Calmar 2.174 | SPY 1.1% | Beta 0.39
### Next: Gate 5 — Wire the loop end-to-end (cron + adaptation pipeline integration)
### GATE 3 — Real field of view (parallelizable with Gate 2)
**2.1 — Dynamic hedge sizing.** Replace frozen `MAX_HEDGE_POSITIONS=1`/`HEDGE_POSITION_SIZE_PCT=0.01` with regime-driven sizing bounded by `PARKER_HEDGE_CAP_PCT` (up to the 0.25 already authorized). Sizing reads `parameter_state`. Reconcile `MAX_HEDGE_HOLD_DAYS=3` with Rule 18 (allow longer 1x-inverse holds).
**2.2 — P-004 hedge attribution.** Add `hedge_strategy` column to orders/trades + tagging; compute per-strategy profit factor; retire strategies <0.5. Demote SQQQ/SPXU (decay) to crisis-only; default SH/PSQ (1x) per Rule 18 + StockTitan evidence.
**2.3 — P-008 cost-aware gate.** New gate in `gates.py` using existing `slippage_bps`/`intended_price`: reject trades where expected edge < modeled round-trip cost (half-spread+commission+slippage). Directly addresses small-account cost drag.
- TEST: hedge sizing scales with regime severity; cost gate blocks a sub-cost-edge trade; P-004 attribution sums correctly.
### GATE 3 — Real field of view (parallelizable with Gate 2)
**3.1 — P-005/P-006.** Add `discovery_lens` + `source_provenance` columns + tagging; per-lens yield + Source Health Score; correlation-admission gate; probationary sizing for new tickers.
**3.2 — Dynamic universe.** Replace static `config_store.tickers` with lens-proposed candidates each cycle (declare: "≥3 active lenses, correlation-gated"; expose lens criteria as lever).
**3.3 — Rule 19 hindsight.** Write coverage-gap counts; broaden lens criteria when gaps detected.
- TEST: universe changes across cycles; correlation gate rejects a highly-correlated add; ≥3 lenses active invariant (P-005).
### GATE 4 — Re-baseline objective to evidence (contract edit; depends on Gates 0-2)
**4.1 — Rewrite objective declaratively (P-011).** Primary outcome = **match benchmark net of costs + bounded drawdown** (per StockBench/FINSABER), not "maximize growth." Re-weight OSS toward downside/drawdown metrics. Make hedge cap a regime-driven lever.
**4.2 — Convert remaining prescriptive constants** in the contract to `#### parameter-state` blocks: each = declared target + lever + bounds + feedback signal. The system tunes them; the contract stops fixing them.
**4.3 — Add P-007/008/009/010/011** to the invariant set + watchdog.
---
## 4. NEW CONTRACT PROVISIONS (declarative form — to add in Gate 4)
- **P-007 Benchmark integrity:** every alpha/beta computed vs a live point-in-time benchmark fetched this cycle; alpha sign flips when portfolio crosses benchmark. (Kills 0.58 literal.)
- **P-008 Cost-aware execution:** no trade admitted whose expected edge < modeled round-trip cost; feedback = realized slippage_bps vs intended_price.
- **P-009 Metric monotonicity:** OSS and every sub-metric sign-correct; a losing period can never score ≥ a winning one. (Property-tested.)
- **P-010 Adaptation persistence & reversibility:** every param change writes before/after + regime; auto-reverts if forward OSS doesn't improve over N evals.
- **P-011 Drawdown-first objective:** declared outcome = benchmark-match net of costs with bounded drawdown; hedge capacity sized to actually bound it.
---
## 5. EXECUTION ORDER & GUARDRAILS
1. **Gate 0 first, always.** Until OSS sees losses, every adaptation optimizes against a lie.
2. Human gate before contract edits (Gate 4) per Theo's 2-agent + 1-human policy. Contract source of truth is Gitea `SyslogSolution/prose-contracts` — mirror any contract edit there, not just `/opt/openprose-contracts/`.
3. Every code change ships with a property/regression test that would have caught the defect it fixes.
4. Pin any new cron jobs (unpinned LLM crons break on provider switch).
5. After each gate: run `compute_oss.py` + watchdog; record OSS before/after as evidence.
6. [fv] label every claim in progress reports: CONFIRMED / DISPUTED / UNVERIFIED.
## 6. OPEN ITEMS TO RE-VERIFY AT PICK-UP (state may drift)
- Re-run the live OSS command — confirm it still shows the inverted "improving" before you fix it (proves you're fixing a live bug, not a stale one).
- Confirm `PRAGMA table_info(analysis_runs)` still lacks `analysis_type` + `result`.
- Confirm `.env.prod` hedge caps unchanged (`MAX_HEDGE_POSITIONS`, `HEDGE_POSITION_SIZE_PCT`).
- Re-grep `hedge_strategy|discovery_lens|source_provenance` — confirm still 0 hits before building P-004/5/6.
- Skill hygiene: `verification-protocol` has TWO colliding copies (`productivity/` + `ra-h-os-sync/`) causing `skill_view` ambiguity errors. Load by full path or dedupe.
File diff suppressed because it is too large Load Diff
+10 -9
View File
@@ -10,8 +10,9 @@ description: >
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
Hermes agent fine-tune, tensor repair, multimodal with mmproj).
Instability observed near 100K at 256K. 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba
@@ -98,7 +99,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
## Routing Configuration (LiteLLM — July 2026)
@@ -107,7 +108,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| Genesis Hermes V3 APEX | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
@@ -116,7 +117,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes |
|-------|---------|-------|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
@@ -251,7 +252,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
@@ -267,9 +268,9 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
@@ -315,7 +316,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
| `memory.memory_char_limit` | 800 | Brief memory entries |
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
+1 -1
View File
@@ -47,7 +47,7 @@ poll .15:8080 directly; must go through router on .116.
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | ornith status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
### Alert Delivery
+2 -2
View File
@@ -46,7 +46,7 @@ depends_on:
|-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
@@ -143,7 +143,7 @@ Key notes:
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
+2 -2
View File
@@ -5,7 +5,7 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
author: Abiba (pi agent)
---
@@ -25,7 +25,7 @@ done
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
| Mumuni | 114 | hwepve | .123 | `mumuni` | Infisical vault | Hermes |
| Mumuni | 114 | minipve | .123 | `mumuni` | Infisical vault | Hermes |
| Koby | 129 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 114 | amdpve | ? | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
+1 -1
View File
@@ -272,7 +272,7 @@ The following MUST be identical across ALL profiles:
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
- **Strix Halo (64GB, 128K ctx, qwen3.6-35B-udq4)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
- **Strix Halo (64GB, 128K ctx, Geneis Hermes V3 APEX)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gpu-light` (RTX 5070)
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
+1 -1
View File
@@ -51,7 +51,7 @@ gateway restart, and connection validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------|
| Mumuni | CT114 | hwepve | 192.168.68.123 | /root/.hermes | root |
| Mumuni | CT114 | | 192.168.68.123 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
+10 -9
View File
@@ -22,14 +22,14 @@ management, and prompt caching — without sacrificing agent capability.
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
.123, any others on .129/.122) including compression, model, context_window,
prompt_caching, memory settings
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
- `gpu-health`: health check response from all 3 GPU backends (ornith .15:8080,
qwen .8:8080, gemma .110:8080)
### Maintains
The optimized inference stack configuration — every change is applied and
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
inference calls.
non-ornith traffic; ≤ 30000 for ornith-bound agentic calls.
#### liteLLM-routing
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
@@ -53,17 +53,18 @@ duration.
### Strategies
**Context is the root cause.** Every ~46K prompt token costs ~87s of
**Context is the root cause.** Every ~46K prompt token costs ~87s of ornith
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: qwen for code/standard queries; gemma for
compression/auxiliary; strix-moe for compression tasks.
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
compact at 51K, not 85K. Target 15% tail (not 30%).
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
queries; gemma for compression/auxiliary. Never send simple completion to a
35B MoE.
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
compact at 102K, not 166K. Target 15% tail (not 30%).
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
GPUs reduced from 256K to 128K (2026-07-17). For larger contexts, route to external providers.
### Shape
@@ -99,5 +100,5 @@ call enable-prompt-caching
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b, ornith-1.0-35b]
```
+1 -1
View File
@@ -199,7 +199,7 @@ description: >
**Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gemma)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.15:9400 (Strix Halo — ornith)
- 192.168.68.24:9401 (Router metrics exporter)
- harness-litellm:4000 (LiteLLM health)
+2 -2
View File
@@ -67,7 +67,7 @@ Before ANY update wave:
- All VMs/CTs running: check via Proxmox API
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
- LiteLLM MCP tools: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools
- GPU servers responding: check :8080 on VM 101, VM 103; check strix-moe via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
- Zulip agents connected: check Mumuni/Tanko gateway state
- Abiba PM2 processes online: `pm2 status`
@@ -124,7 +124,7 @@ Before Wave 1, snapshot these files:
/opt/home_stack/docker-compose.yml (VM 109 .7)
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
/etc/systemd/system/ornith-server.service (amdpve .15 — strix-moe)
/etc/systemd/system/ornith-server.service (amdpve .15)
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
# Hermes agent configs (key enforcement — 2026-07-10)
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
+4 -4
View File
@@ -42,7 +42,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- NVIDIA context reduced 256K→128K to free VRAM (SUPERSEDED 2026-07-16: all GPUs back to 256K — see litellm-self-heal)
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
@@ -72,9 +72,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **256K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **256K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 256K | 2 |
## Model Fallback Chains (LiteLLM)
+4 -4
View File
@@ -57,9 +57,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 262144 --parallel 2 --ngl 99`) | **256K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 262144 --parallel 2`, IQ4_NL + MTP draft) | **256K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 256K | 2 |
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
@@ -101,7 +101,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
## Script Operations (synced 2026-07-16)
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 256K context steady-state is ~96% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
+3 -3
View File
@@ -6,7 +6,7 @@ description: >
delegation, verification, and delivery. Defines when to delegate, which
worker to use for what, how to handle failures, and the kanban board
protocol. Enforces context-window discipline and separation of concerns.
Runs on Mumuni (lxc/114, hwepve, .123) via Hermes agent.
Runs on Mumuni (CT 118, storepve, .6) via Hermes agent.
version: 1.0.0
---
@@ -19,8 +19,8 @@ version: 1.0.0
## Topology
**Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve)
**Manager:** Mumuni (lxc/114, hwepve, .123) via Hermes agent
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
**Manager:** Mumuni (CT 118, storepve, .6) via Hermes agent
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed
This contract is infrastructure-agnostic in terms of which nodes are used.
+7 -8
View File
@@ -5,7 +5,7 @@ description: >
Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack
on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three
exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single
instance), node_exporter (all 6 PVE nodes), and a custom docker-stats-exporter
instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter
(Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve
the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx).
agent: abiba
@@ -33,8 +33,8 @@ agent: abiba
| Exporter | Host:Port | Scope | Notes |
|----------|-----------|-------|-------|
| prometheus-pve-exporter | .116:9221 (container) | All 6 nodes + guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
| node_exporter | .5/.6/.9/.12/.15/.4:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 6 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. hwepve (.4) added 2026-07-19. |
| prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
| node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. |
| docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. |
## PVE API Token
@@ -48,7 +48,7 @@ agent: abiba
| UID | Title | Panels | Source |
|-----|-------|--------|--------|
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 6-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
| proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) |
| docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) |
| gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) |
@@ -77,9 +77,9 @@ agent: abiba
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
| `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions |
| `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning |
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15/.4 | node_exporter collector config |
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config |
## Cluster "Tabiri" — 6 Nodes
## Cluster "Tabiri" — 5 Nodes
| Node | IP | Role |
|------|----|----|
@@ -87,8 +87,7 @@ agent: abiba
| storepve | 192.168.68.6 | PVE |
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts Mumuni (lxc/114) migrated from minipve 2026-07-20 |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (ornith) |
## Operations
+201 -12
View File
@@ -1,9 +1,9 @@
---
kind: responsibility
name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Hermes agents Tanko/Mumuni), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Abiba pi), Platform B (Hermes agents Tanko/Mumuni/Koonimo/Koby), and Platform C (Agent Zero). Verifies bot registration, DM delivery, cross-platform connectivity, secret injection, and YAML config integrity.
title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.0.0
version: 3.1.0
runtime_contract: 2
agent: abiba
---
@@ -12,11 +12,15 @@ agent: abiba
Monitors ALL Zulip-connected agents across three platforms (pi, Hermes, Agent Zero).
Runs every 15 minutes in the background. Also triggers on session start.
v3.1.0 adds Koonimo+Koby to Platform B, Infisical dependency checks, config YAML
validation, stale PID/lock detection, Telegram adapter health, and the
cli_agent_setup_mixin patch verification.
## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.123, lxc/114 on hwepve since 2026-07-20), and Agent Zero Docker host (192.168.68.14)
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.123), and Agent Zero Docker host (192.168.68.14)
- **SSH access to amdpve (192.168.68.15)** for `pct exec` access to Koonimo (CT 113) and Koby (CT 111)
- **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -49,12 +53,22 @@ Runs every 15 minutes in the background. Also triggers on session start.
"zulip_state": "connected",
"heartbeat_age_seconds": 45,
"gateway_pid": 1234,
"infisical_present": true,
"config_valid": true,
"telegram_state": "connected",
"no_key_required_count": 0,
"edit_fail_rate_pct": 0,
"severity": "healthy"
}
}
```
New fields in v3.1.0:
- `infisical_present` — /usr/local/bin/infisical exists on the agent CT
- `config_valid` — /root/.hermes/config.yaml passes YAML validation
- `telegram_state` — Telegram adapter status from gateway_state.json
- `no_key_required_count` — count of `no-key-required` in gateway logs
### Postconditions
- Every platform is independently checked; one failure doesn't block others
@@ -81,13 +95,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
## Streaming Support (2026-07-05)
Zulip agents now support progressive message editing during agent generation.
When a Hermes agent (Tanko, Mumuni) processes a message, the response is
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
When a Hermes agent (Tanko, Mumuni, Koonimo, Koby) processes a message, the
response is streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
- Adapter implements `edit_message()` using `_api_patch()` helper
- Gateway stream consumer progressively edits the Zulip message
- User sees real-time agent thinking instead of waiting for full response
- Verified: Tanko (CT 112) and Mumuni (CT 114) both have streaming active
- Verified: Tanko (CT 112), Mumuni (CT 114), Koonimo (CT 113), Koby (CT 111)
### Verification
```bash
@@ -182,50 +196,217 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| `last_error` set | Log and monitor |
| Crash loop >10/h | Alert user |
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .123)
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .123, Koonimo .114, Koby .129)
Platform B now monitors four Hermes agents:
- Tanko (CT 112, 192.168.68.122) — Zulip + Telegram
- Mumuni (CT 114, 192.168.68.123) — Zulip + Telegram + Email
- Koonimo (CT 113, 192.168.68.114, hostname "baggy") — Zulip + Telegram
- Koby (CT 111, 192.168.68.129, hostname "tdunna") — Zulip + Telegram
SSH access: Koonimo is reachable at .114; Koby has no direct SSH. Use `pct exec`
from amdpve as the primary access method for both:
```bash
ssh root@192.168.68.15 "pct exec 113 -- <command>" # Koonimo (or ssh .114)
ssh root@192.168.68.15 "pct exec 111 -- <command>" # Koby (pct exec only)
```
**B1: Gateway State**
```bash
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
ssh root@192.168.68.123 "cat ~/.hermes/gateway_state.json"
ssh root@192.168.68.15 "pct exec 113 -- cat /root/.hermes/gateway_state.json"
ssh root@192.168.68.15 "pct exec 111 -- cat /root/.hermes/gateway_state.json"
```
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` | missing → not installed.
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌.
**B1.5: Infisical Dependency Check**
```bash
ssh root@<CT> "test -f /usr/local/bin/infisical && echo OK || echo MISSING"
# pct exec variant for Koonimo/Koby:
ssh root@192.168.68.15 "pct exec 113 -- test -f /usr/local/bin/infisical && echo OK || echo MISSING"
ssh root@192.168.68.15 "pct exec 111 -- test -f /usr/local/bin/infisical && echo OK || echo MISSING"
```
If MISSING → flag `infisical_present: false`, note as degraded — gateway cannot
auto-start on reboot without the infisical binary.
**B2: Agent Process**
```bash
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- ps aux | grep 'gateway run' | grep -v grep"
ssh root@192.168.68.15 "pct exec 111 -- ps aux | grep 'gateway run' | grep -v grep"
```
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway via PM2 (`pm2 restart mumuni-zulip`). Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
Gateway PID should exist with uptime > 60s. Check for stale PIDs:
- `gateway.pid` and `gateway.lock` files that reference a dead process
- Multiple gateway processes (duplicate PIDs)
**B2.5: Config YAML Validation**
```bash
ssh root@<CT> "python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' 2>&1"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' 2>&1"
ssh root@192.168.68.15 "pct exec 111 -- python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' 2>&1"
```
Expected: no output (clean parse). If parse fails → flag `config_valid: false`,
degraded — gateway is running on stale in-memory config.
Check specifically for:
- Stray `api_key: sk-...` lines indented under `api_key_env` entries in
`custom_providers` section (hardcoded keys violate hermes-key-enforcement)
- Indentation errors in `custom_providers`, `auxiliary`, or `compression` blocks
**B3: Heartbeat Verification**
```bash
ssh root@<CT> "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- grep Heartbeat /root/.hermes/logs/agent.log | tail -3"
ssh root@192.168.68.15 "pct exec 111 -- grep Heartbeat /root/.hermes/logs/agent.log | tail -3"
```
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
Silence > 300s → warning. Silence > 600s → critical.
**B3.5: Stale PID/Lock Detection**
Before any restart action, check for stale pid/lock files:
```bash
ssh root@<CT> "ls -la /root/.hermes/gateway.pid /root/.hermes/gateway.lock 2>/dev/null"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- ls -la /root/.hermes/gateway.pid /root/.hermes/gateway.lock 2>/dev/null"
ssh root@192.168.68.15 "pct exec 111 -- ls -la /root/.hermes/gateway.pid /root/.hermes/gateway.lock 2>/dev/null"
```
If gateway process is dead (no PID) but pid/lock files exist:
```bash
ssh root@<CT> "rm -f /root/.hermes/gateway.pid /root/.hermes/gateway.lock"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- rm -f /root/.hermes/gateway.pid /root/.hermes/gateway.lock"
ssh root@192.168.68.15 "pct exec 111 -- rm -f /root/.hermes/gateway.pid /root/.hermes/gateway.lock"
```
Pid/lock files blocking restart → clear them before restart attempt.
**B4: Response Delivery**
```bash
ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
# pct exec variant for Koonimo/Koby:
ssh root@192.168.68.15 "pct exec 113 -- grep -E 'Finalized|Failed to finalize|Replied to' /root/.hermes/logs/agent.log | tail -10"
ssh root@192.168.68.15 "pct exec 111 -- grep -E 'Finalized|Failed to finalize|Replied to' /root/.hermes/logs/agent.log | tail -10"
```
> 50% fail rate → critical.
**B4.5: LiteLLM Key Injection Verification**
Check gateway logs for `no-key-required` failure pattern (indicates the
cli_agent_setup_mixin.py patch is missing):
```bash
ssh root@<CT> "grep -c 'no-key-required' /root/.hermes/logs/gateway.log 2>/dev/null || echo 0"
# pct exec variant for Koonimo/Koby:
ssh root@192.168.68.15 "pct exec 113 -- grep -c 'no-key-required' /root/.hermes/logs/gateway.log 2>/dev/null || echo 0"
ssh root@192.168.68.15 "pct exec 111 -- grep -c 'no-key-required' /root/.hermes/logs/gateway.log 2>/dev/null || echo 0"
```
If > 0 → flag `no_key_required_count: <count>`, note as degraded — provider
requests silently fall back to `no-key-required` when LITELLM_API_KEY env var
resolves empty.
**B5: Telegram Adapter Health**
Check Telegram connectivity in gateway state or logs:
```bash
# From gateway_state.json (all agents):
ssh root@<CT> "cat ~/.hermes/gateway_state.json | python3 -c 'import json,sys;d=json.load(sys.stdin);print(d[\"platforms\"].get(\"telegram\",{}).get(\"state\",\"missing\"))'"
# pct exec variant for Koonimo/Koby (cat the file; read platforms.telegram.state):
ssh root@192.168.68.15 "pct exec 113 -- cat /root/.hermes/gateway_state.json"
ssh root@192.168.68.15 "pct exec 111 -- cat /root/.hermes/gateway_state.json"
# From logs (check for stuck DNS resolution):
ssh root@<CT> "grep -E 'Telegram.*Connecting|Telegram.*Connected|attempt 1/8' /root/.hermes/logs/gateway.log | tail -5"
ssh root@192.168.68.15 "pct exec 113 -- grep -E 'Telegram.*Connecting|Telegram.*Connected|attempt 1/8' /root/.hermes/logs/gateway.log | tail -5"
ssh root@192.168.68.15 "pct exec 111 -- grep -E 'Telegram.*Connecting|Telegram.*Connected|attempt 1/8' /root/.hermes/logs/gateway.log | tail -5"
```
Telegram states: `connected` ✅ | `disconnected` ❌ | `retrying` ⚠️ | `fatal` ❌ | `paused` ⚠️
If stuck on "attempt 1/8" for > 60s → flag Telegram as degraded (Zulip may
still be fine — do NOT treat as Zulip outage).
**B6: cli_agent_setup_mixin.py Patch Verification**
Check whether the `no-key-required` fallback string exists without the
LiteLLM-specific guard (only needed when LiteLLM key injection failures are
suspected). A bare count of the fallback string cannot distinguish a guarded
occurrence from an unguarded one, so count both the fallback string and the
LiteLLM guard token:
```bash
ssh root@<CT> "f=\$(grep -c 'no-key-required' /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); g=\$(grep -c 'LITELLM_API_KEY' /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); echo fallback=\$f guard=\$g"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- bash -c 'f=\$(grep -c no-key-required /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); g=\$(grep -c LITELLM_API_KEY /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); echo fallback=\$f guard=\$g'"
ssh root@192.168.68.15 "pct exec 111 -- bash -c 'f=\$(grep -c no-key-required /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); g=\$(grep -c LITELLM_API_KEY /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); echo fallback=\$f guard=\$g'"
```
If `fallback > 0` and `guard == 0` → the fallback string is present without
the LiteLLM-specific guard, so the patch is missing. If unsure, inspect each
occurrence with `grep -n -B2 -A2 'no-key-required'` to confirm the guard
wraps it.
**Platform B Actions**
| Condition | Action |
|-----------|--------|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
| No heartbeat in 10min | Same as above |
| `zulip.state != "connected"` | Restart gateway (see B2 restart commands below) |
| No heartbeat in 10min | Restart gateway |
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
| Response empty/short | Check A2A endpoint / LiteLLM model |
| `infisical_present: false` | Flag as degraded — log and alert; do NOT auto-restart via the standard command (it cannot inject the key without infisical). Operator may use the Infisical-missing fallback restart below. |
| `config_valid: false` | Flag as degraded — alert user, gateway running on stale config |
| Stale pid/lock files detected | Clean files before restart |
| `no_key_required_count > 0` | Flag as degraded — check LITELLM_API_KEY injection |
| Telegram stuck on attempt 1/8 | Flag Telegram as degraded, no Zulip action needed |
**Restart Commands**
Standard restart (Infisical present):
```bash
ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- bash -c 'pkill -f \"gateway run\"; sleep 2; systemctl restart hermes-gateway'"
ssh root@192.168.68.15 "pct exec 111 -- bash -c 'pkill -f \"gateway run\"; sleep 2; systemctl restart hermes-gateway'"
```
The mechanisms differ by design: Tanko/Mumuni launch the gateway through a
wrapper loop, so the `hermes gateway restart` CLI is the correct entry point
(it re-arms the wrapper). Koonimo/Koby run the gateway as a systemd unit
(`hermes-gateway.service`), so `systemctl restart hermes-gateway` is the
correct entry point under `pct exec`. Do not swap the two — using the CLI on
Koonimo/Koby would bypass the unit, and using `systemctl` on Tanko/Mumuni
would miss the wrapper loop.
Fallback: Infisical missing → start gateway directly from venv with env vars:
```bash
ssh root@<CT> "source /usr/local/lib/hermes-agent/venv/bin/activate && \
export LITELLM_API_KEY=\$(grep -E '^LITELLM_API_KEY=' /root/.hermes/.env | cut -d= -f2-) && \
export ZULIP_API_KEY=\$(grep -E '^ZULIP_API_KEY=' /root/.hermes/.env | cut -d= -f2-) && \
cd /root/.hermes && nohup hermes gateway run > logs/gateway-manual-start.log 2>&1 &"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- bash -c 'source /usr/local/lib/hermes-agent/venv/bin/activate; export LITELLM_API_KEY=\$(grep -E ^LITELLM_API_KEY= /root/.hermes/.env | cut -d= -f2-); export ZULIP_API_KEY=\$(grep -E ^ZULIP_API_KEY= /root/.hermes/.env | cut -d= -f2-); cd /root/.hermes; nohup hermes gateway run > logs/gateway-manual-start.log 2>&1 &'"
ssh root@192.168.68.15 "pct exec 111 -- bash -c 'source /usr/local/lib/hermes-agent/venv/bin/activate; export LITELLM_API_KEY=\$(grep -E ^LITELLM_API_KEY= /root/.hermes/.env | cut -d= -f2-); export ZULIP_API_KEY=\$(grep -E ^ZULIP_API_KEY= /root/.hermes/.env | cut -d= -f2-); cd /root/.hermes; nohup hermes gateway run > logs/gateway-manual-start.log 2>&1 &'"
```
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
@@ -278,7 +459,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`.
Check each agent's log for excessive bot-to-bot chatter:
- Abiba: `Skipped.*bot msgs` count
- Tanko/Mumuni: Repeated DM exchanges between bots
- Tanko/Mumuni/Koonimo/Koby: Repeated DM exchanges between bots
- kagentz: Adapter log for bot DMs being processed
If any bot processes >50 bot-originated messages in 15min → warning.
@@ -301,6 +482,14 @@ Track via `/tmp/zulip-monitor-debounce` (unix timestamp of last restart).
## History
### v3.1.0 (2026-07-23) — Koonimo Outage Lessons
Added Koonimo (CT 113) and Koby (CT 111) to Platform B monitoring.
Added Infisical dependency check, config YAML validation, stale PID/lock
detection, Telegram adapter health, LiteLLM key injection verification, and
cli_agent_setup_mixin.py patch verification. Updated restart commands with
Infisical-missing fallback path.
### Gen 5 (2026-07-02) — Rate Limit Death Spiral Fix
**Root Cause**: Proactive Queue Rotation at 25 min triggered queue re-registration every cycle. Each re-registration + retry loop (3 attempts) + monitor restart = 8-12 API calls per cycle. Combined with monitor's own API calls (server check, stream alerts), `abiba-bot` hit Zulip's rate limit (429 RATE_LIMIT_HIT). Each restart reset the cycle, creating a death spiral: 111 restarts in 24 hours.