Commit Graph
7 Commits
Author SHA1 Message Date
root 6253aeb72b Add check-health section with live probes to infrastructure-monitoring contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-08-22 23:01:19 +00:00
root 66f94d14fd Item 6: Auto-compaction at ~60% — update thresholds
- gpu-fleet.prose.md: Update compression.threshold to 0.60 (~77K triggers)
- Add pi compaction.reserveTokens: 52739 (≈60% of 128K)
- Remove stale 256K hardcode references
2026-08-20 07:54:36 +00:00
root 819f2f400d Item 5: Hermes context-detection fix — Rule 14
- hermes-config-template.prose.md: Add Rule 14 — Hermes reads max_model_tokens (128K), NOT max_input_tokens (64K)
- Warning: Using max_input_tokens for Hermes agents causes premature context loss
- Crewmates use max_input_tokens (64K cap), Abiba/Hermes use max_model_tokens (128K)
2026-08-20 07:53:58 +00:00
root b7778c287b Item 4: LiteLLM 3-way cap split — document context caps
- litellm-self-heal.prose.md: Add 3-way cap split (Abiba/Hermes=128K uncapped, Crew=64K via crew-auto alias)
- Note: Do NOT document max_input_tokens as context window — use max_tokens/context_window_size
2026-08-20 07:52:44 +00:00
root ab92e57781 Item 3: gpu-light vision swap to Qwen3.5-9B
- gpu-fleet.prose.md: Update gpu-light row to Qwen3.5-9B (Q5_K_M + mmproj-F16)
- gpu-fleet.prose.md: Update topology diagram, routing tables, VRAM, benchmarks
- gpu-fleet.prose.md: Fix timeout references, backward compat notes
- gpu-self-heal.prose.md: Update gpu-light row and tok/s performance
- Note: Qwen3.5-9B is multimodal (image+text), dedicated vision endpoint
- Remove gemma-4-12b references where Qwen3.5-9B now runs
2026-08-20 07:51:49 +00:00
root 05800d6ccb Item 2: Carnice Q5_K_M on strix-moe — update model row
- gpu-fleet.prose.md: Update strix-moe row to Carnice-Qwen3.6-MoE-35B-A3B (Q5_K_M, ~24.73GB)
- Note: Three-role split (gpu-dense + strix-moe = text; gpu-light = vision)
- Remove mmproj/vision claim from strix-moe (it's text-only MoE)
2026-08-20 07:49:08 +00:00
root 6195b59317 Item 1: Koby report-only — encode HARD RULE in contracts (2026-08-17)
- hermes-config-template.prose.md: Add Rule 17 — Koby is never repaired, full stop
- hermes-agent-baseline.prose.md: Document Koby report-only posture
- contract-registry.yaml: Tag all healing contracts as Koby-eligible (skip heal)
- scripts/agent-health-check.py: Mark Koby as report_only=True, skip repairs
- All healing contracts: Add report_only_agents.koby marker

Captain-approved ship via no-mistakes. PR auto-merges green.
2026-08-20 07:42:22 +00:00