Root cause of the 2026-09-06 incident: 157 client-abandoned 408s against a
backend that was succeeding at 20-70s/call once clients stopped giving up.
Measured grounding: syslog-auto 28.8s avg / 25.4s TTFT over 888 calls;
nginx already allows 600s (Rule 5); the gap was entirely client-side.
Values: >=300s primary path, >=120s gpu-dense delegation, 2 retries with
15s/45s backoff, identified probes at 30s/hourly, batch jobs chunked +
day-scheduled. Verified live: Prometheus metrics, 0.56s syslog-auto probe,
gpu-fleet /health/unified topology.
agent-zero-fix-summary.prose.md has no YAML frontmatter (kind/name/description),
so CI validate fails on every master push that touches it (runs 226-228).
It is a dated session fix-log, not a contract - the durable knowledge already
lives in agent-zero-openrouter-key.prose.md (kind: function, referenced from
the summary). Rename to .md so validate/lint (which scan only *.prose.md)
stop rejecting master; matches repo root docs like cron-prompts-review.md.
Verified live: local validate repro PASS (36 files), prose-lint PASS at
baseline 15 warnings, no new warnings. Intentionally NOT changed: no content
edits, no other files, no contract frontmatter added to non-contract docs.
- Root cause: Agent Zero loaded key from .env.clobbered-by-new-image
- Not main .env file!
- Both files must be updated when changing OpenRouter key
- Added to prose contract for future reference
- run_ui process caches API keys in memory
- supervisorctl restart run_ui was insufficient
- Full container restart (docker restart agent-zero) required
- Added Agent Zero (kagentz .14) to .env fallback inventory
- Documented direct OpenRouter API access (not via LiteLLM proxy)
- Included key prefix, user ID, model config, and rotation procedure
- Cross-referenced with agent-zero-openrouter-key.prose.md
Post-merge verification of PR #54 found:
1. Lifecycle commands in zulip-health/zulip-self-heal used
'sudo systemctl' for hermes-gateway - sudo is NOT installed
on kagentz (no sudoers, no polkit user rules). Correct path:
root SSH invocation (root@192.168.68.14), matching how the Proxmox
host actually manages the unit.
2. hermes-zulip-restore.prose.md line 5 + zulip-resilience-v3 line 423
still said 'Mumuni CT100' in prose - repointed.
Found via independent review + live runtime check (whoami, which sudo,
journalctl, systemctl show). Refs PR #54, relay #738/#739.
Mumuni migrated off Abiba CT100 (192.168.68.24, purged) to dedicated
CT kagentz (192.168.68.14) on minipve. Verified live on kagentz:
- Gateway: systemd unit hermes-gateway.service (User=hermes), active
- Hermes home: /home/hermes/.hermes
- gateway_state.json present at ~/.hermes/gateway_state.json
Updates Mumuni rows/references in 9 contracts. Abiba-owned refs
(.24 pi agent, GPU monitor, infrastructure-control CRITICAL file)
and historical runs/ logs intentionally left unchanged.
Per relay #738 follow-ups. Refs #735-#738.