diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 3f36d62..a6e3d29 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -284,7 +284,7 @@ The following MUST be identical across ALL profiles: - For 128K context window: `threshold: 0.65` (fires at ~85K tokens) - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin -- `max_context_window: 131072` MUST match the model's actual capacity (128K) +- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K) - See `devops-hermes-compression` skill for full reference ### Rule 9: Compression Threshold for 128K Models