no-mistakes(review): Fix monitor-key scope, Strix context, retired delegation names
This commit is contained in:
@@ -4,7 +4,7 @@ kind: responsibility
|
||||
description: >
|
||||
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
|
||||
assignments, agent context management, and prompt caching — to reduce response
|
||||
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
|
||||
times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12).
|
||||
id: 067NC6KP02RG60S50M40E30928
|
||||
---
|
||||
|
||||
@@ -63,7 +63,7 @@ prefill time at 532 tok/s. Fix context first, routing second.
|
||||
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
||||
never change between turns. Single-digit cache hit rate is unacceptable.
|
||||
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
||||
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||
GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||
|
||||
### Shape
|
||||
|
||||
|
||||
Reference in New Issue
Block a user