feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo #20
@@ -1,6 +1,10 @@
|
|||||||
---
|
---
|
||||||
name: inference-optimization
|
name: inference-optimization
|
||||||
kind: responsibility
|
kind: responsibility
|
||||||
|
description: >
|
||||||
|
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
|
||||||
|
assignments, agent context management, and prompt caching — to reduce response
|
||||||
|
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
|
||||||
id: 067NC6KP02RG60S50M40E30928
|
id: 067NC6KP02RG60S50M40E30928
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user