fix(audit): reject retired aliases; sweep gpu-light/gemma-4-12b to gpu-vision
audit-hermes-config.py Rule 8 required auxiliary.vision.model and auxiliary.web_extract.model to equal the retired 'gpu-light', so a config adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was enforcing a dead alias (400 Invalid model name). Rule 8 now requires gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set; the guidance message names the live aliases. Sweep of the remaining references: gpu-self-heal stops canonicalizing gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement, inference-optimization, litellm-client-timeouts and gpu-fleet now use the live gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now points at CT 116 /opt/inference-harness/litellm_config.yaml instead of duplicating it. koby's .129 config is report-only and recorded, not edited. Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts gpu-vision passes while gpu-light and gemma-4-12b fail.
This commit is contained in:
@@ -28,7 +28,7 @@ description: >
|
||||
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
||||
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
|
||||
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
||||
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
|
||||
| gpu-vision | 2.6s | — | RTX 5070, healthy |
|
||||
|
||||
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
||||
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
||||
@@ -53,7 +53,7 @@ proxy queuing.
|
||||
|
||||
### 2. Auxiliary tasks — keep template timeouts, one correction
|
||||
|
||||
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
|
||||
- vision: 60s (keep), web_extract: 30s (keep) — gpu-vision averages 2.6s;
|
||||
these are fine.
|
||||
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
||||
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
||||
|
||||
Reference in New Issue
Block a user