audit-hermes-config.py Rule 8 required auxiliary.vision.model == "gpu-light" and auxiliary.web_extract.model == "gpu-light". gpu-light (and its raw predecessor gemma-4-12b) were retired on 2026-09-12 and return 400 Invalid model name; the live RTX 5070 canonical alias is gpu-vision. So a config that adopted the correct alias failed our own audit — the audit was actively pushing configs toward a dead alias.
Proven by executing the real CLI (before → after):
Model-bearing fields are now derived, not enumerated
Five review rounds each found one more field the hand-written list had missed (custom_providers, then fallback_providers, then x_search). Rather than patch a sixth, the audit now derives the values to check: a recursive walk of the loaded config for every mapping key named model / model_name at any depth, including inside lists, plus a small documented allowlist for the exceptions where a model name sits under a different key (models: lists). The derivation rule is documented in the script so the next person extends the allowlist instead of adding a field by hand.
Fail vs warn — the owner's policy call
Verified through live LiteLLM: qwen3.6-27B-code and qwen3.6-35B-udq4 resolve (200) — live, discouraged but working; gpu-light, gemma-4-12b, crew-auto, ornith-1.0-35b do not resolve. So the audit hard-fails the non-resolving set (naming the live replacement) and warns on the resolving set. The audit's job is to catch broken configs, not to enforce a style preference — failing a working alias would reject valid configs, which is the defect this change exists to fix. Rationale recorded in the script comments.
Fixture results (real audit CLI)
Fixture
Result
gpu-vision config
exit 0 — PASS
gpu-light in delegation.model
exit 1 — FAIL
gpu-light under custom_providers
exit 1 — FAIL
gpu-light under fallback_providers
exit 1 — FAIL
gpu-light in a nested auxiliary block
exit 1 — FAIL
qwen3.6-27B-code (raw but live) in delegation.model
exit 0 — PASS with a warning naming gpu-dense
tests/test_audit_hermes_config_alias.py covers these by executing the CLI (10 passed). Verified it is a genuine regression: against the pre-change audit the gpu-vision fixture exits 1.
Sweep and corrections
gpu-self-heal stops canonicalizing gpu-light; weights replaced by the CT 116 authority pointer.
hermes-config-template, hermes-agent-baseline, hermes-key-enforcement, inference-optimization, gpu-fleet, gpu-monitor, infrastructure-control now use the live gpu-vision alias; the gpu-monitor diagram's swapped .8/.110 labels are corrected.
The Mumuni reference profile's compression entries are aligned to syslog-auto (Rule 7) so a copied profile passes the audit.
The dated latency table keeps the measured name (gemma-4-12b) with a "retired" annotation rather than being re-attributed to an alias that did not produce the measurement.
The key-duration claim is corrected in both hermes-key-enforcement and litellm-api-keys: CT 116 has nodefault_key_generate_params block, so no default lifetime enforcement exists today. Whether agent keys should expire by default is an open policy question for the captain; no block was added.
koby's .129 config still names gpu-light/gemma-4-E4B; .129 is report-only, so it is recorded for its owner and not edited.
Validation
no-mistakes run 01M2B67C734FF4DZ87V22VB4SD — outcome passed after several review rounds.
## The functional break (fixed first)
`audit-hermes-config.py` Rule 8 required `auxiliary.vision.model == "gpu-light"` and `auxiliary.web_extract.model == "gpu-light"`. `gpu-light` (and its raw predecessor `gemma-4-12b`) were retired on 2026-09-12 and return 400 `Invalid model name`; the live RTX 5070 canonical alias is `gpu-vision`. So a config that adopted the *correct* alias **failed our own audit** — the audit was actively pushing configs toward a dead alias.
Proven by executing the real CLI (before → after):
```
BEFORE: gpu-vision config -> FAIL (Rule 8) gpu-light config -> PASS
AFTER: gpu-vision config -> PASS (exit 0) gpu-light config -> FAIL (exit 1)
```
## Model-bearing fields are now derived, not enumerated
Five review rounds each found one more field the hand-written list had missed (`custom_providers`, then `fallback_providers`, then `x_search`). Rather than patch a sixth, the audit now **derives** the values to check: a recursive walk of the loaded config for every mapping key named `model` / `model_name` at any depth, including inside lists, plus a small documented allowlist for the exceptions where a model name sits under a different key (`models:` lists). The derivation rule is documented in the script so the next person extends the allowlist instead of adding a field by hand.
## Fail vs warn — the owner's policy call
Verified through live LiteLLM: `qwen3.6-27B-code` and `qwen3.6-35B-udq4` resolve (200) — live, discouraged but working; `gpu-light`, `gemma-4-12b`, `crew-auto`, `ornith-1.0-35b` do not resolve. So the audit **hard-fails the non-resolving set** (naming the live replacement) and **warns** on the resolving set. The audit's job is to catch broken configs, not to enforce a style preference — failing a working alias would reject valid configs, which is the defect this change exists to fix. Rationale recorded in the script comments.
## Fixture results (real audit CLI)
| Fixture | Result |
|---|---|
| gpu-vision config | **exit 0 — PASS** |
| gpu-light in `delegation.model` | **exit 1 — FAIL** |
| gpu-light under `custom_providers` | **exit 1 — FAIL** |
| gpu-light under `fallback_providers` | **exit 1 — FAIL** |
| gpu-light in a nested `auxiliary` block | **exit 1 — FAIL** |
| `qwen3.6-27B-code` (raw but live) in `delegation.model` | **exit 0 — PASS with a warning** naming `gpu-dense` |
`tests/test_audit_hermes_config_alias.py` covers these by executing the CLI (10 passed). Verified it is a genuine regression: against the pre-change audit the gpu-vision fixture exits 1.
## Sweep and corrections
- `gpu-self-heal` stops canonicalizing `gpu-light`; weights replaced by the CT 116 authority pointer.
- `hermes-config-template`, `hermes-agent-baseline`, `hermes-key-enforcement`, `inference-optimization`, `gpu-fleet`, `gpu-monitor`, `infrastructure-control` now use the live `gpu-vision` alias; the `gpu-monitor` diagram's swapped `.8`/`.110` labels are corrected.
- The Mumuni reference profile's compression entries are aligned to `syslog-auto` (Rule 7) so a copied profile passes the audit.
- The dated latency table keeps the **measured** name (`gemma-4-12b`) with a "retired" annotation rather than being re-attributed to an alias that did not produce the measurement.
- The key-duration claim is corrected in both `hermes-key-enforcement` and `litellm-api-keys`: CT 116 has **no** `default_key_generate_params` block, so no default lifetime enforcement exists today. Whether agent keys should expire by default is an **open policy question** for the captain; no block was added.
- koby's `.129` config still names `gpu-light`/`gemma-4-E4B`; `.129` is report-only, so it is recorded for its owner and not edited.
## Validation
no-mistakes run `01M2B67C734FF4DZ87V22VB4SD` — outcome **passed** after several review rounds.
audit-hermes-config.py Rule 8 required auxiliary.vision.model and
auxiliary.web_extract.model to equal the retired 'gpu-light', so a config
adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was
enforcing a dead alias (400 Invalid model name). Rule 8 now requires
gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set;
the guidance message names the live aliases.
Sweep of the remaining references: gpu-self-heal stops canonicalizing
gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement,
inference-optimization, litellm-client-timeouts and gpu-fleet now use the live
gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now
points at CT 116 /opt/inference-harness/litellm_config.yaml instead of
duplicating it. koby's .129 config is report-only and recorded, not edited.
Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts
gpu-vision passes while gpu-light and gemma-4-12b fail.
The validate job runs with bash -e -o pipefail. `echo "$FM" | grep -q '^name:'`
lets grep exit on first match, which can SIGPIPE the echo; pipefail then reports
the pipeline non-zero and the || branch raises a false "Missing name/description".
The flagged file set varied run to run (and included files untouched by the PR)
while a fresh clone of the same commit passes the identical check. Reproduced:
the old form failed 3 of 5 local runs under the same shell flags, the herestring
form passed 5 of 5. Use herestrings so no pipe can be broken.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
The functional break (fixed first)
audit-hermes-config.pyRule 8 requiredauxiliary.vision.model == "gpu-light"andauxiliary.web_extract.model == "gpu-light".gpu-light(and its raw predecessorgemma-4-12b) were retired on 2026-09-12 and return 400Invalid model name; the live RTX 5070 canonical alias isgpu-vision. So a config that adopted the correct alias failed our own audit — the audit was actively pushing configs toward a dead alias.Proven by executing the real CLI (before → after):
Model-bearing fields are now derived, not enumerated
Five review rounds each found one more field the hand-written list had missed (
custom_providers, thenfallback_providers, thenx_search). Rather than patch a sixth, the audit now derives the values to check: a recursive walk of the loaded config for every mapping key namedmodel/model_nameat any depth, including inside lists, plus a small documented allowlist for the exceptions where a model name sits under a different key (models:lists). The derivation rule is documented in the script so the next person extends the allowlist instead of adding a field by hand.Fail vs warn — the owner's policy call
Verified through live LiteLLM:
qwen3.6-27B-codeandqwen3.6-35B-udq4resolve (200) — live, discouraged but working;gpu-light,gemma-4-12b,crew-auto,ornith-1.0-35bdo not resolve. So the audit hard-fails the non-resolving set (naming the live replacement) and warns on the resolving set. The audit's job is to catch broken configs, not to enforce a style preference — failing a working alias would reject valid configs, which is the defect this change exists to fix. Rationale recorded in the script comments.Fixture results (real audit CLI)
delegation.modelcustom_providersfallback_providersauxiliaryblockqwen3.6-27B-code(raw but live) indelegation.modelgpu-densetests/test_audit_hermes_config_alias.pycovers these by executing the CLI (10 passed). Verified it is a genuine regression: against the pre-change audit the gpu-vision fixture exits 1.Sweep and corrections
gpu-self-healstops canonicalizinggpu-light; weights replaced by the CT 116 authority pointer.hermes-config-template,hermes-agent-baseline,hermes-key-enforcement,inference-optimization,gpu-fleet,gpu-monitor,infrastructure-controlnow use the livegpu-visionalias; thegpu-monitordiagram's swapped.8/.110labels are corrected.syslog-auto(Rule 7) so a copied profile passes the audit.gemma-4-12b) with a "retired" annotation rather than being re-attributed to an alias that did not produce the measurement.hermes-key-enforcementandlitellm-api-keys: CT 116 has nodefault_key_generate_paramsblock, so no default lifetime enforcement exists today. Whether agent keys should expire by default is an open policy question for the captain; no block was added..129config still namesgpu-light/gemma-4-E4B;.129is report-only, so it is recorded for its owner and not edited.Validation
no-mistakes run
01M2B67C734FF4DZ87V22VB4SD— outcome passed after several review rounds.