fix(audit): stop requiring the retired gpu-light alias; derive model fields; sweep retired names #80

Merged
abiba-bot merged 8 commits from fix/retired-alias-sweep-20260912 into master 2026-09-12 17:15:06 +00:00
Owner

The functional break (fixed first)

audit-hermes-config.py Rule 8 required auxiliary.vision.model == "gpu-light" and auxiliary.web_extract.model == "gpu-light". gpu-light (and its raw predecessor gemma-4-12b) were retired on 2026-09-12 and return 400 Invalid model name; the live RTX 5070 canonical alias is gpu-vision. So a config that adopted the correct alias failed our own audit — the audit was actively pushing configs toward a dead alias.

Proven by executing the real CLI (before → after):

BEFORE:  gpu-vision config -> FAIL (Rule 8)      gpu-light config -> PASS
AFTER:   gpu-vision config -> PASS (exit 0)      gpu-light config -> FAIL (exit 1)

Model-bearing fields are now derived, not enumerated

Five review rounds each found one more field the hand-written list had missed (custom_providers, then fallback_providers, then x_search). Rather than patch a sixth, the audit now derives the values to check: a recursive walk of the loaded config for every mapping key named model / model_name at any depth, including inside lists, plus a small documented allowlist for the exceptions where a model name sits under a different key (models: lists). The derivation rule is documented in the script so the next person extends the allowlist instead of adding a field by hand.

Fail vs warn — the owner's policy call

Verified through live LiteLLM: qwen3.6-27B-code and qwen3.6-35B-udq4 resolve (200) — live, discouraged but working; gpu-light, gemma-4-12b, crew-auto, ornith-1.0-35b do not resolve. So the audit hard-fails the non-resolving set (naming the live replacement) and warns on the resolving set. The audit's job is to catch broken configs, not to enforce a style preference — failing a working alias would reject valid configs, which is the defect this change exists to fix. Rationale recorded in the script comments.

Fixture results (real audit CLI)

Fixture Result
gpu-vision config exit 0 — PASS
gpu-light in delegation.model exit 1 — FAIL
gpu-light under custom_providers exit 1 — FAIL
gpu-light under fallback_providers exit 1 — FAIL
gpu-light in a nested auxiliary block exit 1 — FAIL
qwen3.6-27B-code (raw but live) in delegation.model exit 0 — PASS with a warning naming gpu-dense

tests/test_audit_hermes_config_alias.py covers these by executing the CLI (10 passed). Verified it is a genuine regression: against the pre-change audit the gpu-vision fixture exits 1.

Sweep and corrections

  • gpu-self-heal stops canonicalizing gpu-light; weights replaced by the CT 116 authority pointer.
  • hermes-config-template, hermes-agent-baseline, hermes-key-enforcement, inference-optimization, gpu-fleet, gpu-monitor, infrastructure-control now use the live gpu-vision alias; the gpu-monitor diagram's swapped .8/.110 labels are corrected.
  • The Mumuni reference profile's compression entries are aligned to syslog-auto (Rule 7) so a copied profile passes the audit.
  • The dated latency table keeps the measured name (gemma-4-12b) with a "retired" annotation rather than being re-attributed to an alias that did not produce the measurement.
  • The key-duration claim is corrected in both hermes-key-enforcement and litellm-api-keys: CT 116 has no default_key_generate_params block, so no default lifetime enforcement exists today. Whether agent keys should expire by default is an open policy question for the captain; no block was added.
  • koby's .129 config still names gpu-light/gemma-4-E4B; .129 is report-only, so it is recorded for its owner and not edited.

Validation

no-mistakes run 01M2B67C734FF4DZ87V22VB4SD — outcome passed after several review rounds.

## The functional break (fixed first) `audit-hermes-config.py` Rule 8 required `auxiliary.vision.model == "gpu-light"` and `auxiliary.web_extract.model == "gpu-light"`. `gpu-light` (and its raw predecessor `gemma-4-12b`) were retired on 2026-09-12 and return 400 `Invalid model name`; the live RTX 5070 canonical alias is `gpu-vision`. So a config that adopted the *correct* alias **failed our own audit** — the audit was actively pushing configs toward a dead alias. Proven by executing the real CLI (before → after): ``` BEFORE: gpu-vision config -> FAIL (Rule 8) gpu-light config -> PASS AFTER: gpu-vision config -> PASS (exit 0) gpu-light config -> FAIL (exit 1) ``` ## Model-bearing fields are now derived, not enumerated Five review rounds each found one more field the hand-written list had missed (`custom_providers`, then `fallback_providers`, then `x_search`). Rather than patch a sixth, the audit now **derives** the values to check: a recursive walk of the loaded config for every mapping key named `model` / `model_name` at any depth, including inside lists, plus a small documented allowlist for the exceptions where a model name sits under a different key (`models:` lists). The derivation rule is documented in the script so the next person extends the allowlist instead of adding a field by hand. ## Fail vs warn — the owner's policy call Verified through live LiteLLM: `qwen3.6-27B-code` and `qwen3.6-35B-udq4` resolve (200) — live, discouraged but working; `gpu-light`, `gemma-4-12b`, `crew-auto`, `ornith-1.0-35b` do not resolve. So the audit **hard-fails the non-resolving set** (naming the live replacement) and **warns** on the resolving set. The audit's job is to catch broken configs, not to enforce a style preference — failing a working alias would reject valid configs, which is the defect this change exists to fix. Rationale recorded in the script comments. ## Fixture results (real audit CLI) | Fixture | Result | |---|---| | gpu-vision config | **exit 0 — PASS** | | gpu-light in `delegation.model` | **exit 1 — FAIL** | | gpu-light under `custom_providers` | **exit 1 — FAIL** | | gpu-light under `fallback_providers` | **exit 1 — FAIL** | | gpu-light in a nested `auxiliary` block | **exit 1 — FAIL** | | `qwen3.6-27B-code` (raw but live) in `delegation.model` | **exit 0 — PASS with a warning** naming `gpu-dense` | `tests/test_audit_hermes_config_alias.py` covers these by executing the CLI (10 passed). Verified it is a genuine regression: against the pre-change audit the gpu-vision fixture exits 1. ## Sweep and corrections - `gpu-self-heal` stops canonicalizing `gpu-light`; weights replaced by the CT 116 authority pointer. - `hermes-config-template`, `hermes-agent-baseline`, `hermes-key-enforcement`, `inference-optimization`, `gpu-fleet`, `gpu-monitor`, `infrastructure-control` now use the live `gpu-vision` alias; the `gpu-monitor` diagram's swapped `.8`/`.110` labels are corrected. - The Mumuni reference profile's compression entries are aligned to `syslog-auto` (Rule 7) so a copied profile passes the audit. - The dated latency table keeps the **measured** name (`gemma-4-12b`) with a "retired" annotation rather than being re-attributed to an alias that did not produce the measurement. - The key-duration claim is corrected in both `hermes-key-enforcement` and `litellm-api-keys`: CT 116 has **no** `default_key_generate_params` block, so no default lifetime enforcement exists today. Whether agent keys should expire by default is an **open policy question** for the captain; no block was added. - koby's `.129` config still names `gpu-light`/`gemma-4-E4B`; `.129` is report-only, so it is recorded for its owner and not edited. ## Validation no-mistakes run `01M2B67C734FF4DZ87V22VB4SD` — outcome **passed** after several review rounds.
abiba-bot added 7 commits 2026-09-12 16:55:22 +00:00
audit-hermes-config.py Rule 8 required auxiliary.vision.model and
auxiliary.web_extract.model to equal the retired 'gpu-light', so a config
adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was
enforcing a dead alias (400 Invalid model name). Rule 8 now requires
gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set;
the guidance message names the live aliases.

Sweep of the remaining references: gpu-self-heal stops canonicalizing
gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement,
inference-optimization, litellm-client-timeouts and gpu-fleet now use the live
gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now
points at CT 116 /opt/inference-harness/litellm_config.yaml instead of
duplicating it. koby's .129 config is report-only and recorded, not edited.

Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts
gpu-vision passes while gpu-light and gemma-4-12b fail.
no-mistakes(document): Sweep residual gemma labels; align compression rule contradiction
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
78b501798f
abiba-bot added 1 commit 2026-09-12 17:13:58 +00:00
ci(pr-pipeline): fix flaky frontmatter check (grep -q SIGPIPE under pipefail)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
a2edc2f56f
The validate job runs with bash -e -o pipefail. `echo "$FM" | grep -q '^name:'`
lets grep exit on first match, which can SIGPIPE the echo; pipefail then reports
the pipeline non-zero and the || branch raises a false "Missing name/description".
The flagged file set varied run to run (and included files untouched by the PR)
while a fresh clone of the same commit passes the identical check. Reproduced:
the old form failed 3 of 5 local runs under the same shell flags, the herestring
form passed 5 of 5. Use herestrings so no pipe can be broken.
abiba-bot merged commit d9e06863d8 into master 2026-09-12 17:15:06 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#80