The litellm-health contract (the 4-hourly run contract: litellm-health check) carried drift that made its probes and expectations wrong, and it duplicated model/alias/timeout state that lives authoritatively in CT 116's /opt/inference-harness/litellm_config.yaml. After review found the duplication kept regenerating inconsistencies, the task owner directed a deliberate single-source-of-truth re-scope rather than more patches.
Changes
Public vs backend surfaces are now explicit. They serve the same app under different paths. Verified live: public https://litellm.sysloggh.net serves /ui/ and /docs (200) and 404s on the /litellm/ prefix; backend http://192.168.68.116:80 serves /litellm/ui/ and /litellm/docs (200) with /ui/ and /docs as 301 one-hop helpers. Every probe now names its surface.
Retired models removed and the rename completed.gemma-4-12b (returns 400 Invalid model name) and its alias gpu-light are replaced by the live gpu-vision on the RTX 5070 (.110); crew-auto is documented as retired, so the 64K crew cap it enforced no longer exists. Step 7 tests one model per GPU host (qwen3.6-27B-code, gpu-vision, strix-moe), all verified 200.
Duplicated config state deleted, not corrected. The fallback/timeout tables, the weighted-pool and stable-alias rpm tables, and a frozen /v1/models snapshot are gone from litellm-health.prose.md, litellm-self-heal.prose.md and gpu-fleet.prose.md. They are replaced by the minimum agents need (alias, what it serves, where, and whether it is a direct alias or a syslog-auto pool member) plus an explicit pointer to CT 116 /opt/inference-harness/litellm_config.yaml as the single source of truth for models, aliases, rpm caps, weights and fallbacks.
Inference no longer uses the master key. Step 7 authenticates with the dedicated monitor key from /etc/litellm-monitor.envon CT 116, keeping the master key for admin endpoints only. /v1/models is key-scoped (master, monitor and agent keys return different sets), so every quoted reading now names its key.
One owner per check.litellm-health is declared the active owner of the probes (status: active, stale replaced_by/"reference only" frontmatter removed); litellm-self-heal is reduced to a one-line pointer and keeps its remediation rules. The registry cadence for litellm-health is corrected from */10 * * * * to the real 4-hourly staggered dispatch, and the stale derived copies of that cadence are pointed at the registry instead of hand-maintained.
Validation
no-mistakes pipeline run 01M2B2RV27M4NHEWAKN31M4S09 — outcome passed (review, test, document, lint, push), after six fix rounds. Live probes re-verified throughout: surfaces as above, /litellm/health/liveliness 200, Grafana 200, 12/12 containers on CT 116, three model probes 200.
Known follow-ups (out of scope, explicitly tracked)
Retired-name sweep (queued separately, not started):gpu-light still appears in 7 files and gemma-4-12b in 12. Critically, the executable audit-hermes-config.py (Rule 8, lines ~93-100) REQUIRES auxiliary.vision.model == "gpu-light" and web_extract == "gpu-light", so a config adopting the new canonical gpu-vision currently fails our own audit — a functional break to fix first.
koby's config on .129 still names gpu-light (and gemma-4-E4B); .129 is report-only, so it is recorded for its owner rather than edited.
infrastructure-monitoring sources /etc/litellm-monitor.env on the ops host, where it does not exist (the file lives on CT 116) — the cause of its recurring Zulip 401.
## What and why
The `litellm-health` contract (the 4-hourly `run contract: litellm-health` check) carried drift that made its probes and expectations wrong, and it duplicated model/alias/timeout state that lives authoritatively in CT 116's `/opt/inference-harness/litellm_config.yaml`. After review found the duplication kept regenerating inconsistencies, the task owner directed a deliberate **single-source-of-truth re-scope** rather than more patches.
## Changes
- **Public vs backend surfaces are now explicit.** They serve the same app under different paths. Verified live: public `https://litellm.sysloggh.net` serves `/ui/` and `/docs` (200) and 404s on the `/litellm/` prefix; backend `http://192.168.68.116:80` serves `/litellm/ui/` and `/litellm/docs` (200) with `/ui/` and `/docs` as 301 one-hop helpers. Every probe now names its surface.
- **Retired models removed and the rename completed.** `gemma-4-12b` (returns 400 `Invalid model name`) and its alias `gpu-light` are replaced by the live `gpu-vision` on the RTX 5070 (.110); `crew-auto` is documented as retired, so the 64K crew cap it enforced no longer exists. Step 7 tests one model per GPU host (`qwen3.6-27B-code`, `gpu-vision`, `strix-moe`), all verified 200.
- **Duplicated config state deleted, not corrected.** The fallback/timeout tables, the weighted-pool and stable-alias rpm tables, and a frozen `/v1/models` snapshot are gone from `litellm-health.prose.md`, `litellm-self-heal.prose.md` and `gpu-fleet.prose.md`. They are replaced by the minimum agents need (alias, what it serves, where, and whether it is a direct alias or a `syslog-auto` pool member) plus an explicit pointer to CT 116 `/opt/inference-harness/litellm_config.yaml` as the single source of truth for models, aliases, rpm caps, weights and fallbacks.
- **Inference no longer uses the master key.** Step 7 authenticates with the dedicated monitor key from `/etc/litellm-monitor.env` **on CT 116**, keeping the master key for admin endpoints only. `/v1/models` is key-scoped (master, monitor and agent keys return different sets), so every quoted reading now names its key.
- **One owner per check.** `litellm-health` is declared the active owner of the probes (`status: active`, stale `replaced_by`/"reference only" frontmatter removed); `litellm-self-heal` is reduced to a one-line pointer and keeps its remediation rules. The registry cadence for `litellm-health` is corrected from `*/10 * * * *` to the real 4-hourly staggered dispatch, and the stale derived copies of that cadence are pointed at the registry instead of hand-maintained.
## Validation
no-mistakes pipeline run `01M2B2RV27M4NHEWAKN31M4S09` — outcome **passed** (review, test, document, lint, push), after six fix rounds. Live probes re-verified throughout: surfaces as above, `/litellm/health/liveliness` 200, Grafana 200, 12/12 containers on CT 116, three model probes 200.
## Known follow-ups (out of scope, explicitly tracked)
1. **Retired-name sweep (queued separately, not started):** `gpu-light` still appears in 7 files and `gemma-4-12b` in 12. Critically, the executable `audit-hermes-config.py` (Rule 8, lines ~93-100) REQUIRES `auxiliary.vision.model == "gpu-light"` and `web_extract == "gpu-light"`, so a config adopting the new canonical `gpu-vision` currently fails our own audit — a functional break to fix first.
2. **koby's config on .129** still names `gpu-light` (and `gemma-4-E4B`); .129 is report-only, so it is recorded for its owner rather than edited.
3. **`infrastructure-monitoring`** sources `/etc/litellm-monitor.env` on the ops host, where it does not exist (the file lives on CT 116) — the cause of its recurring Zulip 401.
- Execution step 2 now documents the public edge and the backend edge as two
distinct surfaces: public serves /ui/ and /docs (404 on the /litellm/ prefix),
backend http://192.168.68.116 serves /litellm/ui/ and /litellm/docs (with /ui/
and /docs as 301 helpers). Each probe names its surface.
- GPU topology: ocu-llm RTX 5070 now serves gpu-vision (gemma-4-12b retired).
- Fallback/timeout table rewritten to the live router_settings.fallbacks chains.
- Step 7 model list: gemma-4-12b -> gpu-vision, with a key-scoped /v1/models note
and the 2026-09-12 master-key registry snapshot.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What and why
The
litellm-healthcontract (the 4-hourlyrun contract: litellm-healthcheck) carried drift that made its probes and expectations wrong, and it duplicated model/alias/timeout state that lives authoritatively in CT 116's/opt/inference-harness/litellm_config.yaml. After review found the duplication kept regenerating inconsistencies, the task owner directed a deliberate single-source-of-truth re-scope rather than more patches.Changes
https://litellm.sysloggh.netserves/ui/and/docs(200) and 404s on the/litellm/prefix; backendhttp://192.168.68.116:80serves/litellm/ui/and/litellm/docs(200) with/ui/and/docsas 301 one-hop helpers. Every probe now names its surface.gemma-4-12b(returns 400Invalid model name) and its aliasgpu-lightare replaced by the livegpu-visionon the RTX 5070 (.110);crew-autois documented as retired, so the 64K crew cap it enforced no longer exists. Step 7 tests one model per GPU host (qwen3.6-27B-code,gpu-vision,strix-moe), all verified 200./v1/modelssnapshot are gone fromlitellm-health.prose.md,litellm-self-heal.prose.mdandgpu-fleet.prose.md. They are replaced by the minimum agents need (alias, what it serves, where, and whether it is a direct alias or asyslog-autopool member) plus an explicit pointer to CT 116/opt/inference-harness/litellm_config.yamlas the single source of truth for models, aliases, rpm caps, weights and fallbacks./etc/litellm-monitor.envon CT 116, keeping the master key for admin endpoints only./v1/modelsis key-scoped (master, monitor and agent keys return different sets), so every quoted reading now names its key.litellm-healthis declared the active owner of the probes (status: active, stalereplaced_by/"reference only" frontmatter removed);litellm-self-healis reduced to a one-line pointer and keeps its remediation rules. The registry cadence forlitellm-healthis corrected from*/10 * * * *to the real 4-hourly staggered dispatch, and the stale derived copies of that cadence are pointed at the registry instead of hand-maintained.Validation
no-mistakes pipeline run
01M2B2RV27M4NHEWAKN31M4S09— outcome passed (review, test, document, lint, push), after six fix rounds. Live probes re-verified throughout: surfaces as above,/litellm/health/liveliness200, Grafana 200, 12/12 containers on CT 116, three model probes 200.Known follow-ups (out of scope, explicitly tracked)
gpu-lightstill appears in 7 files andgemma-4-12bin 12. Critically, the executableaudit-hermes-config.py(Rule 8, lines ~93-100) REQUIRESauxiliary.vision.model == "gpu-light"andweb_extract == "gpu-light", so a config adopting the new canonicalgpu-visioncurrently fails our own audit — a functional break to fix first.gpu-light(andgemma-4-E4B); .129 is report-only, so it is recorded for its owner rather than edited.infrastructure-monitoringsources/etc/litellm-monitor.envon the ops host, where it does not exist (the file lives on CT 116) — the cause of its recurring Zulip 401.