docs(litellm-health): single source of truth, gpu-vision alias, litellm-health as live owner #79

Merged
abiba-bot merged 8 commits from fix/litellm-health-drift-20260912 into master 2026-09-12 16:05:28 +00:00
Owner

What and why

The litellm-health contract (the 4-hourly run contract: litellm-health check) carried drift that made its probes and expectations wrong, and it duplicated model/alias/timeout state that lives authoritatively in CT 116's /opt/inference-harness/litellm_config.yaml. After review found the duplication kept regenerating inconsistencies, the task owner directed a deliberate single-source-of-truth re-scope rather than more patches.

Changes

  • Public vs backend surfaces are now explicit. They serve the same app under different paths. Verified live: public https://litellm.sysloggh.net serves /ui/ and /docs (200) and 404s on the /litellm/ prefix; backend http://192.168.68.116:80 serves /litellm/ui/ and /litellm/docs (200) with /ui/ and /docs as 301 one-hop helpers. Every probe now names its surface.
  • Retired models removed and the rename completed. gemma-4-12b (returns 400 Invalid model name) and its alias gpu-light are replaced by the live gpu-vision on the RTX 5070 (.110); crew-auto is documented as retired, so the 64K crew cap it enforced no longer exists. Step 7 tests one model per GPU host (qwen3.6-27B-code, gpu-vision, strix-moe), all verified 200.
  • Duplicated config state deleted, not corrected. The fallback/timeout tables, the weighted-pool and stable-alias rpm tables, and a frozen /v1/models snapshot are gone from litellm-health.prose.md, litellm-self-heal.prose.md and gpu-fleet.prose.md. They are replaced by the minimum agents need (alias, what it serves, where, and whether it is a direct alias or a syslog-auto pool member) plus an explicit pointer to CT 116 /opt/inference-harness/litellm_config.yaml as the single source of truth for models, aliases, rpm caps, weights and fallbacks.
  • Inference no longer uses the master key. Step 7 authenticates with the dedicated monitor key from /etc/litellm-monitor.env on CT 116, keeping the master key for admin endpoints only. /v1/models is key-scoped (master, monitor and agent keys return different sets), so every quoted reading now names its key.
  • One owner per check. litellm-health is declared the active owner of the probes (status: active, stale replaced_by/"reference only" frontmatter removed); litellm-self-heal is reduced to a one-line pointer and keeps its remediation rules. The registry cadence for litellm-health is corrected from */10 * * * * to the real 4-hourly staggered dispatch, and the stale derived copies of that cadence are pointed at the registry instead of hand-maintained.

Validation

no-mistakes pipeline run 01M2B2RV27M4NHEWAKN31M4S09 — outcome passed (review, test, document, lint, push), after six fix rounds. Live probes re-verified throughout: surfaces as above, /litellm/health/liveliness 200, Grafana 200, 12/12 containers on CT 116, three model probes 200.

Known follow-ups (out of scope, explicitly tracked)

  1. Retired-name sweep (queued separately, not started): gpu-light still appears in 7 files and gemma-4-12b in 12. Critically, the executable audit-hermes-config.py (Rule 8, lines ~93-100) REQUIRES auxiliary.vision.model == "gpu-light" and web_extract == "gpu-light", so a config adopting the new canonical gpu-vision currently fails our own audit — a functional break to fix first.
  2. koby's config on .129 still names gpu-light (and gemma-4-E4B); .129 is report-only, so it is recorded for its owner rather than edited.
  3. infrastructure-monitoring sources /etc/litellm-monitor.env on the ops host, where it does not exist (the file lives on CT 116) — the cause of its recurring Zulip 401.
## What and why The `litellm-health` contract (the 4-hourly `run contract: litellm-health` check) carried drift that made its probes and expectations wrong, and it duplicated model/alias/timeout state that lives authoritatively in CT 116's `/opt/inference-harness/litellm_config.yaml`. After review found the duplication kept regenerating inconsistencies, the task owner directed a deliberate **single-source-of-truth re-scope** rather than more patches. ## Changes - **Public vs backend surfaces are now explicit.** They serve the same app under different paths. Verified live: public `https://litellm.sysloggh.net` serves `/ui/` and `/docs` (200) and 404s on the `/litellm/` prefix; backend `http://192.168.68.116:80` serves `/litellm/ui/` and `/litellm/docs` (200) with `/ui/` and `/docs` as 301 one-hop helpers. Every probe now names its surface. - **Retired models removed and the rename completed.** `gemma-4-12b` (returns 400 `Invalid model name`) and its alias `gpu-light` are replaced by the live `gpu-vision` on the RTX 5070 (.110); `crew-auto` is documented as retired, so the 64K crew cap it enforced no longer exists. Step 7 tests one model per GPU host (`qwen3.6-27B-code`, `gpu-vision`, `strix-moe`), all verified 200. - **Duplicated config state deleted, not corrected.** The fallback/timeout tables, the weighted-pool and stable-alias rpm tables, and a frozen `/v1/models` snapshot are gone from `litellm-health.prose.md`, `litellm-self-heal.prose.md` and `gpu-fleet.prose.md`. They are replaced by the minimum agents need (alias, what it serves, where, and whether it is a direct alias or a `syslog-auto` pool member) plus an explicit pointer to CT 116 `/opt/inference-harness/litellm_config.yaml` as the single source of truth for models, aliases, rpm caps, weights and fallbacks. - **Inference no longer uses the master key.** Step 7 authenticates with the dedicated monitor key from `/etc/litellm-monitor.env` **on CT 116**, keeping the master key for admin endpoints only. `/v1/models` is key-scoped (master, monitor and agent keys return different sets), so every quoted reading now names its key. - **One owner per check.** `litellm-health` is declared the active owner of the probes (`status: active`, stale `replaced_by`/"reference only" frontmatter removed); `litellm-self-heal` is reduced to a one-line pointer and keeps its remediation rules. The registry cadence for `litellm-health` is corrected from `*/10 * * * *` to the real 4-hourly staggered dispatch, and the stale derived copies of that cadence are pointed at the registry instead of hand-maintained. ## Validation no-mistakes pipeline run `01M2B2RV27M4NHEWAKN31M4S09` — outcome **passed** (review, test, document, lint, push), after six fix rounds. Live probes re-verified throughout: surfaces as above, `/litellm/health/liveliness` 200, Grafana 200, 12/12 containers on CT 116, three model probes 200. ## Known follow-ups (out of scope, explicitly tracked) 1. **Retired-name sweep (queued separately, not started):** `gpu-light` still appears in 7 files and `gemma-4-12b` in 12. Critically, the executable `audit-hermes-config.py` (Rule 8, lines ~93-100) REQUIRES `auxiliary.vision.model == "gpu-light"` and `web_extract == "gpu-light"`, so a config adopting the new canonical `gpu-vision` currently fails our own audit — a functional break to fix first. 2. **koby's config on .129** still names `gpu-light` (and `gemma-4-E4B`); .129 is report-only, so it is recorded for its owner rather than edited. 3. **`infrastructure-monitoring`** sources `/etc/litellm-monitor.env` on the ops host, where it does not exist (the file lives on CT 116) — the cause of its recurring Zulip 401.
abiba-bot added 8 commits 2026-09-12 16:02:41 +00:00
- Execution step 2 now documents the public edge and the backend edge as two
  distinct surfaces: public serves /ui/ and /docs (404 on the /litellm/ prefix),
  backend http://192.168.68.116 serves /litellm/ui/ and /litellm/docs (with /ui/
  and /docs as 301 helpers). Each probe names its surface.
- GPU topology: ocu-llm RTX 5070 now serves gpu-vision (gemma-4-12b retired).
- Fallback/timeout table rewritten to the live router_settings.fallbacks chains.
- Step 7 model list: gemma-4-12b -> gpu-vision, with a key-scoped /v1/models note
  and the 2026-09-12 master-key registry snapshot.
no-mistakes(document): Dedupe cadence copies; registry remains authoritative
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
21f7b6171c
abiba-bot merged commit 1dc040251d into master 2026-09-12 16:05:28 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#79