Aligns syslog-harness with the verified live state of CT 116 and the GPU hosts, and removes model-version names from every client-facing surface.
Why
Kwame's directive: "do not advertise models to downstream that way — we won't have to update those when we update models." Client-visible names should describe capability, not model version. While aligning, the repo also turned out to be stale in four places against live state.
What changed (all verified against running systems, not from documentation)
Area
Change
litellm_config.yaml
client catalogue is now exactly syslog-auto, gpu-dense, strix-moe, gpu-vision; fallbacks + model_cost re-keyed; upstream litellm_params.model set to capability names (backends ignore the model field — verified 200 on all three hosts, so no llama.cpp relaunch, no KV warmth lost)
gpu_roster.yaml
launch args replaced with the verified live command lines read from the running processes (PVE guest agent on acerpve VM 101 = .8, ocupve VM 103 = .110; SSH on amdpve = .15). .8 model_path corrected, max_concurrent 2 → 1, context 262144 → 131072; .15 model_path corrected; keys renamed to capability names
README.md / dashboard
dense entry corrected (1 slot, 131K ctx, actual model file); picker uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070 serves a 9B Qwen3.5) and the gpu-vision → gpu-light id mapping
scripts/
added the three operational monitors as tracked files (they were untracked) with their retired-name references cleared — they were failing every benchmark cycle (150 failed gemma-4-12b calls in 7 days), and gpu-monitor.py's .110 entry wrongly listed .8's model
Verification
/v1/models advertises exactly the four names; each returns 200 through nginx; container healthy.
Health check now reports model-gpu-dense, model-strix-moe, model-gpu-vision, model-syslog-autoall pass (previously two dead names).
11 virtual-key allowlists in the LiteLLM DB were updated in the same change (not part of this repo).
Rollback for the config: litellm_config.yaml.bak-20260912-prerename on CT 116 + docker restart harness-litellm.
Intentionally NOT changed
LITELLM-MIGRATION-PLAN.md — historical planning document (June 14, 2026), not a live-state claim.
Unrelated untracked files (router.py, docker-compose.yml.pre-1991-20260911, nginx/default.conf, dashboard/gpu-monitor.html, scripts/gitea-logger.sh) — out of scope.
Note for review
.8's live server alias is still qwen3.6-27B-code while the loaded weights are Qwen3.8-27B-Uncensored-Q4_K_M.gguf. Renaming the alias requires a llama.cpp relaunch (and loses KV warmth), so it is deliberately left as an explicit follow-up; the roster records the actual value.
Aligns `syslog-harness` with the **verified live state** of CT 116 and the GPU hosts, and removes model-version names from every client-facing surface.
## Why
Kwame's directive: *"do not advertise models to downstream that way — we won't have to update those when we update models."* Client-visible names should describe **capability**, not model version. While aligning, the repo also turned out to be stale in four places against live state.
## What changed (all verified against running systems, not from documentation)
| Area | Change |
|---|---|
| `litellm_config.yaml` | client catalogue is now exactly `syslog-auto`, `gpu-dense`, `strix-moe`, `gpu-vision`; `fallbacks` + `model_cost` re-keyed; upstream `litellm_params.model` set to capability names (backends ignore the model field — verified 200 on all three hosts, so **no llama.cpp relaunch, no KV warmth lost**) |
| `gpu_roster.yaml` | launch args replaced with the **verified live command lines** read from the running processes (PVE guest agent on acerpve VM 101 = .8, ocupve VM 103 = .110; SSH on amdpve = .15). `.8` model_path corrected, `max_concurrent` 2 → 1, `context` 262144 → 131072; `.15` model_path corrected; keys renamed to capability names |
| `README.md` / `dashboard` | dense entry corrected (1 slot, 131K ctx, actual model file); picker uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070 serves a **9B Qwen3.5**) and the `gpu-vision → gpu-light` id mapping |
| `scripts/` | added the three operational monitors as **tracked files** (they were untracked) with their retired-name references cleared — they were failing every benchmark cycle (150 failed `gemma-4-12b` calls in 7 days), and `gpu-monitor.py`'s `.110` entry wrongly listed `.8`'s model |
## Verification
- `/v1/models` advertises exactly the four names; each returns **200** through nginx; container healthy.
- Health check now reports `model-gpu-dense`, `model-strix-moe`, `model-gpu-vision`, `model-syslog-auto` **all pass** (previously two dead names).
- 11 virtual-key allowlists in the LiteLLM DB were updated in the same change (not part of this repo).
- Rollback for the config: `litellm_config.yaml.bak-20260912-prerename` on CT 116 + `docker restart harness-litellm`.
## Intentionally NOT changed
- `LITELLM-MIGRATION-PLAN.md` — historical planning document (June 14, 2026), not a live-state claim.
- `backups/`, `graphify-out/`, `litellm_config.yaml.backup` — historical artifacts.
- Unrelated untracked files (`router.py`, `docker-compose.yml.pre-1991-20260911`, `nginx/default.conf`, `dashboard/gpu-monitor.html`, `scripts/gitea-logger.sh`) — out of scope.
## Note for review
`.8`'s live server alias is still `qwen3.6-27B-code` while the loaded weights are `Qwen3.8-27B-Uncensored-Q4_K_M.gguf`. Renaming the alias requires a llama.cpp relaunch (and loses KV warmth), so it is deliberately left as an explicit follow-up; the roster records the actual value.
litellm_config.yaml
- client-visible model_list reduced to capability names: syslog-auto, gpu-dense, strix-moe, gpu-vision
- retired qwen3.6-27B-code, qwen3.8-27B-uncensored, qwen3.6-35B-udq4 (and the already-dead gemma-4-12b)
- fallbacks and model_cost re-keyed to the surviving names
- upstream litellm_params.model set to the capability names; the backends ignore the model field
(verified HTTP 200 on all three hosts), so this needs no llama.cpp relaunch and loses no KV warmth
- verified live: /v1/models advertises exactly the four names, each returns 200 through nginx
gpu_roster.yaml
- launch args replaced with the VERIFIED live command lines, read from the running processes via the
PVE guest agent (acerpve VM 101 = .8, ocupve VM 103 = .110) and on amdpve (.15)
- .8: model_path corrected to Qwen3.8-27B-Uncensored-Q4_K_M.gguf, max_concurrent 2 -> 1,
context 262144 -> 131072, full arg list recorded (incl. --parallel 1 and the new
--slot-save-path / --metrics added 2026-09-12)
- .15: model_path corrected to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf, full arg list recorded
- keys renamed to the capability names; hosts.current_model aligned
README.md / dashboard
- README dense entry corrected (1 slot, 131K ctx, actual model file)
- dashboard picker now uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070
serves a 9B Qwen3.5) and the gpu-vision -> gpu-light id mapping
scripts/
- added the three operational monitors as tracked files (they were untracked): gpu-monitor.py,
gpu-self-heal.py, litellm-health-check.sh
- cleared their references to retired model names, which were causing failed calls every benchmark
cycle (150 failed gemma-4-12b calls in the last 7 days); gpu-monitor.py's .110 entry also wrongly
listed .8's model
Intentionally NOT changed
- LITELLM-MIGRATION-PLAN.md: historical planning document (June 14), not a live-state claim
- backups/, graphify-out/, litellm_config.yaml.backup: historical artifacts
- unrelated untracked files (router.py, docker-compose.yml.pre-1991-20260911, nginx/default.conf,
dashboard/gpu-monitor.html, scripts/gitea-logger.sh): out of scope for this change
Review verdict: REQUEST CHANGES — from the 2026-09-19 management-brief review (deep diff review via delegated reviewer; security posture verified live by Mumuni).
The alignment work itself checks out: no retired model names survive except the intentional allowlist migration notes, litellm_config.yaml is internally consistent with the capability-name catalogue (syslog-auto / gpu-dense / strix-moe / gpu-vision), Python scripts py_compile clean, shell scripts bash -n clean, live gateway confirms the four advertised model names match.
Blockers:
[URGENT — exposed NOW, merge or not]scripts/gpu-monitor.py (~L303) hardcodes a live Zulip bot API key (cKTD…), and scripts/gpu-self-heal.py (~L387) hardcodes the LiteLLM master key (sk-lit…6796). Verified: this repo is public, Gitea allows anonymous raw reads, and git.sysloggh.net serves this PR branch from the public internet as we speak. Consequences: (a) rotate both credentials immediately — a branch delete after the fact does not un-leak them; (b) before merge, move both to the existing /etc/litellm-monitor.env pattern (no literal values in tracked files).
[minor]scripts/gpu-monitor.py L24-26: DASHBOARD_PATH hardcoded to /root/… — 404s silently if the monitor ever runs as a non-root service user. Make it configurable via env.
[minor]scripts/litellm-health-check.sh L12: MONITOR_KEY silently falls back to the master key when the env file is missing — that hides misconfiguration and encourages master-key use; fail loudly instead.
Repo hygiene note:main and master have no merge base (orphaned histories) — this PR's base ambiguity is a symptom. Recommend picking one branch and deleting the other.
Everything else is good to merge once the secrets are externalized and the two keys are rotated. Owner pinged via RA-H relay. — Mumuni 🦅
**Review verdict: REQUEST CHANGES** — from the 2026-09-19 management-brief review (deep diff review via delegated reviewer; security posture verified live by Mumuni).
The alignment work itself checks out: no retired model names survive except the intentional allowlist migration notes, litellm_config.yaml is internally consistent with the capability-name catalogue (syslog-auto / gpu-dense / strix-moe / gpu-vision), Python scripts py_compile clean, shell scripts bash -n clean, live gateway confirms the four advertised model names match.
**Blockers:**
1. **[URGENT — exposed NOW, merge or not]** `scripts/gpu-monitor.py` (~L303) hardcodes a live Zulip bot API key (`cKTD…`), and `scripts/gpu-self-heal.py` (~L387) hardcodes the LiteLLM **master key** (`sk-lit…6796`). Verified: this repo is **public**, Gitea allows anonymous raw reads, and `git.sysloggh.net` serves this PR branch from the public internet *as we speak*. Consequences: (a) **rotate both credentials immediately** — a branch delete after the fact does not un-leak them; (b) before merge, move both to the existing `/etc/litellm-monitor.env` pattern (no literal values in tracked files).
2. **[minor]** `scripts/gpu-monitor.py` L24-26: `DASHBOARD_PATH` hardcoded to `/root/…` — 404s silently if the monitor ever runs as a non-root service user. Make it configurable via env.
3. **[minor]** `scripts/litellm-health-check.sh` L12: `MONITOR_KEY` silently falls back to the master key when the env file is missing — that hides misconfiguration and encourages master-key use; fail loudly instead.
**Repo hygiene note:** `main` and `master` have no merge base (orphaned histories) — this PR's base ambiguity is a symptom. Recommend picking one branch and deleting the other.
Everything else is good to merge once the secrets are externalized and the two keys are rotated. Owner pinged via RA-H relay. — Mumuni 🦅
Follow-up: reviewer report landed + one correction to my review above (tested 23:32 UTC).
Correction on blocker #1 urgency: I wrote "rotate both credentials immediately." Re-tested by direct API call: both embedded keys are already dead — the sk-lit… master key gets 401 token_not_found from the live LiteLLM (the independent reviewer reproduced this same 401 independently), and the Zulip cKTD… key gets 401 Invalid API key. They're stale remnants of the 2026-09-12 alignment work; rotation evidently already happened. No live exposure, no urgency. What stands: this repo is public and the branch is anonymously fetchable from the internet (git.sysloggh.net/.../raw/branch/fix/harness-align-20260912/... → HTTP 200), so dead-but-confusable literals in public history still must be stripped in favor of the /etc/litellm-monitor.env pattern before merge. Verdict stays REQUEST CHANGES but low priority — the reviewer's independent verdict was APPROVE_WITH_NITS ("core goal achieved cleanly… no blocker-level issues"); the secrets are the only reason I hold it.
Additional findings from the deep review (fold into the same fix pass):
README.md L9 + hardware table L64: retired name qwen3.6-35B-A3B (MoE) survives on the Strix Halo row — PR fixed the 3090 row but missed this one; also "Total: 6 concurrent slots" is stale vs the new roster (dense now 1 slot ⇒ total 5).
gpu-self-heal.py:144 — host_key fallback (name.replace(" ","-").lower()) produces values like nvidia-geforce-rtx-3090 that can never match a GPU_HOSTS key, so unknown GPUs silently get host_key='unknown' and the Rule 7 Prometheus check is skipped — a silent monitoring gap.
litellm-health-check.sh:162-170 — "401 errors" counter actually matches any Received= auth error, inflating the count; label or regex should tighten.
litellm-health-check.sh:343 — stale FEED TO KNOWLEDGE GRAPH heading left above the corrected LOG TO GITEA comment.
litellm-health-check.sh:12 — silent fallback of MONITOR_KEY to the master key contradicts its own "NEVER for inference" comment; log a warning or fail.
gpu-self-heal.py — bare except: at ~704/714/722/1065 swallow KeyboardInterrupt; tighten to except Exception.
gpu-monitor.py:24 — DASHBOARD_PATH=/root/... (already in my first comment) + HISTORY_FILE=/var/log/litellm/ requires write perms — fine under root, note it in README if this ever ships as a service user.
None change the substance: the alignment work is solid. Fix list = env-ize the two dead keys + README row/slot arithmetic + item 2 (the one with real operational teeth). — Mumuni 🦅
**Follow-up: reviewer report landed + one correction to my review above (tested 23:32 UTC).**
**Correction on blocker #1 urgency:** I wrote "rotate both credentials immediately." Re-tested by direct API call: **both embedded keys are already dead** — the `sk-lit…` master key gets `401 token_not_found` from the live LiteLLM (the independent reviewer reproduced this same 401 independently), and the Zulip `cKTD…` key gets `401 Invalid API key`. They're stale remnants of the 2026-09-12 alignment work; rotation evidently already happened. **No live exposure, no urgency.** What stands: this repo is public and the branch is anonymously fetchable from the internet (`git.sysloggh.net/.../raw/branch/fix/harness-align-20260912/...` → HTTP 200), so dead-but-confusable literals in public history still must be stripped in favor of the `/etc/litellm-monitor.env` pattern before merge. Verdict stays REQUEST CHANGES but **low priority** — the reviewer's independent verdict was APPROVE_WITH_NITS ("core goal achieved cleanly… no blocker-level issues"); the secrets are the only reason I hold it.
**Additional findings from the deep review (fold into the same fix pass):**
1. README.md L9 + hardware table L64: retired name `qwen3.6-35B-A3B (MoE)` survives on the Strix Halo row — PR fixed the 3090 row but missed this one; also "Total: 6 concurrent slots" is stale vs the new roster (dense now 1 slot ⇒ total 5).
2. gpu-self-heal.py:144 — `host_key` fallback (`name.replace(" ","-").lower()`) produces values like `nvidia-geforce-rtx-3090` that can never match a GPU_HOSTS key, so unknown GPUs silently get `host_key='unknown'` and the Rule 7 Prometheus check is skipped — a silent monitoring gap.
3. litellm-health-check.sh:162-170 — "401 errors" counter actually matches any `Received=` auth error, inflating the count; label or regex should tighten.
4. litellm-health-check.sh:343 — stale `FEED TO KNOWLEDGE GRAPH` heading left above the corrected `LOG TO GITEA` comment.
5. litellm-health-check.sh:12 — silent fallback of MONITOR_KEY to the master key contradicts its own "NEVER for inference" comment; log a warning or fail.
6. gpu-self-heal.py — bare `except:` at ~704/714/722/1065 swallow KeyboardInterrupt; tighten to `except Exception`.
7. gpu-monitor.py:316-322 — SSH fallback checks `len(parts) >= 8` but indexes `parts[8]`; needs `>= 9` guard.
8. gpu-monitor.py:24 — `DASHBOARD_PATH=/root/...` (already in my first comment) + `HISTORY_FILE=/var/log/litellm/` requires write perms — fine under root, note it in README if this ever ships as a service user.
None change the substance: the alignment work is solid. Fix list = env-ize the two dead keys + README row/slot arithmetic + item 2 (the one with real operational teeth). — Mumuni 🦅
Merged 2026-09-20 under Kwame's decision: "review and merge if aligned with live state."
Pre-merge pass executed on this branch (5 commits, all validated: ast.parse + bash -n + secret-literal scan = 0 leaks):
scripts/gpu-monitor.py — Zulip bot key literal -> os.environ['ZULIP_API_KEY'] with graceful suppression when unset (key was dead anyway; re-issue for abiba-bot is on her).
scripts/gpu-self-heal.py — broken sk-lit...6796 literal -> env var (the literal was a truncated non-working string, so this also fixes the probe rather than just scrubbing).
scripts/litellm-health-check.sh — both Bearer placeholders rewired through LLK="${LITELLM_API_KEY:?}" guard at entry.
Rule-7 blind spot: unmapped GPU names now emit an explicit [warn] ... remote rules (incl. Rule 7) skipped instead of silently inheriting the default threshold.
Reviewer verdict trail: #689 (REQUEST CHANGES, hardcoded keys) -> #690 (correction: keys proven dead, downgrade to strip-before-merge) -> this merge. Nits NOT addressed here (tracked for a follow-up): ERR_401 overcount in health-check.sh, bare except: blocks, HISTORY_FILE perms coupling, hardcoded container threshold 8.
**Merged 2026-09-20 under Kwame's decision: "review and merge if aligned with live state."**
Pre-merge pass executed on this branch (5 commits, all validated: `ast.parse` + `bash -n` + secret-literal scan = 0 leaks):
1. `scripts/gpu-monitor.py` — Zulip bot key literal -> `os.environ['ZULIP_API_KEY']` with graceful suppression when unset (key was dead anyway; re-issue for abiba-bot is on her).
2. `scripts/gpu-self-heal.py` — broken `sk-lit...6796` literal -> env var (the literal was a truncated non-working string, so this also *fixes* the probe rather than just scrubbing).
3. `scripts/litellm-health-check.sh` — both Bearer placeholders rewired through `LLK="${LITELLM_API_KEY:?}"` guard at entry.
4. `README.md` — architecture block + fleet table aligned to deployed roster args (strix-moe naming, vision 1 slot / 131K, total 4 concurrent slots).
5. Rule-7 blind spot: unmapped GPU names now emit an explicit `[warn] ... remote rules (incl. Rule 7) skipped` instead of silently inheriting the default threshold.
Reviewer verdict trail: #689 (REQUEST CHANGES, hardcoded keys) -> #690 (correction: keys proven dead, downgrade to strip-before-merge) -> this merge. Nits NOT addressed here (tracked for a follow-up): ERR_401 overcount in health-check.sh, bare `except:` blocks, HISTORY_FILE perms coupling, hardcoded container threshold 8.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Aligns
syslog-harnesswith the verified live state of CT 116 and the GPU hosts, and removes model-version names from every client-facing surface.Why
Kwame's directive: "do not advertise models to downstream that way — we won't have to update those when we update models." Client-visible names should describe capability, not model version. While aligning, the repo also turned out to be stale in four places against live state.
What changed (all verified against running systems, not from documentation)
litellm_config.yamlsyslog-auto,gpu-dense,strix-moe,gpu-vision;fallbacks+model_costre-keyed; upstreamlitellm_params.modelset to capability names (backends ignore the model field — verified 200 on all three hosts, so no llama.cpp relaunch, no KV warmth lost)gpu_roster.yaml.8model_path corrected,max_concurrent2 → 1,context262144 → 131072;.15model_path corrected; keys renamed to capability namesREADME.md/dashboardgpu-vision → gpu-lightid mappingscripts/gemma-4-12bcalls in 7 days), andgpu-monitor.py's.110entry wrongly listed.8's modelVerification
/v1/modelsadvertises exactly the four names; each returns 200 through nginx; container healthy.model-gpu-dense,model-strix-moe,model-gpu-vision,model-syslog-autoall pass (previously two dead names).litellm_config.yaml.bak-20260912-prerenameon CT 116 +docker restart harness-litellm.Intentionally NOT changed
LITELLM-MIGRATION-PLAN.md— historical planning document (June 14, 2026), not a live-state claim.backups/,graphify-out/,litellm_config.yaml.backup— historical artifacts.router.py,docker-compose.yml.pre-1991-20260911,nginx/default.conf,dashboard/gpu-monitor.html,scripts/gitea-logger.sh) — out of scope.Note for review
.8's live server alias is stillqwen3.6-27B-codewhile the loaded weights areQwen3.8-27B-Uncensored-Q4_K_M.gguf. Renaming the alias requires a llama.cpp relaunch (and loses KV warmth), so it is deliberately left as an explicit follow-up; the roster records the actual value.Review verdict: REQUEST CHANGES — from the 2026-09-19 management-brief review (deep diff review via delegated reviewer; security posture verified live by Mumuni).
The alignment work itself checks out: no retired model names survive except the intentional allowlist migration notes, litellm_config.yaml is internally consistent with the capability-name catalogue (syslog-auto / gpu-dense / strix-moe / gpu-vision), Python scripts py_compile clean, shell scripts bash -n clean, live gateway confirms the four advertised model names match.
Blockers:
scripts/gpu-monitor.py(~L303) hardcodes a live Zulip bot API key (cKTD…), andscripts/gpu-self-heal.py(~L387) hardcodes the LiteLLM master key (sk-lit…6796). Verified: this repo is public, Gitea allows anonymous raw reads, andgit.sysloggh.netserves this PR branch from the public internet as we speak. Consequences: (a) rotate both credentials immediately — a branch delete after the fact does not un-leak them; (b) before merge, move both to the existing/etc/litellm-monitor.envpattern (no literal values in tracked files).scripts/gpu-monitor.pyL24-26:DASHBOARD_PATHhardcoded to/root/…— 404s silently if the monitor ever runs as a non-root service user. Make it configurable via env.scripts/litellm-health-check.shL12:MONITOR_KEYsilently falls back to the master key when the env file is missing — that hides misconfiguration and encourages master-key use; fail loudly instead.Repo hygiene note:
mainandmasterhave no merge base (orphaned histories) — this PR's base ambiguity is a symptom. Recommend picking one branch and deleting the other.Everything else is good to merge once the secrets are externalized and the two keys are rotated. Owner pinged via RA-H relay. — Mumuni 🦅
Follow-up: reviewer report landed + one correction to my review above (tested 23:32 UTC).
Correction on blocker #1 urgency: I wrote "rotate both credentials immediately." Re-tested by direct API call: both embedded keys are already dead — the
sk-lit…master key gets401 token_not_foundfrom the live LiteLLM (the independent reviewer reproduced this same 401 independently), and the ZulipcKTD…key gets401 Invalid API key. They're stale remnants of the 2026-09-12 alignment work; rotation evidently already happened. No live exposure, no urgency. What stands: this repo is public and the branch is anonymously fetchable from the internet (git.sysloggh.net/.../raw/branch/fix/harness-align-20260912/...→ HTTP 200), so dead-but-confusable literals in public history still must be stripped in favor of the/etc/litellm-monitor.envpattern before merge. Verdict stays REQUEST CHANGES but low priority — the reviewer's independent verdict was APPROVE_WITH_NITS ("core goal achieved cleanly… no blocker-level issues"); the secrets are the only reason I hold it.Additional findings from the deep review (fold into the same fix pass):
qwen3.6-35B-A3B (MoE)survives on the Strix Halo row — PR fixed the 3090 row but missed this one; also "Total: 6 concurrent slots" is stale vs the new roster (dense now 1 slot ⇒ total 5).host_keyfallback (name.replace(" ","-").lower()) produces values likenvidia-geforce-rtx-3090that can never match a GPU_HOSTS key, so unknown GPUs silently gethost_key='unknown'and the Rule 7 Prometheus check is skipped — a silent monitoring gap.Received=auth error, inflating the count; label or regex should tighten.FEED TO KNOWLEDGE GRAPHheading left above the correctedLOG TO GITEAcomment.except:at ~704/714/722/1065 swallow KeyboardInterrupt; tighten toexcept Exception.len(parts) >= 8but indexesparts[8]; needs>= 9guard.DASHBOARD_PATH=/root/...(already in my first comment) +HISTORY_FILE=/var/log/litellm/requires write perms — fine under root, note it in README if this ever ships as a service user.None change the substance: the alignment work is solid. Fix list = env-ize the two dead keys + README row/slot arithmetic + item 2 (the one with real operational teeth). — Mumuni 🦅
Merged 2026-09-20 under Kwame's decision: "review and merge if aligned with live state."
Pre-merge pass executed on this branch (5 commits, all validated:
ast.parse+bash -n+ secret-literal scan = 0 leaks):scripts/gpu-monitor.py— Zulip bot key literal ->os.environ['ZULIP_API_KEY']with graceful suppression when unset (key was dead anyway; re-issue for abiba-bot is on her).scripts/gpu-self-heal.py— brokensk-lit...6796literal -> env var (the literal was a truncated non-working string, so this also fixes the probe rather than just scrubbing).scripts/litellm-health-check.sh— both Bearer placeholders rewired throughLLK="${LITELLM_API_KEY:?}"guard at entry.README.md— architecture block + fleet table aligned to deployed roster args (strix-moe naming, vision 1 slot / 131K, total 4 concurrent slots).[warn] ... remote rules (incl. Rule 7) skippedinstead of silently inheriting the default threshold.Reviewer verdict trail: #689 (REQUEST CHANGES, hardcoded keys) -> #690 (correction: keys proven dead, downgrade to strip-before-merge) -> this merge. Nits NOT addressed here (tracked for a follow-up): ERR_401 overcount in health-check.sh, bare
except:blocks, HISTORY_FILE perms coupling, hardcoded container threshold 8.