Health-check script: critical-only gpu-fleet alerting, strix-moe model test. gpu-monitor: systemd unit (was bare &), Strix counted in gpu_count, VRAM thresholds raised (93/97 — 256K ctx steady-state ~96% on RTX 3090). agent-health-check.py: reads live LITELLM_API_KEY from gateway env (no hardcoded keys). daily-infra-report.py: stale SYNTHETIC key → env-based master key.
Verified on ground 2026-07-16 against CT 116 litellm_config.yaml + GPU hosts: - AMD host serves qwen3.6-35B-udq4 (LiteLLM alias strix-moe); ornith-1.0-35b does NOT exist - All 3 GPUs at 256K ctx, parallel 2 (RTX 3090 was listed 128K/parallel 1) - LiteLLM timeouts: qwen 300s, gemma 120s, strix 300s (were stale 90s/120s) - Added LiteLLM model surface + key scoping to litellm-self-heal - Patched health-check script path ref Files: litellm-self-heal, litellm-health, gpu-fleet, gpu-self-heal, zulip-adapter-lessons, abiba-zulip-restore, hermes-agent-baseline, delegation-prose-contract, mumuni-delegation-prose-contract