contract updates: ornith decommissioned, 256K→128K context, Mumuni Discord disabled
Change 1: Strix Halo — ornith decommissioned - gpu-fleet: Genesis Hermes V3 APEX → qwen3.6-35B-udq4 throughout - inference-optimization: ornith→strix-moe/qwen3.6-35B-udq4 - gpu-monitor: ornith status → Strix Halo status - infrastructure-control: Strix Halo — ornith → qwen3.6-35B-udq4 - infrastructure-update: ornith→strix-moe via router - proxmox-monitor: Strix Halo LLM (ornith) → (qwen3.6-35B-udq4, strix-moe) Change 2: GPU context 256K→128K fleet-wide - hermes-agent-baseline: frontmatter description updated - litellm-health: GPU Fleet Topology table 256K→128K - litellm-self-heal: GPU Fleet Topology, engine flags, VRAM alert - inference-optimization: compress threshold 256K→128K compact at 85K - gpu-fleet: instability note updated Change 3: Mumuni Discord platform disabled - gpu-fleet: Mumuni platforms: removed discord
This commit is contained in:
@@ -47,7 +47,7 @@ poll .15:8080 directly; must go through router on .116.
|
||||
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
|
||||
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
||||
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
||||
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | ornith status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
||||
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
||||
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
|
||||
|
||||
### Alert Delivery
|
||||
|
||||
Reference in New Issue
Block a user