From dc42ecc235816bde0b22e12fc9149bf10b37919f Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 15:56:16 +0000 Subject: [PATCH] no-mistakes(document): Align gpu-vision and single-source-of-truth documentation --- README.md | 2 +- docs/AUTHORING-GUIDE.md | 2 +- gpu-fleet.prose.md | 7 +++---- 3 files changed, 5 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index 46b8968..ac61914 100644 --- a/README.md +++ b/README.md @@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4 # Configure an agent with a different auxiliary model -prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b +prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision ``` ### Option B: Manual Execution diff --git a/docs/AUTHORING-GUIDE.md b/docs/AUTHORING-GUIDE.md index f73129e..adf2f78 100644 --- a/docs/AUTHORING-GUIDE.md +++ b/docs/AUTHORING-GUIDE.md @@ -13,7 +13,7 @@ the description: 1. **What system does this contract touch?** Name the hosts, CTs, containers, and services explicitly. "The inference fleet" is vague. "GPU .8 (RTX 3090, - qwen), .110 (RTX 5070, gemma), .15 (Strix Halo, strix-moe), and LiteLLM on CT + qwen), .110 (RTX 5070, gpu-vision), .15 (Strix Halo, strix-moe), and LiteLLM on CT 116" is specific. 2. **Who runs this contract, and when?** State the agent, the trigger (cron, diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 04cd136..fefebba 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -25,7 +25,7 @@ triggers: ## Maintains -- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models +- gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116 `litellm_config.yaml` - router: { status: "healthy", roster_loaded: bool, models: array } - litellm: { status: "healthy", keys: array, models: array } - agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB @@ -97,9 +97,8 @@ is retired and returns 400 `Invalid model name`. ## Routing Configuration (LiteLLM — July 2026) -Single source of truth for models, aliases, rpm caps, weights and fallback chains: -CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in -contracts — read them there. +Model, alias, rpm/weight and fallback values are owned by CT 116 +`/opt/inference-harness/litellm_config.yaml` (see § Stable Role-Based Aliases above). Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.