diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 028bcad..aec6746 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -100,6 +100,10 @@ model: api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). + # ⚠️ MANDATORY: Hermes probes unknown models from 256K + # and falls back to 256K when /v1/models lacks a context + # field (llama-server does). Without this override, agents + # silently run syslog-auto at 256K (verified 2026-08-09). # Set 65536 if using gemma-4-12b directly (tight VRAM). fallback_providers: diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index c55f98e..bb52406 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -7,13 +7,15 @@ description: > from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics via existing /metrics Prometheus endpoint. - DEPLOYMENT STATUS (2026-07-09): + DEPLOYMENT STATUS (2026-08-09): ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters - (via proxmox-monitor contract). Grafana at :3001, 5 scrape targets active. - ❌ GPU exporters NOT deployed: gpu-exporter crash-loops on .15, - NVIDIA sidecar exporters (.8/.110:9400) never installed. - Router falls back to direct GPU /health probes. - ⚠️ This contract is target-state aspirational — not as-built. + (via proxmox-monitor contract). Grafana at :3001, all scrape targets active. + ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on + :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. + ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via + master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. + ⚠️ This contract is target-state aspirational — but GPU export + alerting + are now as-built (verified 2026-08-09). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). version: 1.0.0 --- diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 615ab3c..9b8c904 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -18,6 +18,15 @@ description: > Designed as a reusable contract for any Syslog agent. Source of truth: gpu-fleet.prose.md + + ## Monitoring / Alerting (as-built 2026-08-09) + + - LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]` + (failure_callback alone does NOT mount /metrics — verified 2026-08-09). + Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/. + - Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver + firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09. + - Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). --- ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)