Ship: monitoring contract fixes (2026-08-09) #43
@@ -100,6 +100,10 @@ model:
|
|||||||
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||||
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
|
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
|
||||||
|
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
|
||||||
|
# and falls back to 256K when /v1/models lacks a context
|
||||||
|
# field (llama-server does). Without this override, agents
|
||||||
|
# silently run syslog-auto at 256K (verified 2026-08-09).
|
||||||
# Set 65536 if using gemma-4-12b directly (tight VRAM).
|
# Set 65536 if using gemma-4-12b directly (tight VRAM).
|
||||||
|
|
||||||
fallback_providers:
|
fallback_providers:
|
||||||
|
|||||||
@@ -7,13 +7,15 @@ description: >
|
|||||||
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
|
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
|
||||||
via existing /metrics Prometheus endpoint.
|
via existing /metrics Prometheus endpoint.
|
||||||
|
|
||||||
DEPLOYMENT STATUS (2026-07-09):
|
DEPLOYMENT STATUS (2026-08-09):
|
||||||
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
|
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
|
||||||
(via proxmox-monitor contract). Grafana at :3001, 5 scrape targets active.
|
(via proxmox-monitor contract). Grafana at :3001, all scrape targets active.
|
||||||
❌ GPU exporters NOT deployed: gpu-exporter crash-loops on .15,
|
✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on
|
||||||
NVIDIA sidecar exporters (.8/.110:9400) never installed.
|
:9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09.
|
||||||
Router falls back to direct GPU /health probes.
|
✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via
|
||||||
⚠️ This contract is target-state aspirational — not as-built.
|
master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09.
|
||||||
|
⚠️ This contract is target-state aspirational — but GPU export + alerting
|
||||||
|
are now as-built (verified 2026-08-09).
|
||||||
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
||||||
version: 1.0.0
|
version: 1.0.0
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -18,6 +18,15 @@ description: >
|
|||||||
Designed as a reusable contract for any Syslog agent.
|
Designed as a reusable contract for any Syslog agent.
|
||||||
|
|
||||||
Source of truth: gpu-fleet.prose.md
|
Source of truth: gpu-fleet.prose.md
|
||||||
|
|
||||||
|
## Monitoring / Alerting (as-built 2026-08-09)
|
||||||
|
|
||||||
|
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
|
||||||
|
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
|
||||||
|
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||||
|
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||||
|
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||||
|
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
|
||||||
---
|
---
|
||||||
|
|
||||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||||
|
|||||||
Reference in New Issue
Block a user