Commit Graph
10 Commits
Author SHA1 Message Date
mumuni-bot 80811511e7 fix(harness): align repo with verified live state; retire model-version names from the client surface
litellm_config.yaml
- client-visible model_list reduced to capability names: syslog-auto, gpu-dense, strix-moe, gpu-vision
- retired qwen3.6-27B-code, qwen3.8-27B-uncensored, qwen3.6-35B-udq4 (and the already-dead gemma-4-12b)
- fallbacks and model_cost re-keyed to the surviving names
- upstream litellm_params.model set to the capability names; the backends ignore the model field
  (verified HTTP 200 on all three hosts), so this needs no llama.cpp relaunch and loses no KV warmth
- verified live: /v1/models advertises exactly the four names, each returns 200 through nginx

gpu_roster.yaml
- launch args replaced with the VERIFIED live command lines, read from the running processes via the
  PVE guest agent (acerpve VM 101 = .8, ocupve VM 103 = .110) and on amdpve (.15)
- .8: model_path corrected to Qwen3.8-27B-Uncensored-Q4_K_M.gguf, max_concurrent 2 -> 1,
  context 262144 -> 131072, full arg list recorded (incl. --parallel 1 and the new
  --slot-save-path / --metrics added 2026-09-12)
- .15: model_path corrected to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf, full arg list recorded
- keys renamed to the capability names; hosts.current_model aligned

README.md / dashboard
- README dense entry corrected (1 slot, 131K ctx, actual model file)
- dashboard picker now uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070
  serves a 9B Qwen3.5) and the gpu-vision -> gpu-light id mapping

scripts/
- added the three operational monitors as tracked files (they were untracked): gpu-monitor.py,
  gpu-self-heal.py, litellm-health-check.sh
- cleared their references to retired model names, which were causing failed calls every benchmark
  cycle (150 failed gemma-4-12b calls in the last 7 days); gpu-monitor.py's .110 entry also wrongly
  listed .8's model

Intentionally NOT changed
- LITELLM-MIGRATION-PLAN.md: historical planning document (June 14), not a live-state claim
- backups/, graphify-out/, litellm_config.yaml.backup: historical artifacts
- unrelated untracked files (router.py, docker-compose.yml.pre-1991-20260911, nginx/default.conf,
  dashboard/gpu-monitor.html, scripts/gitea-logger.sh): out of scope for this change
2026-09-12 21:03:48 +00:00
agent-zero 93e7605aa2 chore(ct116): capture production state (LiteLLM 1.99.1, nginx /ui //docs, router decommission, trove agent) 2026-09-11 19:34:30 +00:00
Abiba 85608d7c60 Dashboard + LiteLLM config updates from maintenance 2026-06-07 22:49:50 +00:00
Abiba 9a633583ab fix: Security hardening from CT116 deep-dive review
- API keys moved from hardcoded dict to env var (API_KEYS JSON) with fallback
- Rate limiting added: token bucket per API key (Redis-backed), 429 responses with Retry-After
- Rate limit tiers: enterprise 120/min, professional 60/min, starter 20/min
- X-RateLimit-* headers on all responses
- Dashboard polling reduced from 3s to 5s backend, 10s JS fallback
- SSE detection disables redundant polling when stream is connected
- Deleted ts_patch.py (dead one-shot migration, already applied)
- Added ssl/README.md documenting upstream SSL termination

Ref: Relay #444 (Mumuni CT116 harness deep-dive)
Reviewed-by: Abiba <abiba@sysloggh.com>
2026-06-02 10:37:10 +00:00
Abiba 80362fa528 fix: default performance window to 24h so all models appear immediately 2026-05-26 12:37:52 +00:00
Abiba 7ef9e58f61 fix: restore /api/performance route in dashboard (was overwritten to /api/timeseries) 2026-05-26 12:31:53 +00:00
Abiba f47c3f3304 feat: latency vs prompt size scatter plot on dashboard
Router: new /metrics/scatter endpoint returns individual data points
(prompt_tokens, inference_ms, model, agent, reason, stream)
for scatter visualization.

Dashboard: new panel showing latency vs prompt size by model.
- Log-scale X axis (prompt tokens) with model color coding
- Dropdown to filter by individual model or view all
- Hover tooltips with details per point
- Auto-refresh every 30s

Enables direct observation of context-length vs latency
relationship — validates routing tier decisions.
2026-05-26 12:18:31 +00:00
Abiba b2ec4b0572 fix: throughput panel handles streaming-only models gracefully
- Dashboard: when a model has zero non-streaming records, shows
  "streaming only" instead of misleading 0 tok/s
- Dashboard: minimum bar width enforced (6% avg, 4% p50) so
  low-tps models are always visible
- Router: removed inflated streaming tps estimate (prompt tokens
  skewed results for long conversations)

Fixes Dense model appearing to "register nothing" when Mumuni
sends mostly streaming requests.
2026-05-25 19:45:21 +00:00
Abiba f42747d721 feat: performance analytics panel on dashboard
dashboard/dashboard.py (+61 lines):
- New /api/performance endpoint proxying to router metrics/performance
- Performance Analytics row with 4 panels:
  - Latency distribution (p50/p95/p99 per model) with stacked bars
  - Throughput comparison (avg + p50 tokens/sec per model)
  - Routing effectiveness table by reason
  - Agent performance bars with latency
- 1h/24h window toggle, auto-refresh every 15s
- Color-coded per model (purple=MoE, amber=Dense, green=VLM)
2026-05-25 16:58:15 +00:00
Abiba 28fc57c5c7 May 19, 2026: Full harness update
- Model migration: gemma-4-E4B → qwen3.5-9b-vlm
- Dashboard reorder: Usage Over Time + GPU Metrics to top
- Router counter leak fix (gpu_decr in except handler)
- VLM slot upgrade 1→2
- Automated maintenance cron job
- LiteLLM config update
2026-05-19 15:03:47 +00:00