Phase 0.5: Re-add ADMIN_KEY, /admin/keys endpoints, dual-key logging #1

Open
mumuni-bot wants to merge 22 commits from SyslogSolution/syslog-harness:main into main
22 Commits
Author SHA1 Message Date
mumuni-bot 82464d2158 README table: gpu-vision slots/ctx aligned to roster (parallel 1, 131K) 2026-09-20 15:33:31 +00:00
mumuni-bot 4c6cf66aed Merge pull request 'Align harness repo with verified live state; retire model-version names from the client surface' (#2) from fix/harness-align-20260912 into main 2026-09-20 15:32:07 +00:00
mumuni-bot afccc29034 gpu-self-heal: repair Rule-7 warning line (bad quoting/indent from prior patch) 2026-09-20 15:31:52 +00:00
mumuni-bot 502f2d17c3 health-check: route remaining bearer placeholders through $LLK env guard 2026-09-20 15:29:42 +00:00
mumuni-bot 973335d35b security+alignment pass from PR review: env-var keys (strip dead literals), Rule-7 unmapped-GPU warning, README slots vs live roster args 2026-09-20 15:27:57 +00:00
mumuni-bot b4f94bf672 security+alignment pass from PR review: env-var keys (strip dead literals), Rule-7 unmapped-GPU warning, README slots vs live roster args 2026-09-20 15:27:53 +00:00
mumuni-bot 4f3a28ff10 security+alignment pass from PR review: env-var keys (strip dead literals), Rule-7 unmapped-GPU warning, README slots vs live roster args 2026-09-20 15:27:52 +00:00
mumuni-bot f893e4f611 security+alignment pass from PR review: env-var keys (strip dead literals), Rule-7 unmapped-GPU warning, README slots vs live roster args 2026-09-20 15:27:51 +00:00
mumuni-bot 80811511e7 fix(harness): align repo with verified live state; retire model-version names from the client surface
litellm_config.yaml
- client-visible model_list reduced to capability names: syslog-auto, gpu-dense, strix-moe, gpu-vision
- retired qwen3.6-27B-code, qwen3.8-27B-uncensored, qwen3.6-35B-udq4 (and the already-dead gemma-4-12b)
- fallbacks and model_cost re-keyed to the surviving names
- upstream litellm_params.model set to the capability names; the backends ignore the model field
  (verified HTTP 200 on all three hosts), so this needs no llama.cpp relaunch and loses no KV warmth
- verified live: /v1/models advertises exactly the four names, each returns 200 through nginx

gpu_roster.yaml
- launch args replaced with the VERIFIED live command lines, read from the running processes via the
  PVE guest agent (acerpve VM 101 = .8, ocupve VM 103 = .110) and on amdpve (.15)
- .8: model_path corrected to Qwen3.8-27B-Uncensored-Q4_K_M.gguf, max_concurrent 2 -> 1,
  context 262144 -> 131072, full arg list recorded (incl. --parallel 1 and the new
  --slot-save-path / --metrics added 2026-09-12)
- .15: model_path corrected to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf, full arg list recorded
- keys renamed to the capability names; hosts.current_model aligned

README.md / dashboard
- README dense entry corrected (1 slot, 131K ctx, actual model file)
- dashboard picker now uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070
  serves a 9B Qwen3.5) and the gpu-vision -> gpu-light id mapping

scripts/
- added the three operational monitors as tracked files (they were untracked): gpu-monitor.py,
  gpu-self-heal.py, litellm-health-check.sh
- cleared their references to retired model names, which were causing failed calls every benchmark
  cycle (150 failed gemma-4-12b calls in the last 7 days); gpu-monitor.py's .110 entry also wrongly
  listed .8's model

Intentionally NOT changed
- LITELLM-MIGRATION-PLAN.md: historical planning document (June 14), not a live-state claim
- backups/, graphify-out/, litellm_config.yaml.backup: historical artifacts
- unrelated untracked files (router.py, docker-compose.yml.pre-1991-20260911, nginx/default.conf,
  dashboard/gpu-monitor.html, scripts/gitea-logger.sh): out of scope for this change
2026-09-12 21:03:48 +00:00
abiba-bot 4b86629751 Merge pull request 'Reconcile upstream into CT116 production state' (#1) from sync/ct116-reconcile-20260912-095913 into main 2026-09-12 10:22:57 +00:00
agent-zero e4eced2cf9 merge: reconcile upstream main (937ba2c, 9acabf7) into production state
# Conflicts:
#	docker-compose.yml
#	litellm_config.yaml
#	nginx/nginx.conf
2026-09-12 09:59:14 +00:00
agent-zero 8afe94a7a9 chore(ct116): externalize secrets to .env; prune dead router nginx stubs 2026-09-12 09:59:14 +00:00
agent-zero 93e7605aa2 chore(ct116): capture production state (LiteLLM 1.99.1, nginx /ui //docs, router decommission, trove agent) 2026-09-11 19:34:30 +00:00
Abiba 7cad063e27 fix(routing): restore gpu-dense (.8 qwen3.6-27B-code) as syslog-auto 0.55-weight member
The Aug 24 rename moved the syslog-auto top-weight member to .110/gpu-vision,
bypassing gpu-dense entirely (0 requests/hr while strix-moe and .110 absorbed
everything). Restore the intended 55/30/15 split: gpu-dense, strix-moe, .110.
2026-08-28 14:53:51 +00:00
jerome 937ba2c1ce Archive 688 GPU self-heal logs from shared memory 2026-07-18 00:36:06 -04:00
Abiba 9acabf7ba6 docs: update LiteLLM key distribution with verified keys (2026-07-01)
- Replaced placeholder keys with actual verified LiteLLM virtual keys
- Tanko & Mumuni: confirmed working keys from agent configs
- Abiba, Kagenz0, Koby, Koonimo: keys regenerated after old values lost
- Added key rotation history
- All stale/blocked/duplicate keys purged
- Added verification status column
2026-07-01 22:09:49 +00:00
Abiba ce2b90ed50 feat: daily cost snapshot script with cron 2026-06-28 19:05:32 +00:00
Abiba 51f0840426 feat: per-agent API keys + model pricing + nginx pass-through 2026-06-28 19:03:41 +00:00
Abiba a97c754213 fix(nginx): add default Litellm UI auth key, keep agent pass-through on /v1/ 2026-06-28 17:18:43 +00:00
Abiba 765d53762c fix(nginx): forward Authorization to LiteLLM + remove dead GPU_MOE_URL 2026-06-28 16:27:04 +00:00
Abiba e9aa220ee8 feat(gpu-fleet): roster-driven GPU config + Ornith-1.0-35B 2026-06-28 16:24:54 +00:00
Abiba 8e0f6e407b chore: clean up backup files from tracking, fix remote to main repo 2026-06-28 01:36:12 +00:00