diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index b4ef590..360ba94 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -208,7 +208,7 @@ Plaintext keys removed from this contract post-vault-migration. | Agent | CT | IP | LiteLLM Alias | Key Source | Access | |-------|-----|-----|---------------|------------|--------| | Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome | -| Mumuni | 114 | .123 | `mumuni` | Infisical vault | SSH root | +| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | SSH root | | Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) | | Koby | 111 | ? | `koby` | Infisical vault | Zulip DM | | Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH | @@ -256,7 +256,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. - **LiteLLM timeout tuning (verified 2026-07-27)**: SmartCode-Fable-5-27B 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). -- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. +- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116. - **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback. - **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected. @@ -301,7 +301,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent ### Mumuni Agent Profile -Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile is the reference for all agent configs: +Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This profile is the reference for all agent configs: | Setting | Value | Notes | |---------|-------|-------| @@ -326,7 +326,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i | Agent | Host | Status | |-------|------|--------| -| **Mumuni** | CT114 (.123) | ✅ Updated to stable aliases | +| **Mumuni** | CT100 (.24) | ✅ Updated to stable aliases | | **Tanko** | CT112 (.122) | ✅ Updated to stable aliases | | **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM | | **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM | diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 8653068..f5c557a 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -25,7 +25,7 @@ done | Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform | |-------|-----|------|-----|---------------|------------|----------| | Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes | -| Mumuni | 114 | hwepve | .123 | `mumuni` | Infisical vault | Hermes | +| Mumuni | 100 (abiba) | hwepve | .24 | `mumuni` | Infisical vault | Hermes (Pi gateway) | | Koby | 129 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** | | Koonimo | 114 | amdpve | ? | `koonimo` | Infisical vault | Hermes | | Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) | diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index b3d0dda..4ed9881 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -35,7 +35,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed | Agent | Key Alias | Host | SSH | Sub-Agents | |-------|-----------|------|-----|-----------| | Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — | -| Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ | +| Mumuni | `mumuni` | 192.168.68.24 (CT100 abiba) | root@.24 | 6 profiles ✱ | | Abiba | `abiba-pi` | 192.168.68.24 | local | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — | | Koonimo | `koonimo` | CT 114 (baggy) | SSH root | — | diff --git a/hermes-key-enforcement.prose.md b/hermes-key-enforcement.prose.md index 128f482..7f2bfaf 100644 --- a/hermes-key-enforcement.prose.md +++ b/hermes-key-enforcement.prose.md @@ -189,7 +189,7 @@ litellm_settings: | Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified | |-------|-----|-----|---------------|------------|--------|-----------------|---------------| | Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 | -| Mumuni | 114 | .123 | `mumuni` | Infisical vault | ✅ Fixed | `infisical run` | 01:46 EDT Jul 10 | +| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 | | Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 | | Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 | | Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 | diff --git a/hermes-zulip-restore.prose.md b/hermes-zulip-restore.prose.md index d48991c..32ae78d 100644 --- a/hermes-zulip-restore.prose.md +++ b/hermes-zulip-restore.prose.md @@ -51,7 +51,7 @@ gateway restart, and connection validation. | Host | CT | Proxmox | IP (direct) | Hermes Home | User | |------|-----|---------|-------------|-------------|------| -| Mumuni | CT114 | hwepve | 192.168.68.123 | /root/.hermes | root | +| Mumuni | CT100 (abiba) | hwepve | 192.168.68.24 | /root/.hermes | root | | Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | | Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root | | Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky | diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 2fc6bec..ad1937a 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -72,7 +72,7 @@ description: > | docker-vm (.7) | SSH root | SSH key | ✅ | | CT 116 (syslog-api) | SSH root | SSH key | ✅ | | Tanko CT (.122) | SSH jerome | id_ed25519 | ✅ | -| Mumuni CT (.123) | SSH root | id_ed25519 | ✅ | +| Mumuni (inside CT100 abiba .24) | SSH root | id_ed25519 | ✅ | | Baggy CT (113) | SSH jerome | ❌ no key access | | Netbird (.17) | SSH root | SSH key | ✅ | | Gitea | API token | Infisical vault (`GITEA_BOT_TOKEN`) | ✅ | @@ -90,11 +90,11 @@ description: > ### Reachability Matrix -| From / To | PVE API | docker-vm (.7) | CT 116 | Tanko (.122) | Mumuni (.123) | Baggy (.114) | +| From / To | PVE API | docker-vm (.7) | CT 116 | Tanko (.122) | Mumuni (.24) | Baggy (.114) | |-----------|---------|----------------|--------|-------------|---------------|----------------| | **Abiba** (.24) | ✅ :443 | ✅ SSH | ✅ SSH | ✅ SSH jerome | ✅ SSH root | ❌ SSH | | **Tanko** (.122) | ❌ | ❌ | ❌ via NetBird | ✅ | ❌ | ❌ | -| **Mumuni** (.123) | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ | +| **Mumuni** (.24) | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ | **Conclusion:** Only Abiba has cross-infrastructure access. All monitoring contracts run from Abiba. @@ -114,7 +114,7 @@ description: > > **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on > minipve at .10, not acerpve. Abiba (CT 100) is on hwepve and now runs Mumuni > Zulip gateway internally. CT 114 (old mumuni container) destroyed 2026-07-26. -> Mumuni also has a second instance on minipve at .123 — distinguish by CT ID. +> Mumuni now runs inside Abiba CT100 (.24). Old CT114 destroyed 2026-07-26. ### Checks (every 5 min) diff --git a/litellm-api-keys.prose.md b/litellm-api-keys.prose.md index c492206..71166dd 100644 --- a/litellm-api-keys.prose.md +++ b/litellm-api-keys.prose.md @@ -178,7 +178,7 @@ through its agent wrapper. | Agent | Host | Pattern | Keys | Status | |-------|------|---------|------|--------| | abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed | -| mumuni | .123 | systemd drop-in + while-true wrapper + st.8e848433 | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback | +| mumuni | .24 (CT100 abiba) | Pi Hermes gateway (no systemd) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback | | tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback | | koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed | | koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed | diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 949865f..76dced7 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -102,7 +102,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`). - **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault). -- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.123), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. +- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. - **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM. ## Maintains diff --git a/mumuni-delegation-prose-contract.prose.md b/mumuni-delegation-prose-contract.prose.md index 91d8121..69338d0 100644 --- a/mumuni-delegation-prose-contract.prose.md +++ b/mumuni-delegation-prose-contract.prose.md @@ -6,7 +6,7 @@ description: > delegation, verification, and delivery. Defines when to delegate, which worker to use for what, how to handle failures, and the kanban board protocol. Enforces context-window discipline and separation of concerns. - Runs on Mumuni (lxc/114, hwepve, .123) via Hermes agent. + Runs on Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent (Pi + Hermes Zulip gateway). version: 1.0.0 --- @@ -20,7 +20,7 @@ version: 1.0.0 ## Topology **Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve) -**Manager:** Mumuni (lxc/114, hwepve, .123) via Hermes agent +**Manager:** Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent **Workers:** 6 profiles, all running on the same agent — no separate hosts needed This contract is infrastructure-agnostic in terms of which nodes are used. diff --git a/scripts/agent-health-check.py b/scripts/agent-health-check.py index b4df68f..3f04c14 100755 --- a/scripts/agent-health-check.py +++ b/scripts/agent-health-check.py @@ -17,7 +17,7 @@ Changelog: v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity, vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}). - Fleet roster: tanko (.122), mumuni (.123), koby (.129), koonimo (.114), + Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). """ diff --git a/scripts/swap-gpu-dense-model.sh b/scripts/swap-gpu-dense-model.sh new file mode 100755 index 0000000..68c84d1 --- /dev/null +++ b/scripts/swap-gpu-dense-model.sh @@ -0,0 +1,91 @@ +#!/bin/bash +# swap-gpu-dense-model.sh — Swap RTX 3090 from qwen3.6-27B-code to SmartCode-Fable-5 +# Run when download completes: ssh root@192.168.68.8 'bash -s' < this script +# +# Usage: bash swap-gpu-dense-model.sh +# Requires: new model at /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf + +set -e + +MODEL_PATH="/home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf" +OLD_WRAPPER="/home/llmuser/llama-wrapper.sh" + +echo "═══ Swapping gpu-dense to SmartCode-Fable-5 ═══" + +# 1. Verify model file +if [ ! -f "$MODEL_PATH" ]; then + echo "❌ Model not found at $MODEL_PATH" + echo " Download: curl -L -o $MODEL_PATH " + exit 1 +fi + +MODEL_SIZE=$(ls -lh "$MODEL_PATH" | awk '{print $5}') +echo "✅ Model found: $MODEL_SIZE" + +# 2. Create new wrapper script for SmartCode-Fable-5 +cat > /home/llmuser/llama-fable-wrapper.sh << 'WRAPPER' +#!/bin/bash +# SmartCode-Fable-5 llama-server wrapper for RTX 3090 +# Sampler settings from model card: temp 0.9, top-p 0.95, top-k 60, repeat-penalty off + +PORT=8080 +GHOST_PID=$(ss -tlnp 2>/dev/null | grep -Po ":${PORT}\s+.*pid=\K[0-9]+" | head -1) +if [ -n "$GHOST_PID" ] && [ "$GHOST_PID" != "$$" ]; then + echo "[wrapper] Port $PORT occupied by ghost pid $GHOST_PID — cleaning up" >&2 + kill -9 "$GHOST_PID" 2>/dev/null + sleep 2 +fi + +exec /usr/local/bin/llama-server \ + --model /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf \ + --ctx-size 131072 \ + --cache-type-k q4_0 \ + --cache-type-v q4_0 \ + --flash-attn 1 \ + --cont-batching \ + --parallel 1 \ + --batch-size 2048 \ + --ubatch-size 1024 \ + --n-gpu-layers 99 \ + --temp 0.9 \ + --top-p 0.95 \ + --top-k 60 \ + --min-p 0.0 \ + --repeat-penalty 1.0 \ + --api-key not-needed \ + --port 8080 \ + --host 0.0.0.0 +WRAPPER + +chmod 755 /home/llmuser/llama-fable-wrapper.sh +echo "✅ Created /home/llmuser/llama-fable-wrapper.sh" + +# 3. Update systemd service to use new wrapper +echo "📝 Updating systemd service..." +sed -i 's|ExecStart=/home/llmuser/llama-wrapper.sh|ExecStart=/home/llmuser/llama-fable-wrapper.sh|' /etc/systemd/system/llama-server.service +systemctl daemon-reload + +# 4. Stop old server, start new +echo "🔄 Restarting llama-server..." +systemctl stop llama-server +sleep 3 +systemctl start llama-server +sleep 8 + +# 5. Verify +echo "" +echo "═══ Verification ═══" +systemctl is-active llama-server +echo "" +echo "Port 8080:" +ss -tlnp 2>/dev/null | grep ":8080" | head -1 +echo "" +echo "GPU VRAM:" +nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null +echo "" +echo "=== Health check ===" +curl -s --max-time 5 http://localhost:8080/health 2>/dev/null +echo "" +echo "" +echo "✅ Swap complete. Test via LiteLLM:" +echo " curl -s http://192.168.68.116/v1/chat/completions -H 'Authorization: Bearer ' -H 'Content-Type: application/json' -d '{\"model\":\"gpu-dense\",\"messages\":[{\"role\":\"user\",\"content\":\"write hello world in python\"}],\"max_tokens\":100}'" diff --git a/zulip-health.prose.md b/zulip-health.prose.md index 876fd2e..5737799 100644 --- a/zulip-health.prose.md +++ b/zulip-health.prose.md @@ -16,7 +16,7 @@ Runs every 15 minutes in the background. Also triggers on session start. ## Requires - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` -- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.123, lxc/114 on hwepve since 2026-07-20), and Agent Zero Docker host (192.168.68.14) +- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on hwepve), and Agent Zero Docker host (192.168.68.14) - **PM2** on localhost for pi process management - **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` @@ -182,13 +182,13 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta | `last_error` set | Log and monitor | | Crash loop >10/h | Alert user | -### Step 3: Platform B — Hermes (Tanko .122, Mumuni .123) +### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24) **B1: Gateway State** ```bash ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json" -ssh root@192.168.68.123 "cat ~/.hermes/gateway_state.json" +ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100 ``` Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed. diff --git a/zulip-self-heal.prose.md b/zulip-self-heal.prose.md index 9ee024c..2a7474b 100644 --- a/zulip-self-heal.prose.md +++ b/zulip-self-heal.prose.md @@ -65,7 +65,7 @@ triggers: |------|----|------|---------| | Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` | | Abiba (pi) | localhost | root | PM2: `abiba-zulip` | -| Mumuni | 192.168.68.123 | root | `hermes gateway restart` | +| Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` | | Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` | ## Debounce