Compare commits
10
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
c380196fab | ||
|
|
b17c60f997 | ||
|
|
2f0c3c1850 | ||
|
|
cedbdc465d | ||
|
|
8b09c78efd | ||
|
|
767ab25128 | ||
|
|
75381f9737 | ||
|
|
2b9b545ca9 | ||
|
|
e6c52bf071 | ||
|
|
1c44bf1259 |
@@ -137,7 +137,7 @@ Encode topology, architectural decisions, and lessons. You read them before plan
|
|||||||
|
|
||||||
| Contract | Description |
|
| Contract | Description |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `infrastructure-control` | Full topology and control pattern: 6-node Proxmox cluster, 3 Docker ecosystems, NFS storage, network verification, IP-first configuration doctrine. Live-state fields marked `VERIFY-BEFORE-USE`. |
|
| `infrastructure-control` | Full topology and control pattern: 5-node Proxmox cluster, 3 Docker ecosystems, NFS storage, network verification, IP-first configuration doctrine. Live-state fields marked `VERIFY-BEFORE-USE`. |
|
||||||
| `pi-approval-architecture` | pi's approval model vs Hermes, available commands, architectural constraints. |
|
| `pi-approval-architecture` | pi's approval model vs Hermes, available commands, architectural constraints. |
|
||||||
| `zulip-adapter-lessons` | Failure modes, fixes, and patterns from building the pi Zulip extension and Hermes Zulip plugin. |
|
| `zulip-adapter-lessons` | Failure modes, fixes, and patterns from building the pi Zulip extension and Hermes Zulip plugin. |
|
||||||
|
|
||||||
|
|||||||
@@ -286,7 +286,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
|||||||
| 111 | tdunna | amdpve | ✅ reachable |
|
| 111 | tdunna | amdpve | ✅ reachable |
|
||||||
| 112 | tanko | amdpve | ✅ reachable |
|
| 112 | tanko | amdpve | ✅ reachable |
|
||||||
| 113 | baggy | amdpve | ✅ reachable |
|
| 113 | baggy | amdpve | ✅ reachable |
|
||||||
| 114 | mumuni | hwepve | ✅ reachable |
|
| 114 | mumuni | minipve | ✅ reachable |
|
||||||
| 115 | scottdenya | amdpve | ✅ reachable |
|
| 115 | scottdenya | amdpve | ✅ reachable |
|
||||||
| 116 | syslog-api | minipve | ✅ reachable |
|
| 116 | syslog-api | minipve | ✅ reachable |
|
||||||
| 117 | zulip | storepve | ✅ reachable |
|
| 117 | zulip | storepve | ✅ reachable |
|
||||||
|
|||||||
@@ -185,7 +185,7 @@ what, and why should I care?
|
|||||||
|
|
||||||
```
|
```
|
||||||
❌ "Monitors infrastructure health"
|
❌ "Monitors infrastructure health"
|
||||||
✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker
|
✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker
|
||||||
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
|
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
+9
-10
@@ -10,9 +10,8 @@ description: >
|
|||||||
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
||||||
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
|
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
|
||||||
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
||||||
Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
|
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
|
||||||
Hermes agent fine-tune, tensor repair, multimodal with mmproj).
|
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||||
Instability observed near 100K at 256K. 128K is the stable ceiling.
|
|
||||||
For larger context needs → fall back to external providers (deepseek).
|
For larger context needs → fall back to external providers (deepseek).
|
||||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||||
agent: abiba
|
agent: abiba
|
||||||
@@ -99,7 +98,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
|||||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||||
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
|
||||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
||||||
| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
|
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
|
||||||
|
|
||||||
## Routing Configuration (LiteLLM — July 2026)
|
## Routing Configuration (LiteLLM — July 2026)
|
||||||
|
|
||||||
@@ -108,7 +107,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
|||||||
| Model | GPU | Weight | RPM Cap | Timeout |
|
| Model | GPU | Weight | RPM Cap | Timeout |
|
||||||
|-------|-----|--------|---------|---------|
|
|-------|-----|--------|---------|---------|
|
||||||
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||||
| Genesis Hermes V3 APEX | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
|
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
|
||||||
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
|
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
|
||||||
|
|
||||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
||||||
@@ -117,7 +116,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
|||||||
|
|
||||||
| Model | RPM Cap | Notes |
|
| Model | RPM Cap | Notes |
|
||||||
|-------|---------|-------|
|
|-------|---------|-------|
|
||||||
| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
|
||||||
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
|
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
|
||||||
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
|
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
|
||||||
|
|
||||||
@@ -252,7 +251,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
|
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
|
||||||
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
|
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
|
||||||
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
|
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
|
||||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
|
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||||
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||||
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
|
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
|
||||||
@@ -268,9 +267,9 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
|-----|-------|-----------|--------------|----------|---------|
|
|-----|-------|-----------|--------------|----------|---------|
|
||||||
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
|
||||||
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
|
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
|
||||||
| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
|
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
||||||
|
|
||||||
Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||||
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
||||||
|
|
||||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
||||||
@@ -316,7 +315,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
|
|||||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||||
| `personalities` | `creative` | Creative assistant personality |
|
| `personalities` | `creative` | Creative assistant personality |
|
||||||
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
||||||
| Main model timeout | 300s | LiteLLM global timeout |
|
| Main model timeout | 300s | LiteLLM global timeout |
|
||||||
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
|
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
|
||||||
|
|
||||||
|
|||||||
@@ -47,7 +47,7 @@ poll .15:8080 directly; must go through router on .116.
|
|||||||
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
|
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
|
||||||
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
||||||
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
||||||
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | ornith status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
||||||
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
|
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
|
||||||
|
|
||||||
### Alert Delivery
|
### Alert Delivery
|
||||||
|
|||||||
@@ -46,7 +46,7 @@ depends_on:
|
|||||||
|-------|-----|------|-------|------|-----|-------|------|
|
|-------|-----|------|-------|------|-----|-------|------|
|
||||||
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
|
||||||
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
||||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||||
|
|
||||||
Key notes:
|
Key notes:
|
||||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||||
@@ -143,7 +143,7 @@ Key notes:
|
|||||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
||||||
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
||||||
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
||||||
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
|
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
- If tok/s > baseline → context has headroom, consider increasing
|
- If tok/s > baseline → context has headroom, consider increasing
|
||||||
- If tok/s < 90% baseline → reduce context by 25% and retest
|
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||||
|
|||||||
@@ -5,7 +5,7 @@ version: 1.0.0
|
|||||||
description: >
|
description: >
|
||||||
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
||||||
configuration state, keys, workarounds, and audit procedure. When an agent's
|
configuration state, keys, workarounds, and audit procedure. When an agent's
|
||||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||||
author: Abiba (pi agent)
|
author: Abiba (pi agent)
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -272,7 +272,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
|
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
|
||||||
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
|
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
|
||||||
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
|
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
|
||||||
- **Strix Halo (64GB, 128K ctx, Geneis Hermes V3 APEX)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
|
- **Strix Halo (64GB, 128K ctx, qwen3.6-35B-udq4)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
|
||||||
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||||
- `auxiliary.vision.model: gpu-light` (RTX 5070)
|
- `auxiliary.vision.model: gpu-light` (RTX 5070)
|
||||||
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
|
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
|
||||||
|
|||||||
@@ -22,14 +22,14 @@ management, and prompt caching — without sacrificing agent capability.
|
|||||||
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
|
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
|
||||||
.123, any others on .129/.122) including compression, model, context_window,
|
.123, any others on .129/.122) including compression, model, context_window,
|
||||||
prompt_caching, memory settings
|
prompt_caching, memory settings
|
||||||
- `gpu-health`: health check response from all 3 GPU backends (ornith .15:8080,
|
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
|
||||||
qwen .8:8080, gemma .110:8080)
|
qwen .8:8080, gemma .110:8080)
|
||||||
|
|
||||||
### Maintains
|
### Maintains
|
||||||
|
|
||||||
The optimized inference stack configuration — every change is applied and
|
The optimized inference stack configuration — every change is applied and
|
||||||
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
|
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
|
||||||
non-ornith traffic; ≤ 30000 for ornith-bound agentic calls.
|
inference calls.
|
||||||
|
|
||||||
#### liteLLM-routing
|
#### liteLLM-routing
|
||||||
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
|
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
|
||||||
@@ -53,18 +53,17 @@ duration.
|
|||||||
|
|
||||||
### Strategies
|
### Strategies
|
||||||
|
|
||||||
**Context is the root cause.** Every ~46K prompt token costs ~87s of ornith
|
**Context is the root cause.** Every ~46K prompt token costs ~87s of
|
||||||
prefill time at 532 tok/s. Fix context first, routing second.
|
prefill time at 532 tok/s. Fix context first, routing second.
|
||||||
|
|
||||||
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
|
- **Route by task**: qwen for code/standard queries; gemma for
|
||||||
queries; gemma for compression/auxiliary. Never send simple completion to a
|
compression/auxiliary; strix-moe for compression tasks.
|
||||||
35B MoE.
|
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
|
||||||
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
|
compact at 51K, not 85K. Target 15% tail (not 30%).
|
||||||
compact at 102K, not 166K. Target 15% tail (not 30%).
|
|
||||||
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
||||||
never change between turns. Single-digit cache hit rate is unacceptable.
|
never change between turns. Single-digit cache hit rate is unacceptable.
|
||||||
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
||||||
GPUs reduced from 256K to 128K (2026-07-17). For larger contexts, route to external providers.
|
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||||
|
|
||||||
### Shape
|
### Shape
|
||||||
|
|
||||||
@@ -100,5 +99,5 @@ call enable-prompt-caching
|
|||||||
|
|
||||||
call verify-latency
|
call verify-latency
|
||||||
host: 192.168.68.116
|
host: 192.168.68.116
|
||||||
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b, ornith-1.0-35b]
|
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -12,6 +12,11 @@ description: >
|
|||||||
never mutate infrastructure based on them without first confirming
|
never mutate infrastructure based on them without first confirming
|
||||||
against the live system. Policy fields are authoritative. See the
|
against the live system. Policy fields are authoritative. See the
|
||||||
`verify-before-mutate` skill.
|
`verify-before-mutate` skill.
|
||||||
|
|
||||||
|
**Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110),
|
||||||
|
AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve),
|
||||||
|
Abiba placement (hwepve not amdpve), added hwepve as 6th node,
|
||||||
|
added dns.sysloggh.net route. See data/learnings.md.
|
||||||
---
|
---
|
||||||
|
|
||||||
# Infrastructure Control Pattern
|
# Infrastructure Control Pattern
|
||||||
@@ -99,12 +104,17 @@ description: >
|
|||||||
|
|
||||||
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
||||||
|------|----|-----|-----|---------|------|
|
|------|----|-----|-----|---------|------|
|
||||||
| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, jitsi | Auth, git, messaging |
|
| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||||
| amdpve | .15 | 32C | 62GB | abiba, kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute |
|
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
|
||||||
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, zulip | Docker, storage, chat |
|
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
|
||||||
| acerpve | .9 | 28C | 31GB | llm-gpu, adguard | GPU VMs |
|
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
|
||||||
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
||||||
| hwepve | .4 | 12C | 15GB | mumuni | Huawei Matebook 16, agent host |
|
| hwepve | .4 | ? | ? | abiba, (mumuni CT 114 stopped) | Agents (new node) |
|
||||||
|
|
||||||
|
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
|
||||||
|
> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni
|
||||||
|
> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a
|
||||||
|
> second instance on minipve at .123 — distinguish by CT ID, not hostname.
|
||||||
|
|
||||||
### Checks (every 5 min)
|
### Checks (every 5 min)
|
||||||
|
|
||||||
@@ -118,7 +128,7 @@ description: >
|
|||||||
|
|
||||||
## Checks
|
## Checks
|
||||||
|
|
||||||
- For each node in [minipve, amdpve, storepve, acerpve, ocupve, hwepve]:
|
- For each node in [minipve, amdpve, storepve, acerpve, ocupve]:
|
||||||
- GET /api2/json/nodes/{node}/status → check status == "online"
|
- GET /api2/json/nodes/{node}/status → check status == "online"
|
||||||
- GET /api2/json/nodes/{node}/status → cpu < 0.80
|
- GET /api2/json/nodes/{node}/status → cpu < 0.80
|
||||||
- GET /api2/json/nodes/{node}/status → free_mem > 10%
|
- GET /api2/json/nodes/{node}/status → free_mem > 10%
|
||||||
@@ -200,7 +210,7 @@ description: >
|
|||||||
**Prometheus targets**:
|
**Prometheus targets**:
|
||||||
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
||||||
- 192.168.68.110:9400 (RTX 5070 — gemma)
|
- 192.168.68.110:9400 (RTX 5070 — gemma)
|
||||||
- 192.168.68.15:9400 (Strix Halo — ornith)
|
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
||||||
- 192.168.68.24:9401 (Router metrics exporter)
|
- 192.168.68.24:9401 (Router metrics exporter)
|
||||||
- harness-litellm:4000 (LiteLLM health)
|
- harness-litellm:4000 (LiteLLM health)
|
||||||
|
|
||||||
@@ -356,15 +366,17 @@ fine. Services that resolve directly to a LAN IP are NetBird-independent.
|
|||||||
| Proxmox API | minipve.sysloggh.net:8006 | 192.168.68.12 | LAN IP | No | ✅ |
|
| Proxmox API | minipve.sysloggh.net:8006 | 192.168.68.12 | LAN IP | No | ✅ |
|
||||||
| LiteLLM | litellm.sysloggh.net | 192.168.68.116 | LAN IP | No | ✅ |
|
| LiteLLM | litellm.sysloggh.net | 192.168.68.116 | LAN IP | No | ✅ |
|
||||||
| Authentik | auth.sysloggh.net:443 | 192.168.68.11 | CNAME → netbird | **Yes** | ⚠️ |
|
| Authentik | auth.sysloggh.net:443 | 192.168.68.11 | CNAME → netbird | **Yes** | ⚠️ |
|
||||||
| Gitea | git.sysloggh.net:443 | 192.168.68.110 | CNAME → netbird | **Yes** | ⚠️ |
|
| Gitea | git.sysloggh.net:443 | 192.168.68.17:3000 | CNAME → netbird | **Yes** | ⚠️ |
|
||||||
| Zulip | chat.sysloggh.net:443 | 192.168.68.19 | CNAME → netbird | **Yes** | ⚠️ VERIFY-BEFORE-USE |
|
| Zulip | chat.sysloggh.net:443 | 192.168.68.19 | CNAME → netbird | **Yes** | ⚠️ VERIFY-BEFORE-USE |
|
||||||
| Pulse | pulse.sysloggh.net:443 | 192.168.68.7 | CNAME → netbird | **Yes** | ⚠️ |
|
| Pulse | pulse.sysloggh.net:443 | 192.168.68.7 | CNAME → netbird | **Yes** | ⚠️ |
|
||||||
| DNS UI | dns.sysloggh.net:443 | 192.168.68.102 | CNAME → netbird | **Yes** | ⚠️ |
|
| DNS UI | dns.sysloggh.net:443 | 192.168.68.10:80 | CNAME → netbird | **Yes** | ⚠️ |
|
||||||
| SearXNG | searxng.sysloggh.net:8888 | 192.168.68.7:8888 | LAN IP | No | ✅ |
|
| SearXNG | searxng.sysloggh.net:8888 | 192.168.68.7:8888 | LAN IP | No | ✅ |
|
||||||
| Firecrawl | firecrawl.sysloggh.net:3002 | 192.168.68.7:3002 | LAN IP | No | ✅ |
|
| Firecrawl | firecrawl.sysloggh.net:3002 | 192.168.68.7:3002 | LAN IP | No | ✅ |
|
||||||
|
|
||||||
**Verified 2026-07-02:** NetBird VPS rebooted after a hang; all CNAME'd
|
**Verified 2026-07-24:** NetBird VPS rebooted after a hang; all CNAME'd
|
||||||
services recovered. LAN-IP-direct paths stayed up throughout the outage.
|
services recovered. LAN-IP-direct paths stayed up throughout the outage.
|
||||||
|
Also added `dns.sysloggh.net` route (was missing entirely).
|
||||||
|
See `scripts/netbird-add-domain.sh` for adding new proxy routes.
|
||||||
|
|
||||||
### 5.2 Checks (every 2 min)
|
### 5.2 Checks (every 2 min)
|
||||||
|
|
||||||
@@ -512,8 +524,8 @@ enforced by the `routing-regression.config_url_violations` check in Section
|
|||||||
| LiteLLM API | `http://192.168.68.116:4000` | `https://litellm.sysloggh.net` |
|
| LiteLLM API | `http://192.168.68.116:4000` | `https://litellm.sysloggh.net` |
|
||||||
| LiteLLM (nginx) | `http://192.168.68.116` | — |
|
| LiteLLM (nginx) | `http://192.168.68.116` | — |
|
||||||
| Grafana | `http://192.168.68.116:3001` | — |
|
| Grafana | `http://192.168.68.116:3001` | — |
|
||||||
| Authentik | `https://192.168.68.11` | `https://auth.sysloggh.net` |
|
| Authentik | `https://192.168.68.11:9000` | `https://auth.sysloggh.net` |
|
||||||
| Gitea | `http://192.168.68.110:3000` | `https://git.sysloggh.net` |
|
| Gitea | `http://192.168.68.17:3000` | `https://git.sysloggh.net` |
|
||||||
| Zulip API | `http://192.168.68.19` | `https://chat.sysloggh.net` |
|
| Zulip API | `http://192.168.68.19` | `https://chat.sysloggh.net` |
|
||||||
| SearXNG | `http://192.168.68.7:8888` | — |
|
| SearXNG | `http://192.168.68.7:8888` | — |
|
||||||
| Firecrawl | `http://192.168.68.7:3002` | — |
|
| Firecrawl | `http://192.168.68.7:3002` | — |
|
||||||
@@ -576,25 +588,28 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
|||||||
|
|
||||||
| CT | Name | Node | IP | Role | Agent |
|
| CT | Name | Node | IP | Role | Agent |
|
||||||
|----|------|------|----|------|-------|
|
|----|------|------|----|------|-------|
|
||||||
| 100 | abiba | amdpve | .24 | Pi agent (this host) | ✅ pi |
|
| CT | Name | Node | IP | Role | Agent |
|
||||||
|
|----|------|------|----|------|-------|
|
||||||
|
| 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi |
|
||||||
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
|
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
|
||||||
| 102 | adguard | acerpve | — | DNS | ❌ |
|
| 102 | adguard | **minipve** | **.10** | DNS | ❌ |
|
||||||
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
|
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
|
||||||
| 104 | authentik | minipve | .11 | OIDC | ❌ |
|
| 104 | authentik | minipve | .11 | OIDC | ❌ |
|
||||||
| 105 | kagentz | amdpve | — | Agent Zero | ✅ |
|
| 105 | kagentz | **hwepve** | — | Agent Zero | ✅ |
|
||||||
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
|
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
|
||||||
| 107 | pbs | storepve | — | Backups | ❌ |
|
| 107 | pbs | storepve | — | Backups | ❌ |
|
||||||
| 108 | media | storepve | — | Media | ❌ |
|
| 108 | media | storepve | — | Media | ❌ |
|
||||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||||
| 110 | gitea | minipve | — | Git | ❌ |
|
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
||||||
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
||||||
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
||||||
| 113 | baggy | amdpve | ? | Hermes agent | ✅ |
|
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
||||||
| 114 | mumuni | hwepve | .123 | Hermes agent | ✅ |
|
| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
|
||||||
| 115 | scottdenya | amdpve | — | ? | ❌ |
|
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
||||||
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
||||||
| 117 | zulip | storepve | — | Chat | ❌ |
|
| 117 | zulip | storepve | .19 | Chat | ❌ |
|
||||||
| 118 | jitsi | minipve | — | Video | ❌ |
|
| 118 | jdownloader | storepve | — | JDownloader container | ❌ |
|
||||||
|
| 119 | infisical-vault | minipve | — | Vault | ❌ |
|
||||||
|
|
||||||
## Appendix C: Docker Compose Files Location
|
## Appendix C: Docker Compose Files Location
|
||||||
|
|
||||||
@@ -614,27 +629,28 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
|
|||||||
|
|
||||||
| CT | Name | Node | pct-run |
|
| CT | Name | Node | pct-run |
|
||||||
|-----|------|------|---------|
|
|-----|------|------|---------|
|
||||||
| 100 | abiba | amdpve | `pct-run 100` |
|
| 100 | abiba | hwepve | `pct-run 100` |
|
||||||
| 105 | kagentz | amdpve | `pct-run 105` |
|
| 105 | kagentz | **hwepve** | `pct-run 105` |
|
||||||
| 111 | tdunna | amdpve | `pct-run 111` |
|
| 111 | tdunna | amdpve | `pct-run 111` |
|
||||||
| 112 | tanko | amdpve | `pct-run 112` |
|
| 112 | tanko | amdpve | `pct-run 112` |
|
||||||
| 113 | baggy | amdpve | `pct-run 113` |
|
| 113 | baggy | amdpve | `pct-run 113` |
|
||||||
| 115 | scottdenya | amdpve | `pct-run 115` |
|
| 115 | scottdenya | amdpve | `pct-run 115` |
|
||||||
| 104 | authentik | minipve | `pct-run 104` |
|
| 104 | authentik | minipve | `pct-run 104` |
|
||||||
| 110 | gitea | minipve | `pct-run 110` |
|
| 110 | gitea | minipve | `pct-run 110` |
|
||||||
| 114 | mumuni | hwepve | `pct-run 114` |
|
| 114 | mumuni | **hwepve** | `pct-run 114` |
|
||||||
| 116 | syslog-api | minipve | `pct-run 116` |
|
| 116 | syslog-api | minipve | `pct-run 116` |
|
||||||
| 106 | ra-h-os | storepve | `pct-run 106` |
|
| 106 | ra-h-os | storepve | `pct-run 106` |
|
||||||
| 107 | proxmox-backup | storepve | `pct-run 107` |
|
| 107 | proxmox-backup | storepve | `pct-run 107` |
|
||||||
| 108 | media | storepve | `pct-run 108` |
|
| 108 | media | storepve | `pct-run 108` |
|
||||||
| 117 | zulip | storepve | `pct-run 117` |
|
| 117 | zulip | storepve | `pct-run 117` |
|
||||||
| 102 | adguard | acerpve | `pct-run 102` |
|
| 102 | adguard | **minipve** | `pct-run 102` |
|
||||||
|
|
||||||
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
|
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
|
||||||
```bash
|
```bash
|
||||||
ssh root@192.168.68.8 # RTX 3090
|
ssh root@192.168.68.8 # RTX 3090
|
||||||
ssh root@192.168.68.110 # RTX 5070
|
ssh root@192.168.68.110 # RTX 5070
|
||||||
ssh root@192.168.68.15 # Strix Halo
|
ssh root@192.168.68.15 # Strix Halo
|
||||||
|
ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni)
|
||||||
```
|
```
|
||||||
|
|
||||||
## Section 7: Agent Health Check (consolidated — 2026-07-05)
|
## Section 7: Agent Health Check (consolidated — 2026-07-05)
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: infrastructure-update
|
name: infrastructure-update
|
||||||
description: >
|
description: >
|
||||||
Autonomous system-wide update contract covering all 6 Proxmox nodes,
|
Autonomous system-wide update contract covering all 5 Proxmox nodes,
|
||||||
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
|
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
|
||||||
Docker images, and container stacks in safe waves with health checks
|
Docker images, and container stacks in safe waves with health checks
|
||||||
and automatic rollback on failure.
|
and automatic rollback on failure.
|
||||||
@@ -56,11 +56,10 @@ Before ANY update wave:
|
|||||||
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||||
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||||
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
||||||
| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
|
|
||||||
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
|
||||||
| CT 114 (mumuni, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
| CT 114 (mumuni, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
|
||||||
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
|
||||||
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
|
||||||
|
|
||||||
@@ -68,7 +67,7 @@ Before ANY update wave:
|
|||||||
- All VMs/CTs running: check via Proxmox API
|
- All VMs/CTs running: check via Proxmox API
|
||||||
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
|
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
|
||||||
- LiteLLM MCP tools: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools
|
- LiteLLM MCP tools: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools
|
||||||
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
|
- GPU servers responding: check :8080 on VM 101, VM 103; check strix-moe via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
|
||||||
- Zulip agents connected: check Mumuni/Tanko gateway state
|
- Zulip agents connected: check Mumuni/Tanko gateway state
|
||||||
- Abiba PM2 processes online: `pm2 status`
|
- Abiba PM2 processes online: `pm2 status`
|
||||||
|
|
||||||
@@ -92,17 +91,52 @@ Before ANY update wave:
|
|||||||
- Firecrawl test: `curl :3002/`
|
- Firecrawl test: `curl :3002/`
|
||||||
- SearXNG test: `curl :8888`
|
- SearXNG test: `curl :8888`
|
||||||
|
|
||||||
## Wave 4: Proxmox Kernel Reboot (if needed)
|
## Wave 4: Proxmox Kernel Reboot
|
||||||
|
|
||||||
Only if `[ -f /var/run/reboot-required ]` on any node.
|
Only if `[ -f /var/run/reboot-required ]` on any node.
|
||||||
|
|
||||||
| Target | Action |
|
| Target | Action |
|
||||||
|--------|--------|
|
|--------|--------|
|
||||||
| Affected PVE node | Verify all CTs/VMs migrated or stopped |
|
| Affected PVE node | Verify all CTs/VMs migrated or stopped |
|
||||||
| | `reboot` via PVE API |
|
| | `reboot` via PVE API (or `systemctl reboot -f` if dbus fails) |
|
||||||
| | Wait 120s for node to come back |
|
| | Wait 120s for node to come back |
|
||||||
| | Start any stopped CTs |
|
| | Start any stopped CTs |
|
||||||
|
|
||||||
|
### Post-reboot sweep (known gaps)
|
||||||
|
|
||||||
|
After every node reboot, run these checks:
|
||||||
|
|
||||||
|
1. **CT auto-start sweep** — LXC containers sometimes don't start despite
|
||||||
|
`onboot: 1`. Check every CT on the rebooted node and start any left stopped:
|
||||||
|
```bash
|
||||||
|
pct list | awk '/stopped/{print $1}' | xargs -I{} pct start {}
|
||||||
|
```
|
||||||
|
Known cases: scottdenya (CT 115 on amdpve), authentik (CT 104 on minipve).
|
||||||
|
|
||||||
|
2. **Zulip recovery** — When docker-vm or storepve reboots, the Zulip main
|
||||||
|
container loses its Docker network assignment (SIGKILL during storage
|
||||||
|
outage detaches it from `zulip_default` network). Run:
|
||||||
|
```bash
|
||||||
|
ssh root@192.168.68.19 'docker rm -f zulip-zulip-1 && cd /opt/zulip && docker compose up -d'
|
||||||
|
```
|
||||||
|
The compose restart recreates the container on the correct network.
|
||||||
|
|
||||||
|
3. **docker-vm Docker daemon** — After reboot, Docker can take 3-4 minutes
|
||||||
|
to become `active`. The docker-proxy for Pulse (port 7655) starts early,
|
||||||
|
so Pulse is accessible before `docker ps` reports ready. Wait for Docker
|
||||||
|
before checking other stacks.
|
||||||
|
|
||||||
|
### VPS ↔ docker-vm tunnel
|
||||||
|
|
||||||
|
After any VPS or docker-vm reboot, verify the dedicated WireGuard tunnel:
|
||||||
|
```bash
|
||||||
|
ssh root@72.61.0.17 'wg show wg1' | grep "latest handshake"
|
||||||
|
# If no handshake in >60s:
|
||||||
|
ssh root@72.61.0.17 'wg-quick up wg1'
|
||||||
|
```
|
||||||
|
The tunnel uses PersistentKeepalive=25 and is systemd-enabled, but should
|
||||||
|
be verified after a reboot.
|
||||||
|
|
||||||
## Rollback Protocol
|
## Rollback Protocol
|
||||||
|
|
||||||
If ANY verification fails:
|
If ANY verification fails:
|
||||||
@@ -125,7 +159,7 @@ Before Wave 1, snapshot these files:
|
|||||||
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
||||||
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
||||||
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
||||||
/etc/systemd/system/ornith-server.service (amdpve .15)
|
/etc/systemd/system/ornith-server.service (amdpve .15 — strix-moe)
|
||||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||||
# Hermes agent configs (key enforcement — 2026-07-10)
|
# Hermes agent configs (key enforcement — 2026-07-10)
|
||||||
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
|
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
|
||||||
@@ -184,7 +218,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
|
|||||||
|
|
||||||
## Success Criteria
|
## Success Criteria
|
||||||
|
|
||||||
- [ ] All 6 PVE nodes updated, no reboot-loop
|
- [ ] All 5 PVE nodes updated, no reboot-loop
|
||||||
- [ ] All VMs/CTs running post-update
|
- [ ] All VMs/CTs running post-update
|
||||||
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
|
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
|
||||||
- [ ] LiteLLM inference passing (syslog-auto test)
|
- [ ] LiteLLM inference passing (syslog-auto test)
|
||||||
@@ -200,7 +234,7 @@ After completion, send Zulip DM:
|
|||||||
```
|
```
|
||||||
📋 Infrastructure Update — YYYY-MM-DD
|
📋 Infrastructure Update — YYYY-MM-DD
|
||||||
|
|
||||||
Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers
|
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers
|
||||||
Security fixes: N CVEs patched
|
Security fixes: N CVEs patched
|
||||||
Downtime: <service> <duration>
|
Downtime: <service> <duration>
|
||||||
Failures: none / <details>
|
Failures: none / <details>
|
||||||
|
|||||||
@@ -42,7 +42,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
||||||
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
||||||
- All GPUs at parallel 2 (was parallel 1)
|
- All GPUs at parallel 2 (was parallel 1)
|
||||||
- NVIDIA context reduced 256K→128K to free VRAM (SUPERSEDED 2026-07-16: all GPUs back to 256K — see litellm-self-heal)
|
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
|
||||||
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
|
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
|
||||||
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
|
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
|
||||||
|
|
||||||
@@ -72,9 +72,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||||
|------|-----|----------|---------------|--------|---------|----------|
|
|------|-----|----------|---------------|--------|---------|----------|
|
||||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **256K** | 2 |
|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
|
||||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **256K** | 2 |
|
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
|
||||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 256K | 2 |
|
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||||
|
|
||||||
## Model Fallback Chains (LiteLLM)
|
## Model Fallback Chains (LiteLLM)
|
||||||
|
|
||||||
|
|||||||
@@ -57,9 +57,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||||
|------|-----|----------|---------------|--------|---------|----------|
|
|------|-----|----------|---------------|--------|---------|----------|
|
||||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 262144 --parallel 2 --ngl 99`) | **256K** | 2 |
|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
|
||||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 262144 --parallel 2`, IQ4_NL + MTP draft) | **256K** | 2 |
|
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
|
||||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 256K | 2 |
|
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||||
|
|
||||||
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
||||||
|
|
||||||
@@ -101,7 +101,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
## Script Operations (synced 2026-07-16)
|
## Script Operations (synced 2026-07-16)
|
||||||
|
|
||||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 256K context steady-state is ~96% on RTX 3090, not a fault).
|
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
|
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
|
||||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||||
|
|
||||||
|
|||||||
@@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6,
|
|||||||
## Why This Matters
|
## Why This Matters
|
||||||
|
|
||||||
Without enforced delegation, the manager consumes the full iteration budget
|
Without enforced delegation, the manager consumes the full iteration budget
|
||||||
(60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs —
|
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
|
||||||
leaving no capacity for actual coordination. The result: context overflow
|
leaving no capacity for actual coordination. The result: context overflow
|
||||||
(59K tokens in system prompt), iteration exhaustion, and degraded response
|
(59K tokens in system prompt), iteration exhaustion, and degraded response
|
||||||
quality. This contract exists because I blew through my budget checking
|
quality. This contract exists because I blew through my budget checking
|
||||||
@@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own.
|
|||||||
|
|
||||||
**This is a hard rule, not a recommendation.** Violating it produces the exact
|
**This is a hard rule, not a recommendation.** Violating it produces the exact
|
||||||
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
|
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
|
||||||
"6 nodes present" in the raw data but "5/6 online" in the report — even though
|
"5 nodes present" in the raw data but "5/5 online" in the report — even though
|
||||||
one of those nodes was unreachable. The report lied because it used data the
|
one of those nodes was unreachable. The report lied because it used data the
|
||||||
raw data never provided.
|
raw data never provided.
|
||||||
|
|
||||||
@@ -137,7 +137,7 @@ delegate_task(
|
|||||||
```
|
```
|
||||||
delegate_task(
|
delegate_task(
|
||||||
tasks=[
|
tasks=[
|
||||||
{"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
|
||||||
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
|
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
|
||||||
]
|
]
|
||||||
)
|
)
|
||||||
@@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel:
|
|||||||
{
|
{
|
||||||
"lane_id": "devops-check",
|
"lane_id": "devops-check",
|
||||||
"worker": "syslog-devops",
|
"worker": "syslog-devops",
|
||||||
"goal": "Check all 6 Proxmox nodes",
|
"goal": "Check all 5 Proxmox nodes",
|
||||||
"status": "dispatched|completed|failed",
|
"status": "dispatched|completed|failed",
|
||||||
"output_file": "/tmp/node-report.md"
|
"output_file": "/tmp/node-report.md"
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -20,7 +20,7 @@ agent: abiba
|
|||||||
└───────┬──────────────┬──────────────┬───────────────────────┘
|
└───────┬──────────────┬──────────────┬───────────────────────┘
|
||||||
│ │ │
|
│ │ │
|
||||||
▼ ▼ ▼
|
▼ ▼ ▼
|
||||||
pve-exporter docker-stats (scrapes 6x node_exporter)
|
pve-exporter docker-stats (scrapes 5x node_exporter)
|
||||||
:9221 :9324
|
:9221 :9324
|
||||||
│ │
|
│ │
|
||||||
▼ ▼
|
▼ ▼
|
||||||
@@ -71,7 +71,7 @@ agent: abiba
|
|||||||
| File | Host | Purpose |
|
| File | Host | Purpose |
|
||||||
|------|------|---------|
|
|------|------|---------|
|
||||||
| `/opt/monitoring/docker-compose.yml` | .116 | monitoring stack (prometheus, grafana, pve-exporter, docker-stats) |
|
| `/opt/monitoring/docker-compose.yml` | .116 | monitoring stack (prometheus, grafana, pve-exporter, docker-stats) |
|
||||||
| `/opt/monitoring/prometheus.yml` | .116 | 6 scrape jobs (3 GPU, pve, node x6, docker-stats) |
|
| `/opt/monitoring/prometheus.yml` | .116 | 6 scrape jobs (3 GPU, pve, node x5, docker-stats) |
|
||||||
| `/opt/monitoring/pve.yml` | .116 | PVE API credentials (chmod 644, contains token) |
|
| `/opt/monitoring/pve.yml` | .116 | PVE API credentials (chmod 644, contains token) |
|
||||||
| `/opt/monitoring/docker-stats-exporter.py` | .116 | custom Docker metrics exporter |
|
| `/opt/monitoring/docker-stats-exporter.py` | .116 | custom Docker metrics exporter |
|
||||||
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
|
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
|
||||||
@@ -87,7 +87,7 @@ agent: abiba
|
|||||||
| storepve | 192.168.68.6 | PVE |
|
| storepve | 192.168.68.6 | PVE |
|
||||||
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
||||||
| minipve | 192.168.68.12 | PVE |
|
| minipve | 192.168.68.12 | PVE |
|
||||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (ornith) |
|
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
|
||||||
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts Mumuni (lxc/114) migrated from minipve 2026-07-20 |
|
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts Mumuni (lxc/114) migrated from minipve 2026-07-20 |
|
||||||
|
|
||||||
## Operations
|
## Operations
|
||||||
|
|||||||
@@ -0,0 +1,65 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# Netbird Reverse Proxy — Add a new domain route
|
||||||
|
#
|
||||||
|
# Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]
|
||||||
|
#
|
||||||
|
# Example:
|
||||||
|
# netbird-add-domain.sh dns.sysloggh.net 192.168.68.10 80
|
||||||
|
#
|
||||||
|
# This script adds a domain to the Netbird proxy by inserting records
|
||||||
|
# directly into the management server's SQLite database, then restarting
|
||||||
|
# the proxy stack.
|
||||||
|
#
|
||||||
|
# Prerequisites: SSH root access to 72.61.0.17
|
||||||
|
# sqlite3 available on VPS
|
||||||
|
#
|
||||||
|
# Requires: The domain must already have a DNS CNAME to netbird.sysloggh.net
|
||||||
|
# pointing to 72.61.0.17.
|
||||||
|
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
DOMAIN="${1:?Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]}"
|
||||||
|
BACKEND_IP="${2:?Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]}"
|
||||||
|
PORT="${3:-80}"
|
||||||
|
PROTOCOL="${4:-http}"
|
||||||
|
|
||||||
|
VPS="root@72.61.0.17"
|
||||||
|
DB_VOLUME="/var/lib/docker/volumes/root_netbird_data/_data"
|
||||||
|
DB="$DB_VOLUME/store.db"
|
||||||
|
|
||||||
|
echo "=== Adding Netbird proxy route ==="
|
||||||
|
echo "Domain: $DOMAIN"
|
||||||
|
echo "Backend: $BACKEND_IP:$PORT ($PROTOCOL)"
|
||||||
|
echo ""
|
||||||
|
|
||||||
|
ssh "$VPS" bash << REMOTESCRIPT
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
# Generate unique ID using timestamp hash (Netbird format)
|
||||||
|
ID_SUFFIX=\$(date +%s | md5sum | head -c 16)
|
||||||
|
SVC_ID="d9\${ID_SUFFIX}ptsnc73\$(date +%s | md5sum | head -c 10)"
|
||||||
|
TGT_ID=\$(sqlite3 "$DB" "SELECT COALESCE(MAX(id), 100) + 1 FROM targets;")
|
||||||
|
ACCOUNT_ID="d88av3aptsnc73clmogg"
|
||||||
|
ZONE_ID="d8adqjaptsnc73fro5g0"
|
||||||
|
|
||||||
|
echo "Service ID: \$SVC_ID"
|
||||||
|
echo "Target ID: \$TGT_ID"
|
||||||
|
|
||||||
|
# Insert service
|
||||||
|
sqlite3 "$DB" "INSERT INTO services (id, account_id, name, domain, proxy_cluster, enabled, terminated, pass_host_header, rewrite_redirects, mode, source, port_auto_assigned, private) VALUES (\"\$SVC_ID\", \"\$ACCOUNT_ID\", \"$DOMAIN\", \"$DOMAIN\", \"netbird.sysloggh.net\", 1, 0, 1, 0, \"http\", \"permanent\", 0, 0);"
|
||||||
|
echo "Service: OK"
|
||||||
|
|
||||||
|
# Insert target
|
||||||
|
sqlite3 "$DB" "INSERT INTO targets (id, account_id, service_id, host, port, protocol, target_id, target_type, enabled, skip_tls_verify, request_timeout, session_idle_timeout, agent_network, disable_access_log) VALUES (\$TGT_ID, \"\$ACCOUNT_ID\", \"\$SVC_ID\", \"$BACKEND_IP\", $PORT, \"$PROTOCOL\", \"\$ZONE_ID\", \"subnet\", 1, 0, 0, 0, 0, 0);"
|
||||||
|
echo "Target: OK"
|
||||||
|
|
||||||
|
# Verify
|
||||||
|
sqlite3 -column "$DB" "SELECT s.name, t.host, t.port, t.protocol FROM services s JOIN targets t ON s.id=t.service_id WHERE s.name=\"$DOMAIN\";"
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
echo "Restarting proxy stack..."
|
||||||
|
cd /root && docker compose restart netbird-server 2>/dev/null
|
||||||
|
sleep 15
|
||||||
|
docker compose restart proxy 2>/dev/null
|
||||||
|
echo "Done. Verify with: curl -sI https://$DOMAIN"
|
||||||
|
REMOTESCRIPT
|
||||||
+1
-3
@@ -17,11 +17,10 @@ declare -A CT_NODES=(
|
|||||||
[112]=amdpve # tanko
|
[112]=amdpve # tanko
|
||||||
[113]=amdpve # baggy
|
[113]=amdpve # baggy
|
||||||
[115]=amdpve # scottdenya
|
[115]=amdpve # scottdenya
|
||||||
# hwepve (192.168.68.4) — Huawei Matebook 16
|
|
||||||
[114]=hwepve # mumuni (migrated from minipve 2026-07-20)
|
|
||||||
# minipve (192.168.68.12)
|
# minipve (192.168.68.12)
|
||||||
[104]=minipve # authentik
|
[104]=minipve # authentik
|
||||||
[110]=minipve # gitea
|
[110]=minipve # gitea
|
||||||
|
[114]=minipve # mumuni
|
||||||
[116]=minipve # syslog-api
|
[116]=minipve # syslog-api
|
||||||
# storepve (192.168.68.6)
|
# storepve (192.168.68.6)
|
||||||
[106]=storepve # ra-h-os
|
[106]=storepve # ra-h-os
|
||||||
@@ -49,7 +48,6 @@ declare -A NODE_IPS=(
|
|||||||
[storepve]=192.168.68.6
|
[storepve]=192.168.68.6
|
||||||
[acerpve]=192.168.68.9
|
[acerpve]=192.168.68.9
|
||||||
[ocupve]=192.168.68.5
|
[ocupve]=192.168.68.5
|
||||||
[hwepve]=192.168.68.4
|
|
||||||
)
|
)
|
||||||
|
|
||||||
resolve_node() {
|
resolve_node() {
|
||||||
|
|||||||
@@ -46,13 +46,12 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
|
|||||||
|
|
||||||
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
|
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
|
||||||
|
|
||||||
**Proxmox Cluster "Tabiri" (6 nodes):**
|
**Proxmox Cluster "Tabiri" (5 nodes):**
|
||||||
- amdpve (192.168.68.15): abiba, kagentz, tanko, tdunna, baggy, scottdenya
|
- amdpve (192.168.68.15): abiba, kagentz, tanko, tdunna, baggy, scottdenya
|
||||||
- minipve (192.168.68.12): authentik, gitea, syslog-api, jitsi
|
- minipve (192.168.68.12): authentik, gitea, mumuni, syslog-api, jitsi
|
||||||
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, zulip
|
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, zulip
|
||||||
- acerpve (192.168.68.9): llm-gpu, adguard
|
- acerpve (192.168.68.9): llm-gpu, adguard
|
||||||
- ocupve (192.168.68.5): ocu-llm
|
- ocupve (192.168.68.5): ocu-llm
|
||||||
- hwepve (192.168.68.4): mumuni (migrated from minipve 2026-07-20, Huawei Matebook 16, 12C/15GB)
|
|
||||||
|
|
||||||
**CT IDs (verified 2026-07-04 against PVE API):**
|
**CT IDs (verified 2026-07-04 against PVE API):**
|
||||||
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
|
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
|
||||||
|
|||||||
@@ -90,12 +90,16 @@ echo "── 3. Cross-contract consistency ──"
|
|||||||
|
|
||||||
# Check that contracts referencing each other have correct names
|
# Check that contracts referencing each other have correct names
|
||||||
if [ -f "infrastructure-control.prose.md" ]; then
|
if [ -f "infrastructure-control.prose.md" ]; then
|
||||||
# Any contract that claims to check "all 6 PVE nodes" should name them
|
# Any contract that claims to check "all 5 PVE nodes" should name them
|
||||||
for f in *.prose.md; do
|
for f in *.prose.md; do
|
||||||
[ -f "$f" ] || continue
|
[ -f "$f" ] || continue
|
||||||
if grep -qE "\b5-node\b|\b5 node\b|\b5 Proxmox\b|all 5 PVE" "$f" 2>/dev/null; then
|
if grep -q "5-node\|5 node\|5 Proxmox\|all.*PVE.*node" "$f" 2>/dev/null; then
|
||||||
echo " ⚠️ $f: still references 5-node cluster (migrated to 6 nodes 2026-07-20)"
|
for node in amdpve minipve storepve acerpve ocupve; do
|
||||||
|
grep -q "$node" "$f" || {
|
||||||
|
echo " ⚠️ $f: references 5 nodes but '$node' not mentioned"
|
||||||
WARNINGS=$((WARNINGS + 1))
|
WARNINGS=$((WARNINGS + 1))
|
||||||
|
}
|
||||||
|
done
|
||||||
fi
|
fi
|
||||||
done
|
done
|
||||||
fi
|
fi
|
||||||
|
|||||||
@@ -199,7 +199,7 @@ Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error`
|
|||||||
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
|
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
|
||||||
```
|
```
|
||||||
|
|
||||||
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway: for Mumuni (PM2-managed), use `pm2 restart mumuni-zulip`; for Tanko (systemd/non-PM2), use `hermes gateway restart`. Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway via PM2 (`pm2 restart mumuni-zulip`). Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
|
||||||
|
|
||||||
**B3: Heartbeat Verification**
|
**B3: Heartbeat Verification**
|
||||||
|
|
||||||
@@ -222,7 +222,7 @@ ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/
|
|||||||
|
|
||||||
| Condition | Action |
|
| Condition | Action |
|
||||||
|-----------|--------|
|
|-----------|--------|
|
||||||
| `zulip.state != "connected"` | Mumuni (PM2): `ssh root@<CT> "pm2 restart mumuni-zulip"`<br>Tanko (systemd): `ssh root@<CT> "hermes gateway restart"` |
|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
|
||||||
| No heartbeat in 10min | Same as above |
|
| No heartbeat in 10min | Same as above |
|
||||||
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
|
||||||
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
| Response empty/short | Check A2A endpoint / LiteLLM model |
|
||||||
|
|||||||
Reference in New Issue
Block a user