Compare commits

..
Author SHA1 Message Date
root 8a6c8d80de fix: contract accuracy updates post fleet-wide reboot
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
infrastructure-control.prose.md:
- Add hwepve as 6th Proxmox node
- Fix Gitea IP: .17 (was .110)
- Fix AdGuard IP: .10 on minipve (was .102 on acerpve)
- Fix Abiba placement: hwepve (was amdpve)
- Fix Mumuni placement: hwepve (was minipve)
- Fix Authentik port: add :9000
- Add CT 118 (jdownloader), CT 119 (infisical-vault)
- Add last-verified date (2026-07-24)

infrastructure-update.prose.md:
- Add post-reboot CT sweep procedure
- Add Zulip Docker network recovery steps
- Add WireGuard tunnel verification

scripts/netbird-add-domain.sh:
- New script to register domains in Netbird proxy store.db
2026-07-24 19:56:26 +00:00
6 changed files with 90 additions and 341 deletions
+8 -11
View File
@@ -14,10 +14,6 @@ description: >
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek). For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
UPDATED 2026-07-27: gpu-dense swapped to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled
(UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking
tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning.
VRAM ~22.4/24.6GB (91%).
agent: abiba agent: abiba
triggers: triggers:
- on model add/remove - on model add/remove
@@ -90,7 +86,7 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To | | Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------| |-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q4_K_XL | Whatever runs on RTX 3090 | | `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | | `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
@@ -100,7 +96,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy | | qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy | | gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s | | qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
@@ -110,7 +106,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | | qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** | | gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
@@ -121,7 +117,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload | | strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
| SmartCode-Fable-5-27B-UD-Q3_K_XL | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) | | qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s | | gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
### Stable Aliases (for agent configs — never change) ### Stable Aliases (for agent configs — never change)
@@ -251,9 +247,10 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB. - **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 (2026-07-27)**: Swapped to SmartCode-Fable-5-27B-UD-Q3_K_XL (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`. - **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`. - **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-27)**: SmartCode-Fable-5-27B 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
@@ -268,7 +265,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q3_K_XL | **TBD** | — | — | **128K** | | RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** | | RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** | | Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
+1 -1
View File
@@ -253,7 +253,7 @@ EnvironmentFile=/etc/environment # sources LITELLM_API_KEY
Run the consolidated health check: Run the consolidated health check:
```bash ```bash
python3 /root/scripts/agent-health-check.py # v2: now checks all 5 agents including Koby/Koonimo SSH, CT liveness, config YAML, wrapper integrity, vault non-emptiness python3 /root/scripts/agent-health-check.py
``` ```
This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes), This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes),
verifies gateway liveness, confirms Zulip streaming (`edit_message` present), verifies gateway liveness, confirms Zulip streaming (`edit_message` present),
+20 -28
View File
@@ -13,10 +13,10 @@ description: >
against the live system. Policy fields are authoritative. See the against the live system. Policy fields are authoritative. See the
`verify-before-mutate` skill. `verify-before-mutate` skill.
**Last verified:** 2026-07-27gpu-dense swapped to SmartCode-Fable-5 **Last verified:** 2026-07-24corrected Gitea IP (.17 not .110),
(Qwen3.6-27B distilled, ~50% fewer thinking tokens, improved coding reasoning). AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve),
Model pricing reduced 100x across all models ($0.15/$0.60 per 1M tokens). Abiba placement (hwepve not amdpve), added hwepve as 6th node,
All LiteLLM models set to 128K max_model_tokens. See data/learnings.md. added dns.sysloggh.net route. See data/learnings.md.
--- ---
# Infrastructure Control Pattern # Infrastructure Control Pattern
@@ -48,7 +48,7 @@ description: >
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼
minipve amdpve storepve acerpve ocupve hwepve minipve amdpve storepve acerpve ocupve hwepve
(.12) (.15) (.6) (.9) (.5) (.4) (.12) (.15) (.6) (.9) (.5) (.130)
┌─────────────────────────────────────────────┐ ┌─────────────────────────────────────────────┐
@@ -105,16 +105,16 @@ description: >
| Node | IP | CPU | RAM | VMs/CTs | Role | | Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------| |------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging | | minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute | | amdpve | .15 | 32C | 62GB | kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat | | storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs | | acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | | ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
| hwepve | .4 | ? | ? | abiba, (mumuni CT 114 stopped) | Agents (new node) | | hwepve | .130 | ? | ? | abiba, (mumuni CT 114 stopped) | Agents (new node, offline issues) |
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on > **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve and now runs Mumuni > minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni
> Zulip gateway internally. CT 114 (old mumuni container) destroyed 2026-07-26. > (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a
> Mumuni also has a second instance on minipve at .123 — distinguish by CT ID. > second instance on minipve at .123 — distinguish by CT ID, not hostname.
### Checks (every 5 min) ### Checks (every 5 min)
@@ -604,6 +604,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ | | 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ | | 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ | | 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ | | 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ | | 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
| 117 | zulip | storepve | .19 | Chat | ❌ | | 117 | zulip | storepve | .19 | Chat | ❌ |
@@ -629,14 +630,14 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| CT | Name | Node | pct-run | | CT | Name | Node | pct-run |
|-----|------|------|---------| |-----|------|------|---------|
| 100 | abiba | hwepve | `pct-run 100` | | 100 | abiba | hwepve | `pct-run 100` |
| 105 | kagentz | **hwepve** | `pct-run 105` | | 105 | kagentz | amdpve | `pct-run 105` |
| 111 | tdunna | amdpve | `pct-run 111` | | 111 | tdunna | amdpve | `pct-run 111` |
| 112 | tanko | amdpve | `pct-run 112` | | 112 | tanko | amdpve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` | | 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` | | 115 | scottdenya | amdpve | `pct-run 115` |
| 104 | authentik | minipve | `pct-run 104` | | 104 | authentik | minipve | `pct-run 104` |
| 110 | gitea | minipve | `pct-run 110` | | 110 | gitea | minipve | `pct-run 110` |
| 114 | mumuni | minipve | `pct-run 114` |
| 116 | syslog-api | minipve | `pct-run 116` | | 116 | syslog-api | minipve | `pct-run 116` |
| 106 | ra-h-os | storepve | `pct-run 106` | | 106 | ra-h-os | storepve | `pct-run 106` |
| 107 | proxmox-backup | storepve | `pct-run 107` | | 107 | proxmox-backup | storepve | `pct-run 107` |
@@ -649,35 +650,26 @@ GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use S
ssh root@192.168.68.8 # RTX 3090 ssh root@192.168.68.8 # RTX 3090
ssh root@192.168.68.110 # RTX 5070 ssh root@192.168.68.110 # RTX 5070
ssh root@192.168.68.15 # Strix Halo ssh root@192.168.68.15 # Strix Halo
ssh root@192.168.68.4 # hwepve (abiba, kagentz) — Mumuni runs inside CT 100
``` ```
## Section 7: Agent Health Check v2 (2026-07-26) ## Section 7: Agent Health Check (consolidated — 2026-07-05)
Single non-disruptive check running every 10 minutes via cron Replaces 7 scattered Zulip health scripts with a single non-disruptive check
(`/root/scripts/agent-health-check.py`). v2 fixes critical gaps: running every 10 minutes via cron (`/root/scripts/agent-health-check.py`).
- Koby (.129) and Koonimo (.114) now have SSH hosts — no longer skipped
- Agent-specific vault keys (`{NAME}_LITELLM_API_KEY`) not shared master key
- CT liveness check via `pct status` on PVE nodes
- Config YAML integrity check
- Wrapper/CLI integrity check
- Vault secret non-emptiness check
The script is **read-only** — it never restarts, kills, or modifies anything. The script is **read-only** — it never restarts, kills, or modifies anything.
Disruptive cron-based gateways restarts (like Mumuni's zulip-watchdog.sh,
which was kill+nohup outside systemd) are banned by policy.
### Checks Performed ### Checks Performed
| Check | Frequency | What It Detects | | Check | Frequency | What It Detects |
|-------|-----------|-----------------| |-------|-----------|-----------------|
| LiteLLM key validation | 10 min | All 5 agent-specific keys authenticate (not shared master key) | | LiteLLM key validation | 10 min | All 4 agent keys authenticate and return models |
| GPU port conflict | 10 min | Ghost processes squatting port 8080 (ss vs systemd MainPID) | | GPU port conflict | 10 min | Ghost processes squatting port 8080 (ss vs systemd MainPID) |
| Gateway liveness | 10 min | Gateway process running, state file readable (all 5 agents) | | Gateway liveness | 10 min | Gateway process running, state file readable |
| Zulip streaming | 10 min | `edit_message` present in adapter (streaming supported) | | Zulip streaming | 10 min | `edit_message` present in adapter (streaming supported) |
| Recent errors | 10 min | Error count in journald for last 10 min | | Recent errors | 10 min | Error count in journald for last 10 min |
| CT liveness | 10 min | `pct status` on PVE nodes — catches stopped CTs |
| Config YAML integrity | 10 min | Python `yaml.safe_load()` — catches syntax errors |
| Wrapper/CLI integrity | 10 min | hermes wrapper exists, infisical path correct, hermes-real reachable |
| Vault secrets | 10 min | Agent-specific vault keys are non-empty and start with `sk-` |
### Disabled Scripts ### Disabled Scripts
+3 -1
View File
@@ -117,7 +117,9 @@ After every node reboot, run these checks:
container loses its Docker network assignment (SIGKILL during storage container loses its Docker network assignment (SIGKILL during storage
outage detaches it from `zulip_default` network). Run: outage detaches it from `zulip_default` network). Run:
```bash ```bash
ssh root@192.168.68.19 'docker rm -f zulip-zulip-1 && cd /opt/zulip && docker compose up -d' ssh root@192.168.68.19
docker rm -f zulip-zulip-1
cd /opt/zulip && docker compose up -d
``` ```
The compose restart recreates the container on the correct network. The compose restart recreates the container on the correct network.
+1 -1
View File
@@ -102,7 +102,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`). - **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault). - **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.123), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. - **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM. - **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
## Maintains ## Maintains
+57 -299
View File
@@ -1,10 +1,10 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
""" """
/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v2 /root/scripts/agent-health-check.py — Consolidated Agent Health Verification
Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming, Single non-disruptive health check replacing 7 scattered scripts.
gateway liveness, gateway log health, CT liveness, config YAML integrity, Verifies: LiteLLM keys, GPU port conflicts, agent Zulip streaming,
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything. gateway liveness, and gateway log health. NEVER restarts anything.
Usage: Usage:
python3 /root/scripts/agent-health-check.py # Full check python3 /root/scripts/agent-health-check.py # Full check
@@ -12,40 +12,60 @@ Usage:
python3 /root/scripts/agent-health-check.py --quiet # Only output on failure python3 /root/scripts/agent-health-check.py --quiet # Only output on failure
Cron: */10 * * * * python3 /root/scripts/agent-health-check.py --quiet Cron: */10 * * * * python3 /root/scripts/agent-health-check.py --quiet
Changelog:
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
Fleet roster: tanko (.122), mumuni (.123), koby (.129), koonimo (.114),
abiba (.24).
""" """
import subprocess, json, sys, os, time import subprocess, json, sys, os, time
from datetime import datetime from datetime import datetime
LITELLM = "http://192.168.68.116:80" LITELLM = "http://192.168.68.116:80"
INFISICAL_PROJECT = "322fceab-39da-4854-a55a-568e76c0f13f"
INFISICAL_ENV = "prod"
# PVE node IPs for CT liveness checks def _get_agent_key(agent_name):
PVE_NODES = { """Retrieve agent key from Infisical vault."""
"hwepve": "192.168.68.4", try:
"amdpve": "192.168.68.15", result = subprocess.run(
"minipve": "192.168.68.12", ["infisical", "secrets", "get", "LITELLM_API_KEY",
"storepve": "192.168.68.6", "--project=agents", "--env=production", "--plain"],
"acerpve": "192.168.68.9", capture_output=True, text=True, timeout=10
"ocupve": "192.168.68.5", )
} if result.returncode == 0:
return result.stdout.strip()
except Exception:
pass
# Agent definitions: ct, host, user, pve_node, vault_key_name # Fallback: try exporting all secrets
try:
result = subprocess.run(
["infisical", "export", "--project=agents", "--env=production",
"--format=dotenv"],
capture_output=True, text=True, timeout=10
)
if result.returncode == 0:
for line in result.stdout.splitlines():
if line.startswith(f"LITELLM_API_KEY_{agent_name.upper()}") or \
(line.startswith("LITELLM_API_KEY=") and agent_name == os.uname().nodename):
return line.split("=", 1)[1].strip().strip('"').strip("'")
except Exception:
pass
return None
# Agent keys are pulled from Infisical vault at runtime.
# The 'key' field is populated dynamically below.
AGENTS = { AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"}, "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome"},
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key "mumuni": {"ct": 114, "host": "192.168.68.123", "user": "root"},
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"}, "koby": {"ct": 111, "host": None, "user": None},
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"}, "koonimo": {"ct": 113, "host": None, "user": None},
} }
# Inject keys from vault
for agent_name in AGENTS:
key = _get_agent_key(agent_name)
if key:
AGENTS[agent_name]["key"] = key
else:
AGENTS[agent_name]["key"] = None
GPU_HOSTS = { GPU_HOSTS = {
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"}, "gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"}, "gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"},
@@ -54,11 +74,6 @@ GPU_HOSTS = {
FAIL = [] FAIL = []
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
# ── Helpers ──────────────────────────────────────────────────────────
def ssh(host, cmd, user="root"): def ssh(host, cmd, user="root"):
"""Execute a command on a remote host, return stdout or None.""" """Execute a command on a remote host, return stdout or None."""
try: try:
@@ -97,81 +112,15 @@ def http_json(url, headers=None, timeout=5):
except: except:
return None return None
def run_infisical(args, quiet=True):
"""Run infisical CLI with env-based auth, return stdout or None."""
env = os.environ.copy()
env["INFISICAL_API_URL"] = INFISICAL_API_URL
if INFISICAL_TOKEN:
env["INFISICAL_TOKEN"] = INFISICAL_TOKEN
try:
result = subprocess.run(
["/usr/bin/infisical"] + args,
capture_output=True, text=True, timeout=15, env=env
)
return result.stdout.strip() if result.returncode == 0 else None
except:
return None
# ── KEY LOOKUP FIX ───────────────────────────────────────────────────
def _get_agent_key(agent_name, vault_key_name):
"""Retrieve agent-specific key from Infisical vault.
Uses {NAME}_LITELLM_API_KEY format (e.g., TANKO_LITELLM_API_KEY,
KOONIMO_LITELLM_API_KEY) which matches actual vault key names.
"""
if not vault_key_name:
return None
# Primary: get the agent-specific key by name
key = run_infisical([
"secrets", "get", vault_key_name,
"--projectId=" + INFISICAL_PROJECT,
"--env=" + INFISICAL_ENV,
"--plain",
])
if key and key.startswith("sk-"):
return key
# Fallback: export all and search for the key name
try:
export = run_infisical([
"export",
"--projectId=" + INFISICAL_PROJECT,
"--env=" + INFISICAL_ENV,
"--format=dotenv",
])
if export:
for line in export.splitlines():
if line.startswith(vault_key_name + "="):
value = line.split("=", 1)[1].strip().strip('"').strip("'")
if value.startswith("sk-"):
return value
except:
pass
return None
# Inject keys from vault for each agent
for agent_name in AGENTS:
info = AGENTS[agent_name]
key = _get_agent_key(agent_name, info.get("vault_key"))
AGENTS[agent_name]["key"] = key
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 1: LiteLLM Key Validation (agent-specific keys) # CHECK 1: LiteLLM Key Validation
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_keys(): def check_keys():
for name, agent in AGENTS.items(): for name, agent in AGENTS.items():
key = agent.get("key")
if not key:
print(f"{name}: NO KEY FOUND (vault empty or unreachable)")
FAIL.append(f"key:{name}:no-key")
continue
data = http_json(f"{LITELLM}/v1/models", data = http_json(f"{LITELLM}/v1/models",
headers={"Authorization": f"Bearer {key}"}) headers={"Authorization": f"Bearer {agent['key']}"})
if data and data.get("data"): if data and data.get("data"):
model = data["data"][0].get("id", "?") model = data["data"][0].get("id", "?")
print(f"{name}: key valid → {model}") print(f"{name}: key valid → {model}")
@@ -181,7 +130,7 @@ def check_keys():
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 2: GPU Port Conflict Detection (unchanged) # CHECK 2: GPU Port Conflict Detection
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_gpu_ports(): def check_gpu_ports():
@@ -209,6 +158,7 @@ def check_gpu_ports():
else: else:
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}") print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
else: else:
# Verify health endpoint
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health") health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health")
if health and '"status":"ok"' in health: if health and '"status":"ok"' in health:
print(f"{label}: healthy (pid={port_owner})") print(f"{label}: healthy (pid={port_owner})")
@@ -221,7 +171,7 @@ def check_gpu_ports():
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents) # CHECK 3: Agent Gateway Liveness + Streaming
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_agents(): def check_agents():
@@ -234,11 +184,8 @@ def check_agents():
print(f"{name} (CT {ct}): cannot SSH — skip liveness check") print(f"{name} (CT {ct}): cannot SSH — skip liveness check")
continue continue
# Gateway process # Gateway process (exclude the infisical bash wrapper that contains the same string)
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user) pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
# Try alternate binary name
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid: if not pid:
print(f"{name}: GATEWAY NOT RUNNING") print(f"{name}: GATEWAY NOT RUNNING")
FAIL.append(f"gateway-down:{name}") FAIL.append(f"gateway-down:{name}")
@@ -256,7 +203,7 @@ def check_agents():
else: else:
gw_state, zulip = "no-state-file", "?" gw_state, zulip = "no-state-file", "?"
# Zulip streaming check # Zulip streaming: does adapter have edit_message?
adapter_paths = [ adapter_paths = [
"~/.hermes/plugins/zulip-platform/adapter.py", "~/.hermes/plugins/zulip-platform/adapter.py",
"~/.hermes/plugins/platforms/zulip/adapter.py", "~/.hermes/plugins/platforms/zulip/adapter.py",
@@ -273,183 +220,13 @@ def check_agents():
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null " r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0", r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
user=user) user=user)
recent_errors = (recent_errors or "0").strip().split("\n")[-1] recent_errors = (recent_errors or "0").strip().split("\n")[-1] # take last line
print(f" {'' if gw_state == 'running' and zulip == 'connected' else '⚠️'} " print(f" {'' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} " f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
f"errors_10m={recent_errors.strip() or '0'} pid={pid}") f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 4: CT Liveness (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_ct_liveness():
"""Check that all agent CTs are running on their PVE nodes."""
for name, agent in AGENTS.items():
ct = agent["ct"]
pve_node = agent.get("pve")
if not pve_node:
print(f"{name} (CT {ct}): no PVE node mapped — skip")
continue
pve_ip = PVE_NODES.get(pve_node)
if not pve_ip:
print(f"{name}: unknown PVE node '{pve_node}' — skip")
continue
status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root")
if not status:
print(f"{name} (CT {ct} on {pve_node}): PVE UNREACHABLE")
FAIL.append(f"ct-unreachable:{name}:{pve_ip}")
elif "running" in status:
print(f"{name} (CT {ct} on {pve_node}): running")
elif "stopped" in status:
print(f"{name} (CT {ct} on {pve_node}): STOPPED")
FAIL.append(f"ct-stopped:{name}")
else:
print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 5: Config YAML Integrity (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_config_integrity():
"""Verify agent config.yaml parses as valid YAML."""
for name, agent in AGENTS.items():
host = agent.get("host")
user = agent.get("user")
if not host or not user:
print(f"{name}: cannot SSH — skip config check")
continue
# Check YAML parses
yaml_ok = ssh(host,
"python3 -c "
'"import yaml; yaml.safe_load(open(\'/root/.hermes/config.yaml\')); print(\'OK\')" '
"2>&1 || echo 'FAIL'",
user=user)
if not yaml_ok:
print(f"{name}: SSH UNREACHABLE (config check skipped)")
FAIL.append(f"config-unreachable:{name}")
elif "OK" in yaml_ok:
print(f"{name}: config.yaml valid YAML")
else:
print(f"{name}: config.yaml YAML ERROR — {yaml_ok[:120]}")
FAIL.append(f"config-yaml-error:{name}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 6: Wrapper/CLI Integrity (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_wrapper_integrity():
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
for name, agent in AGENTS.items():
host = agent.get("host")
user = agent.get("user")
if not host or not user:
print(f"{name}: cannot SSH — skip wrapper check")
continue
# Check wrapper exists
wrapper = ssh(host, "ls -la /root/.local/bin/hermes 2>/dev/null", user=user)
if not wrapper:
# Check alternate wrapper locations
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
if not wrapper:
print(f"{name}: NO HERMES CLI WRAPPER FOUND")
FAIL.append(f"wrapper-missing:{name}")
continue
else:
print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)")
# Check wrapper has correct infisical path
infisical_path_valid = ssh(host,
"head -20 /root/.local/bin/hermes 2>/dev/null | grep -q '/usr/bin/infisical' && echo OK || echo MISS",
user=user)
if infisical_path_valid == "MISS":
# Check if infisical exists on path
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
if not inf_actual:
print(f"{name}: INFISICAL NOT INSTALLED (wrapper broken)")
FAIL.append(f"wrapper-no-infisical:{name}")
else:
print(f" ⚠️ {name}: wrapper infisical path may be wrong (infisical at {inf_actual})")
FAIL.append(f"wrapper-infisical-path:{name}")
# Check hermes-real exists
hermes_real = ssh(host,
"ls -la /root/.local/bin/hermes-real 2>/dev/null || echo MISS",
user=user)
if not hermes_real or hermes_real.strip() == "MISS":
# Check venv path
hermes_real = ssh(host,
"ls -la /usr/local/lib/hermes-agent/venv/bin/hermes 2>/dev/null || echo MISS",
user=user)
if not hermes_real or hermes_real.strip() == "MISS":
print(f"{name}: hermes-real NOT FOUND (wrapper broken)")
FAIL.append(f"wrapper-no-hermes-real:{name}")
else:
print(f"{name}: hermes-real at alt path")
# Check the .env file has the key
env_has_key = ssh(host,
"grep -c 'LITELLM_API_KEY' /root/.hermes/.env 2>/dev/null || echo 0",
user=user)
if env_has_key and env_has_key.strip() not in ("", "0"):
print(f"{name}: wrapper + .env key present")
else:
print(f" ⚠️ {name}: .env may be missing LITELLM_API_KEY entry")
# ═══════════════════════════════════════════════════════════════════
# CHECK 7: Vault Secret Non-Emptiness (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_vault_secrets():
"""Verify agent-specific vault secrets are non-empty and start with sk-."""
for name, agent in AGENTS.items():
vault_key_name = agent.get("vault_key")
if not vault_key_name:
continue
key = agent.get("key")
if not key:
print(f"{name}: vault secret {vault_key_name} MISSING or EMPTY")
FAIL.append(f"vault-empty:{name}:{vault_key_name}")
elif not key.startswith("sk-"):
print(f"{name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')")
FAIL.append(f"vault-bad-format:{name}:{vault_key_name}")
else:
print(f"{name}: vault {vault_key_name}=sk-...{key[-4:]}")
# ═══════════════════════════════════════════════════════════════════
# DEPLOY: copy updated script to /root/scripts/ on local host
# ═══════════════════════════════════════════════════════════════════
def deploy_self():
"""Copy this script to /root/scripts/agent-health-check.py if out of date."""
dest = "/root/scripts/agent-health-check.py"
try:
with open(__file__, "r") as f:
current = f.read()
if os.path.isfile(dest):
with open(dest, "r") as f:
existing = f.read()
if current == existing:
return # Already deployed
# Write new version
with open(dest, "w") as f:
f.write(current)
os.chmod(dest, 0o755)
print(f" 📦 Deployed updated script to {dest}")
except:
pass # Not fatal if deploy fails
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# MAIN # MAIN
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
@@ -458,12 +235,8 @@ def main():
quiet = "--quiet" in sys.argv quiet = "--quiet" in sys.argv
as_json = "--json" in sys.argv as_json = "--json" in sys.argv
# Self-deploy to canonical location
if not quiet and "--no-deploy" not in sys.argv:
deploy_self()
if not quiet: if not quiet:
print(f"🏥 Agent Health Check v2 {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") print(f"🏥 Agent Health Check — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}")
print() print()
print("🔑 LiteLLM Keys:") print("🔑 LiteLLM Keys:")
@@ -476,26 +249,11 @@ def main():
print("🤖 Agent Gateways:") print("🤖 Agent Gateways:")
check_agents() check_agents()
print()
print("🖥️ CT Liveness:")
check_ct_liveness()
print()
print("📝 Config Integrity:")
check_config_integrity()
print()
print("🔌 Wrapper/CLI Integrity:")
check_wrapper_integrity()
print()
print("🔐 Vault Secrets:")
check_vault_secrets()
if FAIL: if FAIL:
print(f"\n{len(FAIL)} FAILURE(S): {' | '.join(FAIL)}") print(f"\n{len(FAIL)} FAILURE(S): {' | '.join(FAIL)}")
if quiet: if quiet:
# In quiet mode, only print failures as a single alert line
print(f"ALERT agent-health:{','.join(FAIL)}") print(f"ALERT agent-health:{','.join(FAIL)}")
elif not quiet: elif not quiet:
print("\n✅ All checks passed") print("\n✅ All checks passed")