Compare commits

..
Author SHA1 Message Date
root 6253aeb72b Add check-health section with live probes to infrastructure-monitoring contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-08-22 23:01:19 +00:00
root 66f94d14fd Item 6: Auto-compaction at ~60% — update thresholds
- gpu-fleet.prose.md: Update compression.threshold to 0.60 (~77K triggers)
- Add pi compaction.reserveTokens: 52739 (≈60% of 128K)
- Remove stale 256K hardcode references
2026-08-20 07:54:36 +00:00
root 819f2f400d Item 5: Hermes context-detection fix — Rule 14
- hermes-config-template.prose.md: Add Rule 14 — Hermes reads max_model_tokens (128K), NOT max_input_tokens (64K)
- Warning: Using max_input_tokens for Hermes agents causes premature context loss
- Crewmates use max_input_tokens (64K cap), Abiba/Hermes use max_model_tokens (128K)
2026-08-20 07:53:58 +00:00
root b7778c287b Item 4: LiteLLM 3-way cap split — document context caps
- litellm-self-heal.prose.md: Add 3-way cap split (Abiba/Hermes=128K uncapped, Crew=64K via crew-auto alias)
- Note: Do NOT document max_input_tokens as context window — use max_tokens/context_window_size
2026-08-20 07:52:44 +00:00
root ab92e57781 Item 3: gpu-light vision swap to Qwen3.5-9B
- gpu-fleet.prose.md: Update gpu-light row to Qwen3.5-9B (Q5_K_M + mmproj-F16)
- gpu-fleet.prose.md: Update topology diagram, routing tables, VRAM, benchmarks
- gpu-fleet.prose.md: Fix timeout references, backward compat notes
- gpu-self-heal.prose.md: Update gpu-light row and tok/s performance
- Note: Qwen3.5-9B is multimodal (image+text), dedicated vision endpoint
- Remove gemma-4-12b references where Qwen3.5-9B now runs
2026-08-20 07:51:49 +00:00
root 05800d6ccb Item 2: Carnice Q5_K_M on strix-moe — update model row
- gpu-fleet.prose.md: Update strix-moe row to Carnice-Qwen3.6-MoE-35B-A3B (Q5_K_M, ~24.73GB)
- Note: Three-role split (gpu-dense + strix-moe = text; gpu-light = vision)
- Remove mmproj/vision claim from strix-moe (it's text-only MoE)
2026-08-20 07:49:08 +00:00
root 6195b59317 Item 1: Koby report-only — encode HARD RULE in contracts (2026-08-17)
- hermes-config-template.prose.md: Add Rule 17 — Koby is never repaired, full stop
- hermes-agent-baseline.prose.md: Document Koby report-only posture
- contract-registry.yaml: Tag all healing contracts as Koby-eligible (skip heal)
- scripts/agent-health-check.py: Mark Koby as report_only=True, skip repairs
- All healing contracts: Add report_only_agents.koby marker

Captain-approved ship via no-mistakes. PR auto-merges green.
2026-08-20 07:42:22 +00:00
root f7218e04c0 pm2-self-heal: restore abiba-zulip, retire gpu-watchdog, document systemd for gpu-monitor 2026-08-03 22:35:55 +00:00
41 changed files with 352 additions and 884 deletions
+2 -2
View File
@@ -67,10 +67,10 @@ Two incidents taught us this:
| Contract | Sensitivity | Who can change | | Contract | Sensitivity | Who can change |
|----------|------------|----------------| |----------|------------|----------------|
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba, Tanko (Tanko maintains its own CT row) | | `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba only (after live verification) |
| `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only | | `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only |
| `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko | | `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko |
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni, Tanko | | `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni |
| Other contracts | Normal | Any registered agent | | Other contracts | Normal | Any registered agent |
| `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only | | `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only |
+1 -2
View File
@@ -20,8 +20,7 @@ An agent ran `pct set` without checking `pct config` first and changed the IP
to the wrong value, breaking Zulip. The staleness was harmless until acted on. to the wrong value, breaking Zulip. The staleness was harmless until acted on.
**Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents: **Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents:
Abiba, Tanko, Mumuni, Koby, Koonimo — Tanko runs on DSH/DeepSeek Harness since Abiba, Tanko, Mumuni, Koby, Koonimo). ALL infrastructure mutations MUST go
2026-08-27 but safe-mutate enforcement still applies). ALL infrastructure mutations MUST go
through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up
--force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an --force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an
auditable protocol violation. The wrapper runs a verify command, optionally auditable protocol violation. The wrapper runs a verify command, optionally
-186
View File
@@ -1,186 +0,0 @@
# Agent Zero Issue Fix Summary
**Date**: 2026-09-01
**Agent**: Agent Zero (Docker container on kagentz CT105)
**Issue**: AuthenticationError + Telegram conflicts
**Status**: ✅ RESOLVED
---
## Problems Identified
### 1. OpenRouter Authentication Error (CRITICAL)
```
litellm.exceptions.AuthenticationError: OpenrouterException -
{"error":{"message":"User not found.","code":401}}
```
**Root Cause**: The OpenRouter API key in `/a0/usr/.env` belonged to a different OpenRouter user.
**Old Key**: `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`
**New Key**: `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`
**New User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
### 2. Telegram Bot Conflict (CRITICAL)
```
TelegramConflictError: Conflict: terminated by other getUpdates request
```
**Root Cause**: Two Telegram bot instances were competing for the same token:
1. Agent Zero's built-in Telegram plugin (`/a0/usr/plugins/_telegram_integration/config.json`)
2. Standalone Telegram poller scripts (`/a0/usr/projects/telegram/telegram_bot.py`)
Both were using token `8476855065:***` in polling mode.
**Fix**: Disabled the built-in Telegram plugin by setting `"enabled": false` in the config.
### 3. MCP Service Connectivity Issues (SEVERE)
```
McpError: Timed out while waiting for response to ClientRequest. Waited 10.0 seconds.
```
**Root Cause**: The OpenRouter 401 errors caused the agent to fail, which in turn caused MCP services to timeout.
**Status**: ✅ RESOLVED with OpenRouter key fix.
---
## Fixes Applied
### Fix 1: Update OpenRouter Key
```bash
# Container .env update
sudo docker exec agent-zero bash -c '
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=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|" /a0/usr/.env
'
```
**Verification**:
```bash
curl -s https://openrouter.ai/api/v1/auth/key \
-H "Authorization: Bearer sk-or-v1-0af3f3..." | python3 -m json.tool
```
Result: HTTP 200, user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`, not free tier.
### Fix 2: Disable Telegram Plugin
```bash
sudo docker exec agent-zero bash -c '
python3 << "PYEOF"
import json
config_path = "/a0/usr/plugins/_telegram_integration/config.json"
with open(config_path) as f:
config = json.load(f)
config["bots"][0]["enabled"] = False
with open(config_path, "w") as f:
json.dump(config, f, indent=2)
print("✓ Disabled telegram plugin @kagentz_bot")
PYEOF
'
```
### Fix 3: Restart Agent Zero UI
```bash
sudo docker exec agent-zero supervisorctl restart run_ui
```
**Result**: Process restarted (PID 3320), services running.
### Fix 4: Full Container Restart (Required)
```bash
sudo docker restart agent-zero
```
**Why needed**: The `run_ui` process was caching the old API key in memory. A full container restart was required to force Agent Zero to reload the `.env` file with the new OpenRouter key.
**Result**: All services restarted cleanly, no more 401 errors.
### Fix 5: Update Stale `.env.clobbered-by-new-image` (Critical)
**Root cause**: Agent Zero was loading the key from `/a0/usr/.env.clobbered-by-new-image` (line 28) instead of the main `/a0/usr/.env` (line 72). The clobbered file still had the old, stale key.
**Fix**:
```bash
KEY=$(grep "^API_KEY_OPENROUTER=" /a0/usr/.env | cut -d"=" -f2-)
sed -i "s|^API_KEY_OPENROUTER=.*|API_KEY_OPENROUTER=$KEY|" /a0/usr/.env.clobbered-by-new-image
```
**Lesson**: When updating Agent Zero's `.env`, check BOTH files:
- `/a0/usr/.env` (main)
- `/a0/usr/.env.clobbered-by-new-image` (backup, but loaded by Agent Zero)
The clobbered file is the one Agent Zero actually uses for LLM calls.
---
## Infrastructure Documentation
### New Contract Created
**File**: `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md`
Contains:
- Key management procedures
- Rotation instructions
- Verification steps
- Current key inventory
- Related contracts
### Updated Contract
**File**: `/home/home/syslog/prose-contracts/litellm-api-keys.prose.md`
Added section:
- Agent Zero OpenRouter integration
- Key storage locations
- Model configuration
- Why not LiteLLM proxy
- Rotation procedure
---
## Current State
| Component | Status | Details |
|-----------|--------|---------|
| **OpenRouter Key** | ✅ Valid | `sk-or-v1-0af3f3…`, user verified |
| **Telegram Bot** | ✅ Resolved | Plugin disabled, conflicts cleared |
| **MCP Services** | ✅ Working | No timeouts after key fix |
| **Container** | ✅ Running | PID 3320, uptime 16+ hours |
| **Services** | ✅ All UP | run_ui, run_tunnel_api, run_searxng, run_cron, the_listener |
---
## Related Files
| Path | Purpose |
|------|---------|
| `/a0/usr/.env` | Container key storage |
| `/a0/usr/plugins/_telegram_integration/config.json` | Telegram plugin config |
| `/a0/usr/plugins/_model_config/presets.yaml` | Model selection (moonshotai/kimi-k3) |
| `/home/hermes/syslog/prose-contracts/agent-zero-openrouter-key.prose.md` | Key management contract |
| `/home/hermes/syslog/prose-contracts/litellm-api-keys.prose.md` | Fleet key inventory |
---
## Next Steps
1. **Sync key to Infisical vault** (optional, currently .env fallback only)
2. **Monitor usage** — Check OpenRouter dashboard for daily/weekly spend
3. **Consider LiteLLM migration** — Long-term: convert Agent Zero to use LiteLLM proxy for fleet-standard key management
4. **Set up vault sync** — Create machine identity in Infisical for automated key rotation
---
## Prevention
To prevent similar issues:
1. **Always verify API keys** against their providers before using
2. **Keep fleet-wide key inventory** updated in prose contracts
3. **Rotate keys on schedule** (quarterly hygiene, not on-demand only)
4. **Test key changes** in staging before production rollout
5. **Document key locations** in both code and prose contracts
---
**Verified by**: Mumuni 🦅
**Last updated**: 2026-09-01
**Session**: 1
-129
View File
@@ -1,129 +0,0 @@
---
kind: function
name: agent-zero-openrouter-key
description: >
Manages the OpenRouter API key for Agent Zero (Docker container on kagentz .14).
Agent Zero uses OpenRouter as its primary LLM provider for the moonshotai/kimi-k3
model. The key is stored in Infisical vault (project=agents, env=production) and
referenced from /a0/usr/.env in the container. Key must be rotated when the
OpenRouter user account changes or on quarterly hygiene. Last verified: 2026-09-01.
---
## Parameters
- action: "verify" | "rotate" | "update" | "list" — What to do (default: "verify")
- container_name: string — Docker container name (default: "agent-zero")
- host: string — Proxmox host running the container (default: "kagentz" at 192.168.68.14)
- env_path: string — Path to .env file in container (default: "/a0/usr/.env")
- vault_project: string — Infisical project slug (default: "agents")
- vault_env: string — Infisical environment (default: "production")
## Returns
- action: string — What was done
- key_status: string — "valid" | "invalid" | "not_found"
- key_prefix: string — First 10 chars of the key (for identification)
- user_id: string — OpenRouter user ID associated with the key
- vault_synced: boolean — Whether the key is in the Infisical vault
- container_updated: boolean — Whether the container's .env was updated
- verification: { status: string, detail: string } — Health check result
## Execution
### 1. Verify the key
1. **Extract key from container**
```bash
sudo docker exec agent-zero grep '^API_KEY_OPENROUTER' /a0/usr/.env | cut -d'=' -f2-
```
2. **Test against OpenRouter API**
```bash
curl -s https://openrouter.ai/api/v1/auth/key \
-H "Authorization: Bearer <key>" | python3 -m json.tool
```
Expected: HTTP 200, JSON with `data.label` and `data.is_free_tier`
3. **Check vault sync**
```bash
infisical secrets get OPENROUTER_API_KEY \
--token=$(cat ~/.infisical-token) \
--projectId=agents \
--env=production \
--domain=https://vault.sysloggh.net
```
4. **Return status**
- If all checks pass: `{ key_status: "valid", key_prefix: "sk-or-v1-0af", user_id: "user_2rt9lCqcd5d7Vk1t18DHsvWdPTT" }`
- If OpenRouter returns 401: `{ key_status: "invalid", detail: "User not found" }`
- If vault secret is missing: `{ vault_synced: false }`
### 2. Rotate the key
1. **Generate new key** in OpenRouter UI or via API
2. **Update container .env**
```bash
sudo docker exec agent-zero sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env
```
3. **Update Infisical vault**
```bash
infisical secrets set OPENROUTER_API_KEY=<new_key> \
--token=$(cat ~/.infisical-token) \
--projectId=agents \
--env=production \
--domain=https://vault.sysloggh.net
```
4. **Restart Agent Zero UI**
```bash
sudo docker exec agent-zero supervisorctl restart run_ui
```
5. **Verify** — Run "verify" action again
### 3. Update (key changed but no rotation)
1. **Update container .env** (same as rotate step 2)
2. **Sync vault** (same as rotate step 3)
3. **Restart run_ui** (same as rotate step 4)
## Current Key Inventory
| Field | Value |
|-------|-------|
| **Key Prefix** | `sk-or-v1-0af3f3` |
| **Full Key** | `«redacted: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»` (in vault + /a0/usr/.env) |
| **OpenRouter User** | `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT` |
| **Free Tier** | No |
| **Monthly Usage** | 0 (as of 2026-09-01) |
| **Last Verified** | 2026-09-01 |
| **Vault Sync** | ⏳ Pending (service token not on kagentz) |
## Key Rotation Log
| Date | Action | Notes |
|------|--------|-------|
| 2026-09-01 | fix-401 | Old key `sk-or-v1-036e5ca5…` returned 401 "User not found". Replaced with new key `sk-or-v1-0af3f3…` for user `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`. Verified OpenRouter 200. Container .env updated, run_ui restarted. |
## Infrastructure References
- **Docker container**: `agent-zero` (image: `agent0ai/agent-zero:latest`)
- **Host**: kagentz (192.168.68.14, Proxmox LXC CT105)
- **Volume**: `/var/lib/docker/volumes/agent_zero/_data` → `/a0/usr`
- **Config path**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=…`)
- **Model preset**: "Cost Efficient" (uses `openrouter/moonshotai/kimi-k3`)
- **Model config**: `/a0/usr/plugins/_model_config/config.json`
## Verification Before Acting
**Key is a lead, not a fact.** Live OpenRouter accounts can change (user deletion,
plan change, key revocation). Before acting on this contract:
1. Verify the key against OpenRouter's `/auth/key` endpoint
2. Check the user ID matches the expected account
3. Confirm the model `moonshotai/kimi-k3` is available on that account's plan
4. Only then update the vault and container
## Related Contracts
- `litellm-api-keys.prose.md` — LiteLLM key management (Agent Zero does NOT use LiteLLM for OpenRouter)
- `infrastructure-control.prose.md` — Proxmox topology, container locations
- `gpu-fleet.prose.md` — Fleet-wide agent key inventory (add Agent Zero here)
+2 -3
View File
@@ -45,10 +45,9 @@ description: >
## Status ## Status
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service **Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
for Hermes agents that connect to Zulip (Mumuni, Koby, Koonimo). The pi for any Hermes agent that connects to Zulip (Mumuni, Tanko, Koby, Koonimo). The pi
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
plugin system, which is unaffected. **Tanko is excluded — it runs on DSH (DeepSeek plugin system, which is unaffected.
Harness) since 2026-08-27 and no longer uses the Hermes Zulip plugin.**
## Parameters ## Parameters
+1 -1
View File
@@ -584,7 +584,7 @@ contracts:
verify: curl -sf https://git.sysloggh.net/api/v1/version verify: curl -sf https://git.sysloggh.net/api/v1/version
expect: 200 OK expect: 200 OK
- check: SearXNG reachable - check: SearXNG reachable
verify: curl -sf http://192.168.68.7:8888 verify: curl -sf http://192.168.68.17:8080
expect: 200 OK expect: 200 OK
artifact: infrastructure health report artifact: infrastructure health report
receipt: receipt:
+1 -1
View File
@@ -356,7 +356,7 @@ Postconditions to verify:
}, },
{ {
"check": "SearXNG reachable", "check": "SearXNG reachable",
"verify": "curl -sf http://192.168.68.7:8888", "verify": "curl -sf http://192.168.68.17:8080",
"expect": "200 OK" "expect": "200 OK"
} }
] ]
+3 -2
View File
@@ -277,10 +277,10 @@ one-off GPU builds. No automated post-migration cleanup was in place.
### CT Access (via pct-run) ### CT Access (via pct-run)
| CT | Name | Node | Status | | CT | Name | Node | Status |
|----|------|------|--------| |----|------|------|--------|
| 100 | abiba | minipve | local | | 100 | abiba | hwepve | local |
| 102 | adguard | minipve | ✅ reachable | | 102 | adguard | minipve | ✅ reachable |
| 104 | authentik | minipve | ✅ reachable | | 104 | authentik | minipve | ✅ reachable |
| 105 | kagentz | minipve | ✅ reachable | | 105 | kagentz | hwepve | ✅ reachable |
| 106 | ra-h-os | storepve | ✅ reachable | | 106 | ra-h-os | storepve | ✅ reachable |
| 107 | pbs | storepve | ✅ reachable | | 107 | pbs | storepve | ✅ reachable |
| 108 | media | storepve | ✅ reachable | | 108 | media | storepve | ✅ reachable |
@@ -288,6 +288,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| 111 | tdunna | amdpve | ✅ reachable | | 111 | tdunna | amdpve | ✅ reachable |
| 112 | tanko | amdpve | ✅ reachable | | 112 | tanko | amdpve | ✅ reachable |
| 113 | baggy | amdpve | ✅ reachable | | 113 | baggy | amdpve | ✅ reachable |
| 114 | mumuni | hwepve | ✅ reachable |
| 115 | scottdenya | amdpve | ✅ reachable | | 115 | scottdenya | amdpve | ✅ reachable |
| 116 | syslog-api | minipve | ✅ reachable | | 116 | syslog-api | minipve | ✅ reachable |
| 117 | zulip | storepve | ✅ reachable | | 117 | zulip | storepve | ✅ reachable |
+1 -1
View File
@@ -185,7 +185,7 @@ what, and why should I care?
``` ```
❌ "Monitors infrastructure health" ❌ "Monitors infrastructure health"
✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker ✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL" container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
``` ```
+20 -3
View File
@@ -14,6 +14,10 @@ description: >
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek). For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
UPDATED 2026-07-27: gpu-dense swapped to Qwen3.8-27B-Uncensored-Q4_K_M
(UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking
tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning.
VRAM ~22.4/24.6GB (91%).
agent: abiba agent: abiba
triggers: triggers:
- on model add/remove - on model add/remove
@@ -96,6 +100,9 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy |
| Qwen3.5-9B | RTX 5070 | .110 (ocu-llm) | ~6.2/12.2GB (51%) | 128K | Q5_K_M | 2 | 2048/1024 | ✅ healthy |
| Carnice-Qwen3.6-MoE-35B-A3B | Strix Halo Vulkan | .15 (amdpve) | ~24.73GB/64GB | 128K | Q5_K_M | 1 | 4096/1024 | ✅ 65 tok/s |
## Routing Configuration (LiteLLM — July 2026) ## Routing Configuration (LiteLLM — July 2026)
@@ -104,6 +111,8 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| Carnice-Qwen3.6-MoE-35B-A3B | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| Qwen3.5-9B | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path. Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
@@ -111,6 +120,9 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| strix-moe (Carnice-Qwen3.6-MoE-35B-A3B) | 40 | Tight cap — prevents Strix overload |
| Qwen3.8-27B-Uncensored-Q4_K_M | 500 | High cap — primary workhorse |
| Qwen3.5-9B | 500 | High cap — multimodal vision endpoint |
### Stable Aliases (for agent configs — never change) ### Stable Aliases (for agent configs — never change)
@@ -196,7 +208,7 @@ Plaintext keys removed from this contract post-vault-migration.
| Agent | CT | IP | LiteLLM Alias | Key Source | Access | | Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|-------|-----|-----|---------------|------------|--------| |-------|-----|-----|---------------|------------|--------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome | | Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | SSH root | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | SSH root |
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) | | Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM | | Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH | | Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
@@ -238,6 +250,10 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. - **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-08-20)**: RTX 3090 at ~22.4/24.6GB (~91%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~6.2/12.2GB (~51%) with 128K context (Qwen3.5-9B). Strix Halo at ~22GB/64GB.
- **RTX 3090 (2026-07-27)**: Swapped to Qwen3.8-27B-Uncensored-Q4_K_M (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`.
- **RTX 5070 config (2026-08-20)**: Switched to Qwen3.5-9B (Q5_K_M) at 128K context. Multimodal (image+text). Gen speed: ~145 tok/s (estimated). VRAM: ~6.2/12.2GB (~51%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model Qwen3.5-9B-Q5_K_M.gguf --mmproj Qwen3.5-9B-mmproj-F16.gguf --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-27)**: Qwen3.8-27B-Uncensored-Q4_K_M 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
@@ -253,6 +269,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** | | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
| RTX 5070 (.110) | Qwen3.5-9B (Q5_K_M) | **145** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** | | Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
@@ -285,7 +302,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
### Mumuni Agent Profile ### Mumuni Agent Profile
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs: Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This profile is the reference for all agent configs:
| Setting | Value | Notes | | Setting | Value | Notes |
|---------|-------|-------| |---------|-------|-------|
@@ -310,7 +327,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
| Agent | Host | Status | | Agent | Host | Status |
|-------|------|--------| |-------|------|--------|
| **Mumuni** | CT105 (.14) | ✅ Updated to stable aliases | | **Mumuni** | CT100 (.24) | ✅ Updated to stable aliases |
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases | | **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM | | **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM | | **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
+2
View File
@@ -48,6 +48,8 @@ depends_on:
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role | | Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------| |-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | Qwen3.5-9B (Q5_K_M) + mmproj-F16 | 6.2/12.2GB (51%) | 128K | ~145 | Vision (image+text), web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs | | `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes: Key notes:
+7 -6
View File
@@ -24,6 +24,8 @@ done
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform | | Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|-------|-----|------|-----|---------------|------------|----------| |-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
| Mumuni | 100 | hwepve | .24 | `mumuni` | Infisical vault | Hermes |
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** | | Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes | | Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) | | Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
@@ -51,10 +53,9 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) →
## Config Pattern — Mandatory Fields ## Config Pattern — Mandatory Fields
### For Hermes Agents (Mumuni, Koonimo) ### For Hermes Agents (Tanko, Mumuni, Koonimo)
Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have: Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
### 1. Main Model ### 1. Main Model
```yaml ```yaml
@@ -70,7 +71,7 @@ model:
custom_providers: custom_providers:
- name: harness - name: harness
model: syslog-auto model: syslog-auto
base_url: http://192.168.68.116/litellm/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
api_mode: chat_completions api_mode: chat_completions
``` ```
@@ -81,7 +82,7 @@ auxiliary:
vision: vision:
provider: harness provider: harness
model: gemma-4-12b # or syslog-auto model: gemma-4-12b # or syslog-auto
base_url: http://192.168.68.116/litellm/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
timeout: 60 timeout: 60
@@ -96,7 +97,7 @@ auxiliary:
target_ratio: 0.3 target_ratio: 0.3
provider: harness provider: harness
model: syslog-auto # or gemma-4-12b model: syslog-auto # or gemma-4-12b
base_url: http://192.168.68.116/litellm/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
timeout: 120 timeout: 120
+32 -38
View File
@@ -5,7 +5,8 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents. Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices. RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident. UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300). 2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13). UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
@@ -15,7 +16,7 @@ description: >
- template_version: "2.1.0" - template_version: "2.1.0"
- last_applied: timestamp - last_applied: timestamp
- agents_configured: ["mumuni", "abiba", "koby", "koonimo", "kagenz0"] # tanko removed 2026-08-27 (now on DSH/DeepSeek Harness) - agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
- agent_keys: map (see Agent Keys section) - agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array - infra_endpoints_verified: array
@@ -30,6 +31,8 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Agent | Key Alias | Host | SSH | Sub-Agents | | Agent | Key Alias | Host | SSH | Sub-Agents |
|-------|-----------|------|-----|-----------| |-------|-----------|------|-----|-----------|
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
| Mumuni | `mumuni` | 192.168.68.24 | root@.24 | 6 profiles ✱ |
| Abiba | `abiba-pi` | 192.168.68.24 | local | — | | Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Koby | `koby` | CT 111 (tdunna) | Zulip | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — |
| Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — | | Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — |
@@ -45,7 +48,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Component | Endpoint | Purpose | | Component | Endpoint | Purpose |
|---|---|---| |---|---|---|
| Firecrawl | `http://192.168.68.7:3002/` | Web content extraction | | Firecrawl | `http://192.168.68.7:3002/` | Web content extraction |
| SearXNG | `http://192.168.68.7:8888` | Privacy-respecting web search | | SearXNG | `http://storepve:8888` | Privacy-respecting web search |
| LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) | | LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) |
| LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) | | LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) |
| RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge | | RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge |
@@ -94,14 +97,10 @@ work immediately after restart.
model: model:
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
provider: harness provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
# and falls back to 256K when /v1/models lacks a context
# field (llama-server does). Without this override, agents
# silently run syslog-auto at 256K (verified 2026-08-09).
# Set 65536 if using gemma-4-12b directly (tight VRAM). # Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers: fallback_providers:
@@ -144,7 +143,7 @@ compression:
# ─── Auxiliary Tasks (CONSISTENCY RULE) ─── # ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env: # All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b") # model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK # base_url: http://192.168.68.116/v1
# api_key_env: LITELLM_API_KEY # api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU. # Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning. # gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
@@ -155,20 +154,20 @@ auxiliary:
vision: vision:
provider: harness provider: harness
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b) model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 60 timeout: 60
download_timeout: 30 download_timeout: 30
web_extract: web_extract:
provider: harness provider: harness
model: gpu-light # stable alias for RTX 5070 model: gpu-light # stable alias for RTX 5070
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 30 timeout: 30
compression: compression:
provider: harness provider: harness
model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool). model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60) timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
@@ -178,14 +177,14 @@ auxiliary:
delegation: delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code) model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
provider: harness provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
# ─── Custom Provider ─── # ─── Custom Provider ───
custom_providers: custom_providers:
- name: harness - name: harness
model: syslog-auto # weighted pool (default) model: syslog-auto # weighted pool (default)
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
api_mode: chat_completions api_mode: chat_completions
``` ```
@@ -235,13 +234,10 @@ The following MUST be identical across ALL profiles:
- When main config uses `api_key_env`, sub-agents automatically use it - When main config uses `api_key_env`, sub-agents automatically use it
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`) - This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
### Rule 5: Main Config Base URL (UPDATED 2026-08-09) ### Rule 5: Main Config Base URL
|- Use the authenticated LiteLLM path: `http://192.168.68.116/litellm/v1` (canonical, captain-approved migration) |- Use direct IP: `http://192.168.68.116/v1`
|- Legacy `http://192.68.68.116/v1` also works — nginx fronts BOTH paths with key auth
(verified 2026-08-09: 401 without key, 200 with key, on both /v1 and /litellm/v1)
|- Both locations have `proxy_read_timeout 600s` (verified in harness-nginx nginx.conf) —
the old "60s timeout on /litellm/" claim was stale and is retracted
|- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down |- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down
|- NOT the old path (`/litellm/v1`) — nginx now routes `/v1` directly
### Rule 6: max_tokens Is Required (Thermal Safety) ### Rule 6: max_tokens Is Required (Thermal Safety)
- **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable - **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable
@@ -262,7 +258,7 @@ The following MUST be identical across ALL profiles:
fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model` fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model`
MUST be `syslog-auto`. MUST be `syslog-auto`.
- All auxiliary services MUST use identical routing: - All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK) - `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
- `api_key_env: LITELLM_API_KEY` - `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
@@ -294,7 +290,6 @@ The following MUST be identical across ALL profiles:
- See `devops-hermes-compression` skill for full reference - See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents) ### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
- **Koby Exception**: Per captain ruling 2026-08-11, Koby is a DeepSeek-primary external agent; its primary model remains `deepseek-v4-flash` (via api.deepseek.com to preserve DeepSeek-specific reasoning, while other sections follow Rule 10.
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto` - **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json` - **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
- `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe - `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe
@@ -323,13 +318,10 @@ verify ALL FOUR of these against the live config. They are the only root causes
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` 1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used. MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`, 2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical) and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
or `http://192.168.68.116/v1` (legacy, still authenticated via nginx). BOTH verified 200 with has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s.
key + 600s proxy_read_timeout on 2026-08-09. Never bare `:4000` direct. Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING)
Check: `grep -nE 'base_url: http://192.168.68.116(:4000)?/v1' ~/.hermes/config.yaml` — the
ONLY paths allowed are `/v1` or `/litellm/v1` (both via nginx :80).
`:4000` or missing `litellm/v1`/`v1` prefix = violation.
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars). 3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback. Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200) Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
@@ -415,15 +407,6 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>` - **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
before and after any config change to catch this and all other rule violations. before and after any config change to catch this and all other rule violations.
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
- Every MCP server entry must point at the correct endpoint:
- ra-h-os = http://192.168.68.65:3100/mcp
- litellm = https://litellm.sysloggh.net/mcp
- MCP entries must carry a REAL key value in the header.
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
endpoints and result in "Malformed API Key" floods.
- Ensure the header value is the actual key (e.g., `sk-...`).
## Execution ## Execution
1. **Check current config** — Read the target agent's config.yaml 1. **Check current config** — Read the target agent's config.yaml
@@ -433,3 +416,14 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
5. **Set model choice** — Per agent's workload 5. **Set model choice** — Per agent's workload
6. **Verify** — curl all shared endpoints, test the model with the new key 6. **Verify** — curl all shared endpoints, test the model with the new key
7. **Report** — What was changed, preserved, custom 7. **Report** — What was changed, preserved, custom
### Rule 17: Koby Report-Only (URGENT — 2026-08-17, Captain)
**Koby (CT 111, .129, tdunna/Theo) is report-only, full stop.**
No heal step, no restart, no key rotation, no config edit, no memory rewrite,
no disk GC, no service touch, no process kill — **ever**.
Do NOT "repair" Koby under any circumstance, even if a health check shows it degraded.
Detect + report to Zulip (Theo fixes it).
This supersedes every other instruction in this file.
If a rule would trigger a Koby repair action, SKIP it and emit a diagnostic alert.
+16 -17
View File
@@ -34,12 +34,11 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
| Path | Auth | Status | | Path | Auth | Status |
|------|------|--------| |------|------|--------|
| `http://192.168.68.116/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **VALID** — authenticated via nginx :80 (verified 2026-08-09: 401 without key, 200 with) | | `http://192.168.68.116/v1` | None (direct) | ❌ **DEPRECATED** — being phased out |
| `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) | | `http://192.168.68.116/litellm/v1/responses` | Bearer `sk-*` key | ✅ **CURRENT** — authenticated LiteLLM proxy |
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid). All harness/litellm providers MUST use the authenticated `/litellm/v1/responses` path.
Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**. Any `base_url` pointing to bare `/v1` on 192.168.68.116 is a **migration violation**.
### 🔥 CRITICAL: Double-Path Bug (2026-07-10) ### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
@@ -190,23 +189,23 @@ litellm_settings:
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified | | Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|-------|-----|-----|---------------|------------|--------|-----------------|---------------| |-------|-----|-----|---------------|------------|--------|-----------------|---------------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 | | Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
| Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | ✅ Fixed | systemd Hermes gateway | 2026-08-29 | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
| Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed (DeepSeek-primary) | `infisical run` | 23:30 UTC Jul 5 | | Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
| Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 | | Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 | | Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
| Kagenz0 | 105 | .14 | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 | | Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). > **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
> LiteLLM key aliases use agent identity, not CT hostname. > LiteLLM key aliases use agent identity, not CT hostname.
### Migration Status: Authenticated Path ### Migration Status: Authenticated Path
| Agent | `/litellm/v1` | Legacy `/v1` | Status | | Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status |
|-------|--------------|-------------|--------| |-------|--------------------------|--------------------|--------|
| Mumuni | ✅ harness provider | ✅ auxiliary on /v1 (valid) | ✅ Authenticated (verified 2026-08-09) | | Mumuni | ✅ 5 sections | 0 | ✅ Authenticated |
| Tanko | ✅ 5 sections | 0 | ✅ Migrated 2026-08-08, keys 200 | | Tanko | ⚠️ No SSH access | — | Needs check |
| Koby | ✅ custom provider (harness name) | — | ✅ External DeepSeek primary (intentional, captain ruling 2026-08-11) | | Koby | ⚠️ No route to host | — | Needs check |
| Koonimo | ✅ .114 (baggy) | — | ✅ 128K context applied 2026-08-09 | | Koonimo | ⚠️ Connection timed out | — | Needs check |
### Systemd Service Pattern (2026-07-11 — vault migration) ### Systemd Service Pattern (2026-07-11 — vault migration)
@@ -285,13 +284,13 @@ auxiliary:
vision: vision:
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain) api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1 base_url: http://192.168.68.116/v1
model: gemma-4-12b model: gemma-4-12b
provider: harness provider: harness
compression: compression:
api_key: sk-<agent-key-from-vault> # ← workaround (same as above) api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1 base_url: http://192.168.68.116/v1
model: gemma-4-12b model: gemma-4-12b
provider: harness provider: harness
``` ```
+4 -4
View File
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
| Param | Type | Required | Default | Description | | Param | Type | Required | Default | Description |
|-------|------|----------|---------|-------------| |-------|------|----------|---------|-------------|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) | | `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) | | `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
## Maintains ## Maintains
@@ -55,7 +55,8 @@ connectivity recovery including end-to-end DM validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User | | Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------| |------|-----|---------|-------------|-------------|------|
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* | | Mumuni | CT100 | — | 192.168.68.24 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root | | Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky | | Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
@@ -120,8 +121,7 @@ cp plugins/platforms/zulip/adapter.py \
plugins/platforms/zulip/plugin.yaml \ plugins/platforms/zulip/plugin.yaml \
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/ {{hermes_home}}/hermes-agent/plugins/platforms/zulip/
# Fix ownership (was Tanko-only, runs as jerome user) # Fix ownership (Tanko only — runs as jerome user)
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \ [ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/ {{hermes_home}}/hermes-agent/plugins/platforms/zulip/
+8 -4
View File
@@ -4,6 +4,8 @@ report_only_agents:
kind: function kind: function
name: hermes-zulip-restore name: hermes-zulip-restore
description: > description: >
Restores Zulip connectivity for any Hermes agent (Mumuni CT100, Tanko CT112,
Koby CT111, Shumba on Lucky's mini PC). Deploys the zulip-platform adapter to the correct bundled plugin
path, verifies env credentials, restarts the gateway, and confirms Zulip path, verifies env credentials, restarts the gateway, and confirms Zulip
connects. Run this whenever a Hermes agent stops responding on Zulip or after connects. Run this whenever a Hermes agent stops responding on Zulip or after
a fresh agent deployment. a fresh agent deployment.
@@ -24,7 +26,7 @@ gateway restart, and connection validation.
| Param | Type | Required | Default | Description | | Param | Type | Required | Default | Description |
|-------|------|----------|---------|-------------| |-------|------|----------|---------|-------------|
| `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) | | `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
## Maintains ## Maintains
@@ -37,7 +39,7 @@ gateway restart, and connection validation.
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py` - `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path - All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko, historical — DSH since 2026-08-27) - Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko)
- Gateway restarted and zulip platform reports state `connected` - Gateway restarted and zulip platform reports state `connected`
- HTML stripping enabled for `/approve` and `/deny` slash command support - HTML stripping enabled for `/approve` and `/deny` slash command support
@@ -52,6 +54,8 @@ gateway restart, and connection validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User | | Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------| |------|-----|---------|-------------|-------------|------|
| Mumuni | CT100 (abiba) | hwepve | 192.168.68.24 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root | | Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky | | Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
@@ -93,8 +97,8 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \ zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
# Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin) # Fix ownership (Tanko only)
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical) chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only
# Clean up # Clean up
rm -rf /tmp/zulip-deploy rm -rf /tmp/zulip-deploy
+2 -2
View File
@@ -87,8 +87,8 @@ call apply-liteLLM-routing
call apply-agent-compression call apply-agent-compression
agent: mumuni agent: mumuni
host: 192.168.68.14 host: 192.168.68.24
config_path: /home/hermes/.hermes/config.yaml config_path: /root/.hermes/config.yaml
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts -- Phase 4: Enable llama.cpp prompt caching on GPU hosts
+27 -34
View File
@@ -3,7 +3,7 @@ kind: pattern
name: infrastructure-control name: infrastructure-control
description: > description: >
Full infrastructure monitoring and control pattern covering the Full infrastructure monitoring and control pattern covering the
5-node Proxmox cluster, 3 Docker ecosystems (22 containers), 6-node Proxmox cluster, 3 Docker ecosystems (22 containers),
NFS storage, and network services. Defines monitors, remediations, NFS storage, and network services. Defines monitors, remediations,
and the access matrix for all environments. and the access matrix for all environments.
@@ -13,11 +13,10 @@ description: >
against the live system. Policy fields are authoritative. See the against the live system. Policy fields are authoritative. See the
`verify-before-mutate` skill. `verify-before-mutate` skill.
**Last verified:** 2026-08-15 — hwepve removed from Tabiri cluster **Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110),
(now 5 nodes: minipve, amdpve, storepve, acerpve, ocupve). hwepve AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve),
(192.168.68.4) is a standalone PVE node + NetBird routing peer; Abiba placement (hwepve not amdpve), added hwepve as 6th node,
London relocation pending. CTs 100 (abiba) and 105 (kagentz) moved added dns.sysloggh.net route.
to minipve.
--- ---
# Infrastructure Control Pattern # Infrastructure Control Pattern
@@ -35,21 +34,21 @@ description: >
┌─────────────┐ ┌──────────────┐ ┌──────────────┐ ┌─────────────┐ ┌──────────────┐ ┌──────────────┐
│ Abiba │ │ Tanko │ │ Mumuni │ │ Abiba │ │ Tanko │ │ Mumuni │
│ (pi) │ │ (Hermes) │ │ (Hermes) │ │ (pi) │ │ (Hermes) │ │ (Hermes) │
│ CT 100 │ │ CT 112 │ │ CT 100 │ │ CT 100 │ │ CT 112 │ │ CT 114 │
└──────┬──────┘ └──────┬───────┘ └──────┬───────┘ └──────┬──────┘ └──────┬───────┘ └──────┬───────┘
│ │ │ │ │ │
└──────────────────┼────────────────────┘ └──────────────────┼────────────────────┘
▼ ▼
┌────────────────────────────────────┐ ┌──────────────────────────────────────┐
│ Proxmox Cluster API │ │ Proxmox Cluster API │
│ minipve.sysloggh.net:443 │ │ minipve.sysloggh.net:443 │
│ (monitoring@pve!mumuni token) │ │ (monitoring@pve!mumuni token) │
└────┬──────┬──────┬──────┬──────────┘ └────┬──────┬──────┬──────┬──────┬─────┘
│ │ │ │ │ │ │ │ │
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼
minipve amdpve storepve acerpve ocupve minipve amdpve storepve acerpve ocupve hwepve
(.12) (.15) (.6) (.9) (.5) (.12) (.15) (.6) (.9) (.5) (.4)
▼ ▼
┌─────────────────────────────────────────────┐ ┌─────────────────────────────────────────────┐
@@ -101,29 +100,21 @@ description: >
## Section 2: Proxmox Cluster — Monitoring ## Section 2: Proxmox Cluster — Monitoring
### Nodes (5) ### Nodes (6)
| Node | IP | CPU | RAM | VMs/CTs | Role | | Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------| |------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | abiba, kagentz, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging | | minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute | | amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat | | storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs | | acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | | ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
| hwepve | .4 | 12C | 15GB | abiba, kagentz, (mumuni CT 114 stopped) | Agents (new node) |
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on > **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
> minipve at .10, not acerpve. Abiba (CT 100) and kagentz (CT 105) are on > minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni
> minipve (moved from hwepve 2026-08-15). Mumuni runs inside Abiba CT100 > (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a
> (.24); CT 114 (mumuni) no longer exists in the cluster. > second instance on minipve at .123 — distinguish by CT ID, not hostname.
>
> **hwepve (192.168.68.4) — STANDALONE (removed from Tabiri 2026-08-15):**
> Huawei MateBook 16 (KLVL-WXX9), pve-manager/9.2.10, kernel 7.0.14-8-pve.
> Zero VMs/CTs. Being relocated to London as a standalone PVE node + NetBird
> routing peer (relocation pending). Localizations applied: timezone
> Europe/London, lid-switch ignore, sleep/suspend/hibernate targets masked,
> cluster-shared storage removed (remaining: local, local-lvm, storage,
> mediastore). prometheus-node-exporter active on :9100; net.ipv4.ip_forward=1;
> NetBird client not yet installed (enrollment pending setup key).
### Checks (every 5 min) ### Checks (every 5 min)
@@ -601,20 +592,21 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| CT | Name | Node | IP | Role | Agent | | CT | Name | Node | IP | Role | Agent |
|----|------|------|----|------|-------| |----|------|------|----|------|-------|
| 100 | abiba | minipve | .24 | Pi agent | ✅ pi | | 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi |
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ | | 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
| 102 | adguard | **minipve** | **.10** | DNS | ❌ | | 102 | adguard | **minipve** | **.10** | DNS | ❌ |
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ | | 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
| 104 | authentik | minipve | .11 | OIDC | ❌ | | 104 | authentik | minipve | .11 | OIDC | ❌ |
| 105 | kagentz | minipve | — | Agent Zero | ✅ | | 105 | kagentz | **hwepve** | — | Agent Zero | ✅ |
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP | | 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
| 107 | pbs | storepve | — | Backups | ❌ | | 107 | pbs | storepve | — | Backups | ❌ |
| 108 | media | storepve | — | Media | ❌ | | 108 | media | storepve | — | Media | ❌ |
| 109 | docker-vm | storepve | .7 | Docker host | ❌ | | 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | **.17** | Git | ❌ | | 110 | gitea | minipve | **.17** | Git | ❌ |
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ | | 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ | | 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ | | 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ | | 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ | | 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
| 117 | zulip | storepve | .19 | Chat | ❌ | | 117 | zulip | storepve | .19 | Chat | ❌ |
@@ -639,14 +631,15 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| CT | Name | Node | pct-run | | CT | Name | Node | pct-run |
|-----|------|------|---------| |-----|------|------|---------|
| 100 | abiba | minipve | `pct-run 100` | | 100 | abiba | hwepve | `pct-run 100` |
| 105 | kagentz | minipve | `pct-run 105` | | 105 | kagentz | hwepve | `pct-run 105` |
| 111 | tdunna | amdpve | `pct-run 111` | | 111 | tdunna | amdpve | `pct-run 111` |
| 112 | tanko | amdpve | `pct-run 112` | | 112 | tanko | amdpve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` | | 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` | | 115 | scottdenya | amdpve | `pct-run 115` |
| 104 | authentik | minipve | `pct-run 104` | | 104 | authentik | minipve | `pct-run 104` |
| 110 | gitea | minipve | `pct-run 110` | | 110 | gitea | minipve | `pct-run 110` |
| 114 | mumuni | hwepve | `pct-run 114` |
| 116 | syslog-api | minipve | `pct-run 116` | | 116 | syslog-api | minipve | `pct-run 116` |
| 106 | ra-h-os | storepve | `pct-run 106` | | 106 | ra-h-os | storepve | `pct-run 106` |
| 107 | proxmox-backup | storepve | `pct-run 107` | | 107 | proxmox-backup | storepve | `pct-run 107` |
@@ -659,7 +652,7 @@ GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use S
ssh root@192.168.68.8 # RTX 3090 ssh root@192.168.68.8 # RTX 3090
ssh root@192.168.68.110 # RTX 5070 ssh root@192.168.68.110 # RTX 5070
ssh root@192.168.68.15 # Strix Halo ssh root@192.168.68.15 # Strix Halo
ssh root@192.168.68.4 # hwepve — standalone London node + NetBird routing peer (relocation pending) ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni)
``` ```
## Section 7: Agent Health Check (consolidated — 2026-07-05) ## Section 7: Agent Health Check (consolidated — 2026-07-05)
+1
View File
@@ -140,6 +140,7 @@ After ALL updates (apt + images + restarts), verify every critical service is ba
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK | | Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK | | Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` | | PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
| Hermes gateways | SSH to Mumuni CT 100, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here. Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
+47 -9
View File
@@ -7,15 +7,13 @@ description: >
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
via existing /metrics Prometheus endpoint. via existing /metrics Prometheus endpoint.
DEPLOYMENT STATUS (2026-08-09): DEPLOYMENT STATUS (2026-07-09):
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
(via proxmox-monitor contract). Grafana at :3001, all scrape targets active. (via proxmox-monitor contract). Grafana at :3001, 5 scrape targets active.
✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on ❌ GPU exporters NOT deployed: gpu-exporter crash-loops on .15,
:9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. NVIDIA sidecar exporters (.8/.110:9400) never installed.
✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via Router falls back to direct GPU /health probes.
master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. ⚠️ This contract is target-state aspirational — not as-built.
⚠️ This contract is target-state aspirational — but GPU export + alerting
are now as-built (verified 2026-08-09).
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
version: 1.0.0 version: 1.0.0
--- ---
@@ -102,6 +100,47 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
- Stack persists across reboots (systemd for exporters, Docker restart policy) - Stack persists across reboots (systemd for exporters, Docker restart policy)
## Execution ## Execution
### Phase 1: GPU Exporters
**NVIDIA (.8 and .110)**:
1. Download `nvidia_gpu_exporter` binary
2. Create systemd service `nvidia-gpu-exporter.service`
3. Start and enable
**AMD (.15)**:
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
2. Parses `amdgpu_top --json -d 1000` output
3. Exposes key metrics at `:9400/metrics` via Python http.server
4. Create systemd service
5. Start and enable
### Phase 2: Prometheus
1. Create `/opt/monitoring/` directory on CT 116
2. Write `prometheus.yml` with scrape configs for all targets
3. Add to docker-compose (or separate compose file)
4. Start container
### Phase 3: Grafana
1. Create `/opt/monitoring/grafana/` directories
2. Provision Prometheus datasource
3. Provision GPU fleet dashboard JSON
4. Provision LiteLLM dashboard JSON
5. Add to docker-compose
6. Start container
### Phase 4: Verification
1. Verify all 3 GPU exporters return 200 at :9400/metrics
2. Verify Prometheus targets all UP at :9090/targets
3. Verify Grafana accessible at :3001 with dashboards
4. Verify LiteLLM metrics flowing to Prometheus
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
## Execution
### check-health ### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
@@ -135,7 +174,6 @@ curl -s http://192.168.68.116:4001/metrics | head -20
**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert. **Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert.
### Phase 1: GPU Exporters ### Phase 1: GPU Exporters
**NVIDIA (.8 and .110)**: **NVIDIA (.8 and .110)**:
+8 -7
View File
@@ -2,7 +2,7 @@
kind: responsibility kind: responsibility
name: infrastructure-update name: infrastructure-update
description: > description: >
Autonomous system-wide update contract covering all 5 Proxmox nodes, Autonomous system-wide update contract covering all 6 Proxmox nodes,
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages, 15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
Docker images, and container stacks in safe waves with health checks Docker images, and container stacks in safe waves with health checks
and automatic rollback on failure. and automatic rollback on failure.
@@ -56,10 +56,11 @@ Before ANY update wave:
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min | | CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min | | CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min | | CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 105 (kagentz, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min | | CT 100 (mumuni/abiba, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min | | VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min | | VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
@@ -162,9 +163,9 @@ Before Wave 1, snapshot these files:
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe) /etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110) /etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
# Hermes agent configs (key enforcement — 2026-07-10) # Hermes agent configs (key enforcement — 2026-07-10)
/home/hermes/.hermes/config.yaml (Mumuni kagentz CT105; Tanko CT112 uses /home/jerome/.hermes) /root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
/etc/systemd/system/hermes-gateway.service (Mumuni kagentz CT105 — system unit, User=hermes) /root/.config/systemd/user/hermes-gateway.service (Mumuni CT 114 — EnvironmentFile fixed)
/etc/environment (LITELLM_API_KEY — legacy path, Mumuni now keys via Infisical) /etc/environment (Mumuni CT 114 — LITELLM_API_KEY)
``` ```
## MCP Gateway (2026-07-10) ## MCP Gateway (2026-07-10)
@@ -218,7 +219,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
## Success Criteria ## Success Criteria
- [ ] All 5 PVE nodes updated, no reboot-loop - [ ] All 6 PVE nodes updated, no reboot-loop
- [ ] All VMs/CTs running post-update - [ ] All VMs/CTs running post-update
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117) - [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
- [ ] LiteLLM inference passing (syslog-auto test) - [ ] LiteLLM inference passing (syslog-auto test)
@@ -234,7 +235,7 @@ After completion, send Zulip DM:
``` ```
📋 Infrastructure Update — YYYY-MM-DD 📋 Infrastructure Update — YYYY-MM-DD
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers
Security fixes: N CVEs patched Security fixes: N CVEs patched
Downtime: <service> <duration> Downtime: <service> <duration>
Failures: none / <details> Failures: none / <details>
+4 -46
View File
@@ -178,7 +178,7 @@ through its agent wrapper.
| Agent | Host | Pattern | Keys | Status | | Agent | Host | Pattern | Keys | Status |
|-------|------|---------|------|--------| |-------|------|---------|------|--------|
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed | | abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
| mumuni | .14 (kagentz CT105) | systemd unit hermes-gateway.service (user hermes) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback | | mumuni | .24 (CT100 abiba) | Pi Hermes gateway (no systemd) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback | | tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed | | koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed | | koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
@@ -257,51 +257,9 @@ reads use per-agent identities. This eliminates the single shared token risk.
| Agent | .env Keys | | Agent | .env Keys |
|-------|-----------| |-------|-----------|
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY | | Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
|| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY | | Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
|| Koby | (wrapper injects from vault — .env has Telegram token) | | Koby | (wrapper injects from vault — .env has Telegram token) |
|| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY | | Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
|| Agent Zero (kagentz .14) | OPENROUTER_API_KEY (direct OpenRouter access) |
### Agent Zero (kagentz .14) — OpenRouter Integration (2026-09-01)
Agent Zero runs in Docker on kagentz (CT105) and uses **direct OpenRouter API access**,
not via the LiteLLM proxy. This is because Agent Zero's workflow (self-update manager,
UI bootstrap, model selection) is built around OpenRouter's native authentication.
**Key Storage:**
- **Container**: `/a0/usr/.env` (line ~72: `API_KEY_OPENROUTER=sk-or-v1-…`)
- **Vault**: Infisical secret `OPENROUTER_API_KEY` (project=agents, env=production)
- **Fallback**: The container's .env is the primary source; vault sync is optional
(unlike fleet agents which require vault injection)
**Current Key (2026-09-01):**
- **Prefix**: `sk-or-v1-0af3f3…`
- **User**: `user_2rt9lCqcd5d7Vk1t18DHsvWdPTT`
- **Plan**: Paid (not free tier)
- **Usage**: 0 (as of 2026-09-01)
**Model Configuration:**
- **Preset**: "Cost Efficient" (`/a0/usr/plugins/_model_config/presets.yaml`)
- **Model**: `openrouter/moonshotai/kimi-k3`
- **API Base**: (empty — uses OpenRouter default)
**Why not LiteLLM proxy?**
Agent Zero's architecture was designed before the fleet adopted the LiteLLM proxy
standard. The container runs `/exe/self_update_manager.py` and `/a0/run_ui.py` which
directly call OpenRouter via Python's requests library. Converting would require:
1. Refactoring all LLM calls to use `litellm` library
2. Adding vault wrapper injection
3. Updating self_update_manager to use proxy-aware key handling
**Rotation Procedure:**
1. Generate new key in OpenRouter UI
2. Update container: `sed -i 's/^API_KEY_OPENROUTER=.*/API_KEY_OPENROUTER=<new_key>/' /a0/usr/.env`
3. Update vault: `infisical secrets set OPENROUTER_API_KEY=<new_key> --projectId=agents --env=production`
4. Restart container: `sudo docker exec agent-zero supervisorctl restart run_ui`
5. Verify: `curl -s https://openrouter.ai/api/v1/auth/key -H "Authorization: Bearer <new_key>"`
**Related Contract:**
- `agent-zero-openrouter-key.prose.md` — Full agent-zero key management contract
## Key Rotation Log ## Key Rotation Log
-117
View File
@@ -1,117 +0,0 @@
---
kind: pattern
name: litellm-client-timeouts
description: >
Standard client timeout and retry policy for ALL agents calling LiteLLM
(CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident:
a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures
(82% from abiba-pi, 13 from mumuni) against a backend that was actually
succeeding at 20-70s per call once clients stopped giving up. Grounded in
measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes
to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h;
nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely
client-side. Blast radius if wrong: agents fall back to DeepSeek silently
(key/timeout failures present as model degradation, not errors) or abandon
healthy-but-slow reasoning calls, fragmenting long tasks.
---
## Maintains
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
## The measured numbers these values come from
| Model | avg latency | avg TTFT | p-profile (24h) |
|---|---|---|---|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
| strix-moe | 7.5s | — | Strix Halo, healthy |
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
backend), full recovery 07:00-08:00 with ZERO client failures once requests
tolerated 20-70s — 392 successful slow calls in the three hours after recovery.
LiteLLM's internal queue time is ~0s; the latency is model inference, not
proxy queuing.
## Parameters
### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy
28.8s average + 120-300s tail. Every 408 in the incident was a client
abandoning a request the backend would have answered.
- If the transport exposes a timeout setting for the main model, set it to
**300s or more**. If it does not (current Hermes custom-provider path has no
timeout knob), that is acceptable ONLY because nginx holds the request for
600s — but any wrapper, script, or direct API call you write MUST set its own
timeout >= 300s for syslog-auto/qwen-class calls.
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast
here is what caused the incident.
### 2. Auxiliary tasks — keep template timeouts, one correction
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
these are fine.
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
syslog-auto; delegation defaults that assume fast responses will 408 the
same way.
### 3. Retry policy — backoff, not repetition
- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff
(**15s, 45s**) before giving up.
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour
from one key) was a batch job retrying without backoff while the backend was
down — it multiplied load during recovery.
- On 401/403: do NOT retry — that is a key/permission problem (see
litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
- On 429: honor the retry-after header if present, else back off 60s.
### 4. Health probes — identify yourself and time out sanely
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12
such orphans appeared in the incident window and cost investigation time).
Send a real model name and use a real (probe-designated) key.
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
report "backend slow (>30s)" rather than hanging.
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
standard; sub-hourly synthetic traffic distorts latency baselines.
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
- The incident window showed bulk clients amplifying a backend stall 5:1.
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
- Any batch loop over N requests MUST sleep >= 5s between requests and honor
the retry policy in section 3.
## Returns
- A single standard any agent or script can cite: timeouts >= 300s on the
syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x,
15s/45s), identified probes at 30s/hourly, batch jobs chunked and
day-scheduled.
- Failure signature recognition: bulk 408s from multiple keys in one window =
backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated);
single-key 408s = that client's timeout is too short.
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s,
auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent,
stable aliases), litellm-api-keys.prose.md (key/permission failures).
## Intentionally NOT changed
- No server-side LiteLLM timeout/cooldown changes proposed — the incident
self-recovered and the server is healthy (0.56s live probe); changing
server behavior without process-level root cause (CT116 requires root;
not reachable from kagentz) would be guessing.
- No change to the template's vision/web_extract/compression timeouts —
measured data says they are correct.
- No per-agent key permission changes — those are litellm-api-keys.prose.md
territory (and the open gpu-vision/gemma 403 items are already filed with
the key owners).
- No model routing changes — syslog-auto's weighted pool behaved correctly
throughout the incident.
-9
View File
@@ -18,15 +18,6 @@ description: >
Designed as a reusable contract for any Syslog agent. Designed as a reusable contract for any Syslog agent.
Source of truth: gpu-fleet.prose.md Source of truth: gpu-fleet.prose.md
## Monitoring / Alerting (as-built 2026-08-09)
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
--- ---
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
+5 -3
View File
@@ -3,7 +3,7 @@ report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129 - koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
name: memory-audit-maintenance name: memory-audit-maintenance
kind: responsibility kind: responsibility
description: Shared memory audit and maintenance contract for Hermes agents (Mumuni, Koby, Koonimo). Tanko is no longer a Hermes agent (now on DSH/DeepSeek Harness since 2026-08-27) and uses DSH-native memory, so it is excluded from this Hermes roster. Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens. description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
id: 067NC4KG01RG50R40M30E20918 id: 067NC4KG01RG50R40M30E20918
--- ---
--- ---
@@ -18,10 +18,11 @@ Autonomously audit and reorganize an agent's native memory (MEMORY.md, USER.md,
### Scope ### Scope
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across Hermes agents (Mumuni, Koby, Koonimo). **Tanko is excluded — it migrated to DSH (DeepSeek Harness) on 2026-08-27 and now uses DSH-native memory (mnemon), not `~/.hermes/memories/`.** Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data. This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
**Agent Roster (Hermes):** **Agent Roster:**
- Mumuni - Mumuni
- Tanko
- Koby (CT 111 / tdunna) - Koby (CT 111 / tdunna)
- Koonimo (CT 113 / baggy) - Koonimo (CT 113 / baggy)
@@ -345,3 +346,4 @@ return {
### Per-Agent Notes ### Per-Agent Notes
Each Hermes agent (Mumuni, Tanko, Tdunna, Baggy) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory.
+44 -150
View File
@@ -4,176 +4,70 @@ report_only_agents:
kind: pattern kind: pattern
name: memory-fixer name: memory-fixer
description: > description: >
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations. Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations.
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at). Escalate anything that needs Kwame's input.
version: 2.0.0 version: 1.1.0
--- ---
--- ---
# Memory Fixer # Memory Fixer
> **Canonical copy:** `/root/.hermes/contracts/memory-fixer-v3.md` (used by the `memory-fixer-daily` cron job). This file is the institutional record of the same contract. When the two diverge, treat the v3 source in `/root/.hermes/contracts/` as executable truth.
## Purpose ## Purpose
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever. Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input.
**Type:** Write-only (Level 1 fixes only) ## Level 0 Auto-Deletes (Allowed Without Approval)
**Scope:** RA-H OS knowledge graph (192.168.68.65) Ephemeral heartbeat and log nodes that violate "Logs NEVER go in the graph":
**Schedule:** Daily at 8 AM ET
**Escalation:** Level 2+ to Kwame as task items
## Key Design Decision - `[LITELLM-HEALTH]`, `[GPU-SELF-HEAL]`, `[PM2-SELF-HEAL]`
- `[PROXMOX-MONITOR]`, `[GPU-MONITOR]`, `[INFRA-MONITOR]`, `[AGENT-HEALTH]`, `[DISK-GC]`
- `[WAL]` entries older than 30 days
The `updateNode` tool's `metadata` field performs a **restricted merge** — the `state` key only accepts `'processed'` or `'not_processed'`. Additionally, new metadata keys cannot be added via the merge. **Condition:** node must be an orphan (no edges). Deleting a connected node risks breaking other nodes.
**Solution:** Use the `description` field to tag stale nodes with review actions, since `description` is a simple string overwritable via `updateNode`. **Method:** direct SQLite on `.65` (MCP has no delete tool):
```bash
**Tag Format:** `[REVIEW: action] original description text...` ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"
DELETE FROM nodes WHERE id IN (
Where `action` is one of: SELECT id FROM nodes WHERE id NOT IN (SELECT from_node_id FROM edges)
- `archive` — node is stale and should be archived AND id NOT IN (SELECT to_node_id FROM edges)
- `refresh` — node is stale and should be refreshed (infrastructure) AND title LIKE '[LITELLM-HEALTH]%' -- add more prefixes as needed
- `keep` — node has been confirmed as current );\""
- `merge` — node is a duplicate candidate
**Query for finding review-tagged nodes:**
```sql
SELECT id, title, description
FROM nodes
WHERE description LIKE '[REVIEW:%';
``` ```
## Level 1 Auto-Fixes (No Kwame Decision Needed) ## Level 1 Auto-Fixes (No Judgment Required)
### 1. Missing `type` Auto-Classification ### 1. Missing `type` Field
For nodes with content but no `metadata.type`:
- Title contains "Proxmox" or "infrastructure" → `type: infrastructure`
- Title contains "skill" or "how to" or "guide" → `type: skill`
- Title contains "doc" or "template" or "brand" → `type: documentation`
- Title starts with "WAL:" or "TASK:" → `type: note`
- Title starts with "[LEARN]" → `type: documentation`
- Otherwise → `type: note` (default)
### 2. Missing `tenant` / `namespace`
For any node with NULL tenant or namespace:
```sql ```sql
SELECT id, title, UPDATE nodes
CASE SET metadata = json_set(
WHEN title LIKE '%infrastructure%' OR title LIKE '%proxmox%' OR title LIKE '%setup%' THEN 'infrastructure' COALESCE(metadata, '{}'),
WHEN title LIKE '%skill%' OR title LIKE '%how to%' OR title LIKE '%guide%' THEN 'skill' '$.tenant', 'syslogsolution',
WHEN title LIKE '%doc%' OR title LIKE '%template%' OR title LIKE '%brand%' THEN 'documentation' '$.namespace', 'syslogsolution'
WHEN title LIKE 'WAL:%' OR title LIKE 'TASK:%' THEN 'note'
WHEN title LIKE '%[LEARN]%' THEN 'documentation'
ELSE 'note'
END as auto_type
FROM nodes
WHERE json_extract(metadata, '$.type') IS NULL;
```
### 2. Missing `namespace` Auto-Population
```sql
SELECT id, title, json_extract(metadata, '$.tenant') as tenant
FROM nodes
WHERE json_extract(metadata, '$.namespace') IS NULL;
```
### 3. Staleness Review Tagging
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
**Exclusion Rules:**
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
- Nodes whose `description` already starts with `[REVIEW:` are NOT re-processed
```sql
SELECT id, title, json_extract(metadata, '$.type') as node_type,
CAST(julianday('now') - julianday(updated_at) AS INTEGER) as days_stale,
CASE
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN 'refresh'
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN 'refresh'
ELSE 'archive'
END as suggested_action
FROM nodes
WHERE updated_at < datetime('now',
CASE
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN '-14 days'
WHEN json_extract(metadata, '$.type') IN ('skill', 'documentation', 'template', 'protocol-enforcement', 'prd', 'architecture') THEN '-90 days'
WHEN json_extract(metadata, '$.type') IN ('note', 'wal', 'WAL', 'task', 'TASK', 'event') THEN '-30 days'
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN '-120 days'
WHEN json_extract(metadata, '$.type') IN ('deprecated-relay', 'audit', 'audit-report', 'incident', 'incident-report') THEN '-3650 days'
ELSE '-45 days'
END
) )
AND json_extract(metadata, '$.state') NOT IN ('review_pending', 'deprecated', 'archived', 'not_processed') WHERE json_extract(metadata, '$.tenant') IS NULL
AND (description IS NULL OR description NOT LIKE '[REVIEW:%') OR json_extract(metadata, '$.namespace') IS NULL;
ORDER BY days_stale ASC
LIMIT 10;
``` ```
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`. ### 3. Staleness State Transitions
Using the type-based windows from the memory-monitor contract:
- Nodes stale > their window → transition to `state: review_pending`
- Nodes in `review_pending` for >7 days → escalate to Kwame (Level 2)
## Level 2 Escalations (Kwame Decision Required) ## Level 2 Escalations (Kwame Decision Required)
1. **Nodes in `review_pending` >7 days** — Archive, refresh, or keep?
1. **Stale nodes** flagged with `[REVIEW: …]` — Archive, refresh, or keep? 2. **Orphan Nodes >90 days old** — Delete or Connect?
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep? 3. **Potential Duplicate Nodes** — Same title or >70% overlap. Merge or Keep?
3. **Orphan Nodes >90 days old** — Archive or connect? 4. **Conflicting Metadata** — Content suggests one tenant but metadata says another.
## Reporting Format
The fixer reports to Kwame via this Zulip DM:
```
🦅 Memory Fixer — [HH:MM UTC]
Level 1 fixes applied:
- Missing type: X nodes classified
- Missing namespace: Y nodes populated
Stale nodes needing review (max 10):
1. [Node #XXX] Title — X days stale, SUGGEST: refresh
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
...
Duplicates needing decision:
1. [Node #AAA] vs [Node #BBB] — Same title
Orphans >90 days:
1. [Node #EEE] Title — X days stale, orphaned
Reply with:
- "archive #XXX, #YYY" to mark for archive
- "archive all" to archive all stale nodes listed
- "keep #XXX" to confirm a node is current
- "merge #AAA into #BBB" to merge duplicates
- "refresh #XXX" to mark as current
```
## Execution on Next Run
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
> ⚠️ `updateNode` cannot set `state` to non-standard values (restricted to `processed`/`not_processed`) and cannot add metadata keys. For state transitions and `updated_at` bumps, use **direct SSH + SQLite** on the bridge host:
> ```bash
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
> ```
> Use `updateNode` only for description/source/title/link edits.
Decision → completed action mapping:
| Kwame reply | Description change | State | `updated_at` |
|---|---|---|---|
| `archive #XXX` | replace `[REVIEW: archive] ` → `[ARCHIVED] ` prefix | `archived` | bumped to now |
| `keep #XXX` / `refresh #XXX` | **clear the `[REVIEW: …]` tag entirely** | `active` | bumped to now |
| `merge #AAA into #BBB` | set `[REVIEW: merge_into #BBB]` on #AAA, then follow manual merge workflow | handled manually | bumped to now |
| `archive all` | apply the archive row to every node listed in the prior report | `archived` | bumped to now |
**Why `updated_at` must be bumped (critical):** the Level-1 staleness query keys off `updated_at < now - window`. If the fixer clears the tag but leaves a stale `updated_at`, the node is immediately re-flagged on the very next run and the cycle repeats forever. Bumping `updated_at` to now pushes the node back to the front of the window.
**Exclusion after action:** once an action is applied, the node's description no longer starts with `[REVIEW:` (archive → `[ARCHIVED]`, refresh/keep → original text), so it is not re-processed.
After all actions are applied, verify with:
```sql
SELECT id, json_extract(metadata, '$.state') FROM nodes WHERE description LIKE '[REVIEW:%';
```
The result must be 0 rows when all decisions are executed. Report what was done.
## Checks
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
- **Timestamps:** every executed decision bumps `updated_at`, so the node exits the stale window on the next run.
## Logging ## Logging
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md` Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
+7 -7
View File
@@ -6,7 +6,7 @@ description: >
delegation, verification, and delivery. Defines when to delegate, which delegation, verification, and delivery. Defines when to delegate, which
worker to use for what, how to handle failures, and the kanban board worker to use for what, how to handle failures, and the kanban board
protocol. Enforces context-window discipline and separation of concerns. protocol. Enforces context-window discipline and separation of concerns.
Runs on Mumuni (kagentz CT105, minipve, .14) via Hermes agent (Zulip gateway via systemd). Runs on Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent (Pi + Hermes Zulip gateway).
version: 1.0.0 version: 1.0.0
--- ---
@@ -19,8 +19,8 @@ version: 1.0.0
## Topology ## Topology
**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve) **Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve)
**Manager:** Mumuni (kagentz CT105, minipve, .14) via Hermes agent **Manager:** Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed **Workers:** 6 profiles, all running on the same agent — no separate hosts needed
This contract is infrastructure-agnostic in terms of which nodes are used. This contract is infrastructure-agnostic in terms of which nodes are used.
@@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6,
## Why This Matters ## Why This Matters
Without enforced delegation, the manager consumes the full iteration budget Without enforced delegation, the manager consumes the full iteration budget
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs — (60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs —
leaving no capacity for actual coordination. The result: context overflow leaving no capacity for actual coordination. The result: context overflow
(59K tokens in system prompt), iteration exhaustion, and degraded response (59K tokens in system prompt), iteration exhaustion, and degraded response
quality. This contract exists because I blew through my budget checking quality. This contract exists because I blew through my budget checking
@@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own.
**This is a hard rule, not a recommendation.** Violating it produces the exact **This is a hard rule, not a recommendation.** Violating it produces the exact
type of discrepancy the kanban pipeline exists to prevent: a review worker finds type of discrepancy the kanban pipeline exists to prevent: a review worker finds
"5 nodes present" in the raw data but "5/5 online" in the report — even though "6 nodes present" in the raw data but "6/6 online" in the report — even though
one of those nodes was unreachable. The report lied because it used data the one of those nodes was unreachable. The report lied because it used data the
raw data never provided. raw data never provided.
@@ -137,7 +137,7 @@ delegate_task(
``` ```
delegate_task( delegate_task(
tasks=[ tasks=[
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"}, {"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"}, {"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
] ]
) )
@@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel:
{ {
"lane_id": "devops-check", "lane_id": "devops-check",
"worker": "syslog-devops", "worker": "syslog-devops",
"goal": "Check all 5 Proxmox nodes", "goal": "Check all 6 Proxmox nodes",
"status": "dispatched|completed|failed", "status": "dispatched|completed|failed",
"output_file": "/tmp/node-report.md" "output_file": "/tmp/node-report.md"
} }
+1 -1
View File
@@ -54,7 +54,7 @@ garbled input. When a user types these commands to Abiba:
- **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately." - **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately."
- **`/approve session`** → Same response - **`/approve session`** → Same response
- **`/deny`** → Same response + "For agents (Mumuni on Hermes, Tanko on DSH), these work with their built-in approval system." - **`/deny`** → Same response + "For Hermes agents (Tanko, Mumuni), these work with their built-in approval system."
This keeps the UX consistent across agents — users can type `/approve` anywhere This keeps the UX consistent across agents — users can type `/approve` anywhere
without getting confused by LLM responses. without getting confused by LLM responses.
+10 -9
View File
@@ -2,17 +2,23 @@
kind: responsibility kind: responsibility
name: pm2-self-heal name: pm2-self-heal
description: > description: >
Monitors critical PM2 processes (abiba-zulip, abiba-telegram) and
auto-restarts any that are stopped or errored. Logs every action to
the knowledge graph and alerts the owner via Zulip DM on failures.
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
report_only_agents:
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
--- ---
## Maintains ## Maintains
- abiba-telegram: { status: "online", uptime: string, restarts: number } - abiba-telegram: { status: "online", uptime: string, restarts: number }
- abiba-zulip: { status: "online", uptime: string, restarts: number } - abiba-zulip: { status: "online", uptime: string, restarts: number }
- gpu-monitor: { status: "online", uptime: string, restarts: number } (systemd-managed, PM2-tracked)
- gitea-runner: { status: "online", uptime: string, restarts: number } - gitea-runner: { status: "online", uptime: string, restarts: number }
- spoton-service: { status: "online", uptime: string, restarts: number }
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp - last_check: timestamp
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd (stable); PM2 tracks it for status reporting only. `gpu-watchdog` retired from PM2.
## Continuity ## Continuity
@@ -27,15 +33,10 @@ description: >
- **Verify**: Re-check status after 5 seconds - **Verify**: Re-check status after 5 seconds
- **Escalate**: If still failed after 2 retries, send Zulip DM to owner - **Escalate**: If still failed after 2 retries, send Zulip DM to owner
### Rule 2: Process Restarting Too Often (crash-loop guard) ### Rule 2: Process Restarting Too Often
- **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter) **or** a - **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter)
process reporting a restart count > 1000 while showing "online" (a crash-loop mask)
- **Note**: PM2 counter never decrements; only full delete+re-add resets it - **Note**: PM2 counter never decrements; only full delete+re-add resets it
- **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>` - **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>`
- **Script guard (2026-08-16)**: `scripts/pm2-self-heal.sh` now restarts `abiba-telegram`
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
spoton incident). Alerts include the restart count.
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM - **Escalate**: Only when restarts > 30 — alerts to Zulip DM
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's - **Historical fix**: Previous cycles were caused by abiba-zulip extension's
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2). stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
+7 -6
View File
@@ -5,7 +5,7 @@ description: >
Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack
on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three
exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single
instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter instance), node_exporter (all 6 PVE nodes), and a custom docker-stats-exporter
(Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve (Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve
the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx). the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx).
agent: abiba agent: abiba
@@ -33,8 +33,8 @@ agent: abiba
| Exporter | Host:Port | Scope | Notes | | Exporter | Host:Port | Scope | Notes |
|----------|-----------|-------|-------| |----------|-----------|-------|-------|
| prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). | | prometheus-pve-exporter | .116:9221 (container) | All 6 nodes + guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
| node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. | | node_exporter | .5/.6/.9/.12/.15/.4:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 6 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. hwepve (.4) added 2026-07-19. |
| docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. | | docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. |
## PVE API Token ## PVE API Token
@@ -48,7 +48,7 @@ agent: abiba
| UID | Title | Panels | Source | | UID | Title | Panels | Source |
|-----|-------|--------|--------| |-----|-------|--------|--------|
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries | | proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 6-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
| proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) | | proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) |
| docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) | | docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) |
| gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) | | gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) |
@@ -77,9 +77,9 @@ agent: abiba
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator | | `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
| `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions | | `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions |
| `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning | | `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning |
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config | | `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15/.4 | node_exporter collector config |
## Cluster "Tabiri" — 5 Nodes ## Cluster "Tabiri" — 6 Nodes
| Node | IP | Role | | Node | IP | Role |
|------|----|----| |------|----|----|
@@ -88,6 +88,7 @@ agent: abiba
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE | | minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) | | amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts abiba (lxc/100), kagentz (lxc/105), mumuni (lxc/114). CTs 100/105 migrated from amdpve, CT 114 from minipve 2026-07-20 |
## Operations ## Operations
+5 -22
View File
@@ -30,6 +30,7 @@ INFISICAL_ENV = "prod"
# PVE node IPs for CT liveness checks # PVE node IPs for CT liveness checks
PVE_NODES = { PVE_NODES = {
"hwepve": "192.168.68.4",
"amdpve": "192.168.68.15", "amdpve": "192.168.68.15",
"minipve": "192.168.68.12", "minipve": "192.168.68.12",
"storepve": "192.168.68.6", "storepve": "192.168.68.6",
@@ -39,10 +40,10 @@ PVE_NODES = {
# Agent definitions: ct, host, user, pve_node, vault_key_name # Agent definitions: ct, host, user, pve_node, vault_key_name
AGENTS = { AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"}, "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "report_only": False},
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None, "report_only": False},
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"}, "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY", "report_only": True}, # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17)
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"}, "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY", "report_only": False},
} }
GPU_HOSTS = { GPU_HOSTS = {
@@ -240,16 +241,6 @@ def check_agents():
ct = agent["ct"] ct = agent["ct"]
report_only = agent.get("report_only", False) report_only = agent.get("report_only", False)
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
if agent.get("runtime") == "dsh":
live = ssh(host, "true", user=user)
print(f" {'✅' if live is not None else '❌'} {name}: DSH (DeepSeek Harness) — "
f"no Hermes gateway since 2026-08-27 (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
if live is None:
FAIL.append(f"unreachable:{name}")
continue
if not host or not user: if not host or not user:
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check") print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
continue continue
@@ -343,10 +334,6 @@ def check_ct_liveness():
def check_config_integrity(): def check_config_integrity():
"""Verify agent config.yaml parses as valid YAML.""" """Verify agent config.yaml parses as valid YAML."""
for name, agent in AGENTS.items(): for name, agent in AGENTS.items():
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
if agent.get("runtime") == "dsh":
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
continue
host = agent.get("host") host = agent.get("host")
user = agent.get("user") user = agent.get("user")
if not host or not user: if not host or not user:
@@ -376,10 +363,6 @@ def check_config_integrity():
def check_wrapper_integrity(): def check_wrapper_integrity():
"""Verify the hermes CLI wrapper exists and can reach hermes-real.""" """Verify the hermes CLI wrapper exists and can reach hermes-real."""
for name, agent in AGENTS.items(): for name, agent in AGENTS.items():
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
if agent.get("runtime") == "dsh":
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
continue
host = agent.get("host") host = agent.get("host")
user = agent.get("user") user = agent.get("user")
if not host or not user: if not host or not user:
+16 -11
View File
@@ -296,16 +296,21 @@ def collect():
"pm2_uptime": pm2.get("uptime", "?"), "pm2_uptime": pm2.get("uptime", "?"),
} }
# Tanko (CT 112, IP 192.168.68.122) — DSH (DeepSeek Harness), no Hermes gateway # Tanko (CT 122)
# since 2026-08-27. There is no ~/.hermes/gateway_state.json on CT 112 anymore; tanko_state = ssh_jerome("192.168.68.122", "cat ~/.hermes/gateway_state.json 2>/dev/null")
# Zulip/Telegram connectivity is managed by the DSH harness, not the Hermes gateway. tanko_data = {}
try:
tanko_data = json.loads(tanko_state) if tanko_state else {}
except:
tanko_data = {}
platforms = tanko_data.get("platforms", {})
report["agents"]["tanko"] = { report["agents"]["tanko"] = {
"platform": "dsh", "ct": 112, "ip": "192.168.68.122", "platform": "hermes", "ct": 112, "ip": "192.168.68.122",
"gateway_state": "n/a (DSH)", "gateway_state": tanko_data.get("gateway_state", "unknown"),
"zulip_state": "unknown", "zulip_state": platforms.get("zulip", {}).get("state", "unknown"),
"telegram_state": "unknown", "telegram_state": platforms.get("telegram", {}).get("state", "unknown"),
"gateway_pid": None, "gateway_pid": tanko_data.get("pid"),
"updated_at": "", "updated_at": tanko_data.get("updated_at"),
} }
# Mumuni (CT 100, IP 192.168.68.24) # Mumuni (CT 100, IP 192.168.68.24)
@@ -317,7 +322,7 @@ def collect():
mumuni_data = {} mumuni_data = {}
mumuni_platforms = mumuni_data.get("platforms", {}) mumuni_platforms = mumuni_data.get("platforms", {})
report["agents"]["mumuni"] = { report["agents"]["mumuni"] = {
"platform": "hermes", "ct": 100, "ip": "192.168.68.24", "platform": "hermes", "ct": 114, "ip": "192.168.68.24",
"gateway_state": mumuni_data.get("gateway_state", "unknown"), "gateway_state": mumuni_data.get("gateway_state", "unknown"),
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"), "telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"), "zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
@@ -538,7 +543,7 @@ th {{ color: #8b949e; font-weight: normal; }}
elif name == "tanko": elif name == "tanko":
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜") zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
gateway = agent.get("gateway_state", "?") gateway = agent.get("gateway_state", "?")
processed = "DSH" processed = agent.get("updated_at", "")[:10]
elif name == "mumuni": elif name == "mumuni":
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜") zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
gateway = agent.get("gateway_state", "?") gateway = agent.get("gateway_state", "?")
+6 -3
View File
@@ -28,8 +28,10 @@ declare -A CT_NODES=(
[117]=storepve # zulip [117]=storepve # zulip
[118]=storepve # jdownloader [118]=storepve # jdownloader
# acerpve (192.168.68.9) — no CTs (bare metal GPU .8) # acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
[100]=minipve # abiba (was hwepve) # hwepve (192.168.68.4)
[105]=minipve # kagentz (was hwepve) [100]=hwepve # abiba (was amdpve)
[105]=hwepve # kagentz (was amdpve)
[114]=hwepve # mumuni (was minipve)
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110) # ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
# #
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs): # REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
@@ -48,6 +50,7 @@ declare -A NODE_IPS=(
[storepve]=192.168.68.6 [storepve]=192.168.68.6
[acerpve]=192.168.68.9 [acerpve]=192.168.68.9
[ocupve]=192.168.68.5 [ocupve]=192.168.68.5
[hwepve]=192.168.68.4
) )
resolve_node() { resolve_node() {
@@ -74,7 +77,7 @@ main() {
if [[ $# -lt 1 ]]; then if [[ $# -lt 1 ]]; then
echo "Usage: pct-run <CT_ID> [command...]" >&2 echo "Usage: pct-run <CT_ID> [command...]" >&2
echo " pct-run 112 cat /etc/hostname" >&2 echo " pct-run 112 cat /etc/hostname" >&2
echo " pct-run 100 systemctl status hermes-gateway" >&2 echo " pct-run 114 systemctl status hermes-gateway" >&2
echo "" echo ""
echo "Known CTs:" >&2 echo "Known CTs:" >&2
for ct in $(echo "${!CT_NODES[@]}" | tr ' ' '\n' | sort -n); do for ct in $(echo "${!CT_NODES[@]}" | tr ' ' '\n' | sort -n); do
+1 -1
View File
@@ -33,7 +33,7 @@ TEL_LINE=$(echo "$STATUS" | grep "abiba-telegram")
TEL_STATUS=$(echo "$TEL_LINE" | awk -F'│' '{print $10}' | xargs) TEL_STATUS=$(echo "$TEL_LINE" | awk -F'│' '{print $10}' | xargs)
TEL_RESTARTS=$(echo "$TEL_LINE" | awk -F'│' '{print $9}' | xargs) TEL_RESTARTS=$(echo "$TEL_LINE" | awk -F'│' '{print $9}' | xargs)
if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then if [ "$TEL_STATUS" != "online" ]; then
pm2 restart abiba-telegram > /dev/null 2>&1 pm2 restart abiba-telegram > /dev/null 2>&1
sleep 3 sleep 3
TEL_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-telegram") TEL_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-telegram")
+5 -4
View File
@@ -46,17 +46,18 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology: The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (5 nodes):** **Proxmox Cluster "Tabiri" (6 nodes):**
- amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya - amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
- minipve (192.168.68.12): abiba, kagentz, adguard, authentik, gitea, syslog-api, infisical-vault - minipve (192.168.68.12): adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip - storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
- acerpve (192.168.68.9): llm-gpu - acerpve (192.168.68.9): llm-gpu
- ocupve (192.168.68.5): ocu-llm - ocupve (192.168.68.5): ocu-llm
- hwepve (192.168.68.4): abiba, kagentz, mumuni
**CT IDs (verified 2026-07-24 against PVE API):** **CT IDs (verified 2026-07-24 against PVE API):**
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os 100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
107:pbs 108:media 110:gitea 111:tdunna 112:tanko 107:pbs 108:media 110:gitea 111:tdunna 112:tanko
113:baggy 115:scottdenya 116:syslog-api 117:zulip 113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip
118:jdownloader 119:infisical-vault 118:jdownloader 119:infisical-vault
**NO CT 122, CT 123, or .19 exist in the cluster.** **NO CT 122, CT 123, or .19 exist in the cluster.**
@@ -64,7 +65,7 @@ The infrastructure-control.prose.md contract is the canonical reference for the
**CRITICAL RULES (never regress):** **CRITICAL RULES (never regress):**
1. NO /grafana/ nginx route — it was tried and reverted on 2026-07-02. Grafana is direct LAN at :3001. 1. NO /grafana/ nginx route — it was tried and reverted on 2026-07-02. Grafana is direct LAN at :3001.
2. NO .19 IP — Zulip is CT 117 on storepve. 2. NO .19 IP — Zulip is CT 117 on storepve.
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere. 3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 114.
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24). 4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only. 5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only.
+4 -4
View File
@@ -12,10 +12,10 @@ echo ""
# Authorized agents for restricted contracts # Authorized agents for restricted contracts
# Format: contract_pattern|authorized_agents (comma-separated) # Format: contract_pattern|authorized_agents (comma-separated)
declare -A RESTRICTED declare -A RESTRICTED
RESTRICTED["infrastructure-control.prose.md"]="abiba,abiba-bot,tanko,tanko-bot,mumuni,mumuni-bot" RESTRICTED["infrastructure-control.prose.md"]="abiba"
RESTRICTED["proxmox-monitor.prose.md"]="abiba,abiba-bot" RESTRICTED["proxmox-monitor.prose.md"]="abiba"
RESTRICTED["hermes-config-template.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot" RESTRICTED["hermes-config-template.prose.md"]="abiba,mumuni,tanko"
RESTRICTED["zulip-health.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot" RESTRICTED["zulip-health.prose.md"]="abiba,mumuni"
RESTRICTED["scripts/pm2-self-heal.sh"]="abiba" RESTRICTED["scripts/pm2-self-heal.sh"]="abiba"
RESTRICTED["scripts/prose-lint.sh"]="abiba" RESTRICTED["scripts/prose-lint.sh"]="abiba"
RESTRICTED["scripts/prose-ai-review.sh"]="abiba" RESTRICTED["scripts/prose-ai-review.sh"]="abiba"
+26 -11
View File
@@ -9,11 +9,10 @@ ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8" ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
OWNER_ZULIP_ID="9" OWNER_ZULIP_ID="9"
# Email config
LOG="/root/zulip-health-monitor.log" GMAIL_USER="jtabiri@gmail.com"
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC') GMAIL_PASS="rgbuomwcydxwbszd"
ISSUES=0 EMAIL_TO="jerome@sysloggh.com"
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
notify() { notify() {
local severity="$1" msg="$2" local severity="$1" msg="$2"
@@ -25,14 +24,30 @@ notify() {
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \ curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \ -u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
-d "${form}" > /dev/null 2>&1 || true -d "${form}" > /dev/null 2>&1 || true
# Zulip stream post to #agent-hub on topic 'zulip-health'
local stream_content="${severity} Zulip Monitor: ${msg}" # Email alert
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \ local subject="${severity} Zulip Monitor Alert"
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \ python3 -c "
-d "type=stream\&to=%5B7%5D\&topic=zulip-health\&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote(str()))")" \ import smtplib
> /dev/null 2>&1 || true from email.mime.text import MIMEText
m = MIMEText('''${msg}''')
m['From'] = 'abiba@sysloggh.com'
m['To'] = '${EMAIL_TO}'
m['Subject'] = '${subject}'
s = smtplib.SMTP('smtp.gmail.com', 587)
s.starttls()
s.login('${GMAIL_USER}', '${GMAIL_PASS}')
s.sendmail('abiba@sysloggh.com', ['${EMAIL_TO}'], m.as_string())
s.quit()
" 2>/dev/null || true
} }
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
ISSUES=0
LOG="/root/zulip-health-monitor.log"
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
# ── Global: Zulip Server ── # ── Global: Zulip Server ──
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \ SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
https://chat.sysloggh.net/api/v1/server_settings \ https://chat.sysloggh.net/api/v1/server_settings \
+15 -19
View File
@@ -1,7 +1,7 @@
--- ---
kind: responsibility kind: responsibility
name: zulip-health name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity. description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Hermes agents Tanko/Mumuni), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
title: Zulip Mesh Health Monitor — Multi-Platform title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.0.0 version: 3.0.0
runtime_contract: 2 runtime_contract: 2
@@ -12,13 +12,13 @@ report_only_agents:
# Zulip Mesh Health Monitor # Zulip Mesh Health Monitor
Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero). Monitors ALL Zulip-connected agents across three platforms (pi, Hermes, Agent Zero).
Runs every 15 minutes in the background. Also triggers on session start. Runs every 15 minutes in the background. Also triggers on session start.
## Requires ## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.14, kagentz CT105 on minipve), and Agent Zero Docker host (192.168.68.14) - **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on hwepve), and Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management - **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -47,7 +47,7 @@ Runs every 15 minutes in the background. Also triggers on session start.
"severity": "healthy" "severity": "healthy"
}, },
"tanko": { "tanko": {
"platform": "dsh", "platform": "hermes",
"zulip_state": "connected", "zulip_state": "connected",
"heartbeat_age_seconds": 45, "heartbeat_age_seconds": 45,
"gateway_pid": 1234, "gateway_pid": 1234,
@@ -83,13 +83,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
## Streaming Support (2026-07-05) ## Streaming Support (2026-07-05)
Zulip agents now support progressive message editing during agent generation. Zulip agents now support progressive message editing during agent generation.
When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is When a Hermes agent (Tanko, Mumuni) processes a message, the response is
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API: streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
- Adapter implements `edit_message()` using `_api_patch()` helper - Adapter implements `edit_message()` using `_api_patch()` helper
- Gateway stream consumer progressively edits the Zulip message - Gateway stream consumer progressively edits the Zulip message
- User sees real-time agent thinking instead of waiting for full response - User sees real-time agent thinking instead of waiting for full response
- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active - Verified: Tanko (CT 112) and Mumuni (CT 114) both have streaming active
### Verification ### Verification
```bash ```bash
@@ -184,14 +184,15 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| `last_error` set | Log and monitor | | `last_error` set | Log and monitor |
| Crash loop >10/h | Alert user | | Crash loop >10/h | Alert user |
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24)
**B1: Gateway State** **B1: Gateway State**
```bash ```bash
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100
``` ```
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so there is no `~/.hermes/gateway_state.json` on CT 112 (.122). Verify Tanko's Zulip connectivity via the DSH harness bot status instead.
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed. Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
**B2: Agent Process** **B2: Agent Process**
@@ -200,26 +201,21 @@ Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error`
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep" ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
``` ```
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling): Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway via PM2 (`pm2 restart mumuni-zulip`). Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
| Agent | Restart command | Notes | **B3: Heartbeat Verification**
|-------|-----------------|-------|
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
**B3: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
```bash ```bash
ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3" ssh root@<CT> "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
``` ```
Expected: recent heartbeat (within 5 min), `polls=N` incrementing. Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
Silence > 300s → warning. Silence > 600s → critical. Silence > 300s → warning. Silence > 600s → critical.
**B4: Response Delivery** (Hermes agent Mumuni only) **B4: Response Delivery**
```bash ```bash
ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10" ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
``` ```
> 50% fail rate → critical. > 50% fail rate → critical.
@@ -228,7 +224,7 @@ ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.her
| Condition | Action | | Condition | Action |
|-----------|--------| |-----------|--------|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service | | `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
| No heartbeat in 10min | Same as above | | No heartbeat in 10min | Same as above |
| `Failed to finalize` > 50% | Check PATCH API, Zulip server | | `Failed to finalize` > 50% | Check PATCH API, Zulip server |
| Response empty/short | Check A2A endpoint / LiteLLM model | | Response empty/short | Check A2A endpoint / LiteLLM model |
+1 -1
View File
@@ -10,7 +10,7 @@ description: >
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and > **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability > PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
> monitoring now happens through Telegram. Agents (Mumuni on Hermes, Tanko on DSH) and > monitoring now happens through Telegram. Hermes agents (Tanko, Mumuni) and
> Agent Zero (kagentz) continue to use Zulip. > Agent Zero (kagentz) continue to use Zulip.
## Maintains ## Maintains
+1 -1
View File
@@ -420,7 +420,7 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S
|-------|----------|-------------|--------------|-------------| |-------|----------|-------------|--------------|-------------|
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s | | **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) | | **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
| **Mumuni** | Hermes (kagentz CT 105, migrated 2026-08-29) | ✅ Connected | No issues found | None needed | | **Mumuni** | Hermes (CT 114) | ✅ Connected | No issues found | None needed |
### Key Fixes Applied ### Key Fixes Applied
+6 -5
View File
@@ -13,8 +13,7 @@ triggers:
> **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code > **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code
> (`performHealthCheck()`). That code has been removed. Zulip self-healing for > (`performHealthCheck()`). That code has been removed. Zulip self-healing for
> agents (Mumuni on Hermes, Tanko on DSH) continues through their own platform > Hermes agents (Tanko, Mumuni) continues through their own gateway monitoring.
> monitoring.
## Maintains ## Maintains
@@ -66,7 +65,8 @@ triggers:
|------|----|------|---------| |------|----|------|---------|
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` | | Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` | | Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
| Tanko | 192.168.68.122 (CT 112) | jerome | DSH (DeepSeek Harness) — restart via DSH service, not `hermes gateway restart` (no longer a Hermes agent since 2026-08-27) | | Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` |
| Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` |
## Debounce ## Debounce
@@ -75,8 +75,9 @@ Track via `/tmp/zulip-heal-debounce-<agent>` (unix timestamp of last restart).
## Reporting ## Reporting
This contract is RETIRED — health-check logs are NOT knowledge graph content. Every cycle produces a knowledge graph node:
No graph nodes are created. Logs go to Gitea (SyslogSolution/health-logs). - Title: `[LEARN] zulip-self-heal: <timestamp>`
- metadata: { type: "remediation", status: "fixed" | "escalated" | "healthy" }
- Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>" - Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>"
- Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention" - Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention"