Compare commits
14
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
97977f320c | ||
|
|
a3786ab289 | ||
|
|
15e8998289 | ||
|
|
25c6c76fb3 | ||
|
|
4819247e39 | ||
|
|
1dc4155d9a | ||
|
|
09efbdf4e5 | ||
|
|
4058010e57 | ||
|
|
776df418ca | ||
|
|
c0191c9edc | ||
|
|
f6e8791352 | ||
|
|
b1ef4cbe9b | ||
|
|
41ef546d21 | ||
|
|
8f89603c31 |
@@ -21,7 +21,7 @@ jobs:
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
run: |
|
||||
git clone --depth=50 "http://git.sysloggh.net/${{ gitea.repository }}.git" .
|
||||
git clone --depth=50 "http://192.168.68.17:3000/${{ gitea.repository }}.git" .
|
||||
git fetch origin "${{ gitea.ref }}" --depth=50
|
||||
git checkout "${{ gitea.sha }}"
|
||||
|
||||
@@ -34,7 +34,7 @@ jobs:
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
run: |
|
||||
git clone --depth=50 "http://git.sysloggh.net/${{ gitea.repository }}.git" .
|
||||
git clone --depth=50 "http://192.168.68.17:3000/${{ gitea.repository }}.git" .
|
||||
git fetch origin "${{ gitea.ref }}" --depth=50
|
||||
git checkout "${{ gitea.sha }}"
|
||||
|
||||
@@ -64,7 +64,7 @@ jobs:
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
run: |
|
||||
git clone --depth=50 "http://git.sysloggh.net/${{ gitea.repository }}.git" .
|
||||
git clone --depth=50 "http://192.168.68.17:3000/${{ gitea.repository }}.git" .
|
||||
git fetch origin "${{ gitea.ref }}" --depth=50
|
||||
git checkout "${{ gitea.sha }}"
|
||||
|
||||
@@ -77,7 +77,7 @@ jobs:
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
run: |
|
||||
git clone --depth=50 "http://git.sysloggh.net/${{ gitea.repository }}.git" .
|
||||
git clone --depth=50 "http://192.168.68.17:3000/${{ gitea.repository }}.git" .
|
||||
git fetch origin "${{ gitea.ref }}" --depth=50
|
||||
git checkout "${{ gitea.sha }}"
|
||||
|
||||
|
||||
@@ -21,9 +21,11 @@ description: >
|
||||
- connected == true — Establishes and maintains Zulip event queue
|
||||
- reconnect_on_failure == true — Reconnects after queue expiry
|
||||
- recovers_from_network_loss == true — Handles temporary network drops
|
||||
- stream_subscription_check == true — Bot is subscribed to primary streams (Gen 5+)
|
||||
|
||||
*Message Flow*
|
||||
- dm_response_time_ms < 10000 — DMs get a response within 10 seconds
|
||||
- dm_response_time_offline_verifiable == true — simulate_dm_latency() passes without live Zulip (Gen 6+)
|
||||
- stream_mention_detection == true — @mentions in streams are detected and routed
|
||||
- all_bots_detection == true — @all-bots mentions are detected
|
||||
- echo_loop_prevention == true — Never replies to its own messages
|
||||
@@ -38,6 +40,14 @@ description: >
|
||||
- graceful_disconnect == true — shutdown doesn't crash
|
||||
- handles_malformed_messages == true — Bad JSON doesn't kill the poll loop
|
||||
- typing_indicators_work == true — Sends typing notifications
|
||||
- known_issues_tracked == true — add_known_issue() persists problems across generations (Gen 6+)
|
||||
|
||||
## Status
|
||||
|
||||
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
|
||||
for any Hermes agent that connects to Zulip (Mumuni, Tanko, Koby, Koonimo). The pi
|
||||
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
|
||||
plugin system, which is unaffected.
|
||||
|
||||
## Parameters
|
||||
|
||||
@@ -90,13 +100,28 @@ Each generation stores:
|
||||
|
||||
## Execution
|
||||
|
||||
1. **Read state** — Check previous generations from knowledge graph
|
||||
1. **Read state** — Check previous generations from knowledge graph node `build-zulip-plugin:state`
|
||||
2. **Generate or improve** — Based on generation count:
|
||||
- Generation 1: Scaffold full plugin from scratch
|
||||
- Generation N: Read previous code, apply improvements
|
||||
3. **Validate** — Syntax check, structure check, dependency check
|
||||
4. **Log** — Save generation result to knowledge graph as [LEARN]
|
||||
5. **Report** — Output what was generated, what changed, what needs review
|
||||
4. **Run offline tests** — Always run `simulate_dm_latency()` and selftest() if no live Zulip
|
||||
5. **Log** — Save generation result to knowledge graph as [LEARN] node. Update the
|
||||
`build-zulip-plugin:state` node with current generation_count, plugin_version,
|
||||
and known_issues.
|
||||
6. **Report** — Output what was generated, what changed, what needs review
|
||||
|
||||
## State Persistence (Gen 6+)
|
||||
|
||||
All generation state is stored in a knowledge graph node titled `build-zulip-plugin:state`:
|
||||
- `generation_count` — incremented on each verified run
|
||||
- `plugin_version` — current version string
|
||||
- `last_generation` — ISO timestamp of most recent run
|
||||
- `known_issues` — array of unresolved issues carried forward
|
||||
- `last_postcondition_results` — pass/fail map from most recent verification
|
||||
|
||||
Agents executing this contract MUST read this node on startup and update it after
|
||||
validation. If the node does not exist, create it with generation_count=1.
|
||||
|
||||
## Example Output
|
||||
|
||||
|
||||
@@ -3,10 +3,11 @@ kind: responsibility
|
||||
name: disk-gc-threat-response
|
||||
description: >
|
||||
Recurring disk health scan, garbage collection, and threat response
|
||||
across all 19 Proxmox CTs and Docker hosts. Triggered by incident
|
||||
across 15 Proxmox CTs + 3 GPU bare-metal hosts. Triggered by incident
|
||||
2026-07-04 where CT 105 (kagentz) hit 87% disk (49G/59G) from
|
||||
Docker image bloat — 5 dangling images, 15 build cache layers.
|
||||
Recovered 35.67GB via `docker system prune -a --force`.
|
||||
Recovered 35.67GB. Second incident 2026-07-09: amdpve (.15) Docker
|
||||
bloat (15.5GB abandoned HIP image), reclaimed 11.56GB.
|
||||
id: 067NV8KJ03ZG71S44N41F31022
|
||||
version: 1.0.0
|
||||
---
|
||||
@@ -21,15 +22,21 @@ logged within 5 minutes of discovery.
|
||||
|
||||
## Scope
|
||||
|
||||
All CTs in the Proxmox cluster, with special attention to Docker hosts:
|
||||
All 15 CTs via `pct-run` + 3 GPU bare-metal hosts via direct SSH.
|
||||
Docker hosts get special attention:
|
||||
|
||||
| Host | CT | Disk Risk | GC Strategy |
|
||||
|------|----|-----------|-------------|
|
||||
| kagentz | 105 | HIGH — Agent Zero builds images | `docker system prune -a` |
|
||||
| syslog-api | 116 | HIGH — Prometheus data, 8 containers | `docker system prune`, log rotate |
|
||||
| docker-vm | 109 | MED — 11 containers, NFS mounts | `docker system prune`, check mounts |
|
||||
| syslog-api | 116 | HIGH — Prometheus data, 10 containers | `docker system prune`, log rotate |
|
||||
| docker-vm | 109 | HIGH — 16 containers across 4 stacks, NFS mounts | `docker system prune`, check mounts |
|
||||
| amdpve | — | MED — GPU bare metal, Docker for one-off builds | `docker system prune -a` |
|
||||
| abiba | 100 | LOW — local docker, go cache | apt/docker/log prune |
|
||||
| All others | — | LOW — no Docker | apt clean, log rotate |
|
||||
| All other CTs | — | LOW — no Docker | apt clean, log rotate |
|
||||
| GPU bare metal (.8, .110) | — | LOW — no Docker on GPU hosts | log rotate |
|
||||
|
||||
> **Decommissioned:** CT 118 (jitsi) — intentionally stopped, not scanned.
|
||||
> **Migrated:** CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
|
||||
|
||||
## Threat Levels
|
||||
|
||||
@@ -123,7 +130,7 @@ call summary-reporter
|
||||
|
||||
## GC Strategies by Host Type
|
||||
|
||||
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109)
|
||||
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
||||
|
||||
```bash
|
||||
# Phase 1: Safe prune (won't touch running containers' images)
|
||||
@@ -224,32 +231,76 @@ dangling images and orphaned build cache. No automated GC was in place.
|
||||
- This contract now runs disk GC fleet-wide every 6 hours
|
||||
- Docker hosts get `docker system prune` on amber, `-a --force` on red
|
||||
- kagentz is flagged as HIGH risk due to development activity
|
||||
- amdpve is flagged for Docker bloat monitoring — abandoned build images accumulate
|
||||
```
|
||||
|
||||
Done. Here's what we have:
|
||||
## Incident Log: 2026-07-09 — amdpve docker bloat
|
||||
|
||||
## Storage Summary
|
||||
### Discovery
|
||||
Scheduled fleet disk scan via `pct-run` across all 15 CTs + 3 GPU bare-metal hosts.
|
||||
amdpve (.15) flagged at 78% (AMBER threshold: 75%).
|
||||
|
||||
| System | Used | Free | Health |
|
||||
|--------|------|------|--------|
|
||||
| **Abiba** (CT 100) | 12G / 59G (21%) | 45G | ✅ Green |
|
||||
| **kagentz** (CT 105) | 16G / 59G (27%) | 41G | ✅ Resolved (was 87%) |
|
||||
| **/var/lib/docker** (abiba) | 2.8G | — | Fine |
|
||||
| **/root/go** (abiba) | 773M | — | Fine |
|
||||
### Diagnosis
|
||||
- amdpve (.15, Strix Halo host): 70G used / 94G total (78%)
|
||||
- Docker images: 1 image, 0 containers running, 15.47GB (100% reclaimable)
|
||||
- Image: `llama-strix-hip:latest` — abandoned ROCm/HIP Docker build from 7 days ago
|
||||
- Root cause: Strix Halo migrated from Docker-based HIP path to bare-metal Vulkan
|
||||
(`/root/llama.cpp/build-vk/`) but the old Docker image was never cleaned up
|
||||
- Not a running service — zero containers, zero active volumes
|
||||
|
||||
## New Contract: `disk-gc-threat-response.prose.md`
|
||||
### Resolution
|
||||
```
|
||||
docker system prune -a --force
|
||||
→ Reclaimed 11.56GB
|
||||
→ Post: 55G / 94G (62%), 35G free
|
||||
→ 1 image removed (llama-strix-hip:latest, 15.5GB)
|
||||
→ 8 build cache layers removed
|
||||
→ amdpve now GREEN
|
||||
```
|
||||
|
||||
Created at `/root/prose-contracts/disk-gc-threat-response.prose.md`. Here's what it does:
|
||||
### Root Cause
|
||||
Technology migration (Docker HIP → bare-metal Vulkan) left orphaned build
|
||||
artifacts. Docker on amdpve serves no running purpose — it's only used for
|
||||
one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
|
||||
### Three-in-One
|
||||
### Preventive Measures
|
||||
- amdpve added to Docker GC scan list
|
||||
- Post-migration cleanup step added: after any GPU backend migration, prune
|
||||
the old backend's Docker images within 24 hours
|
||||
- Contract now scans GPU bare-metal hosts alongside CTs
|
||||
- Access via `pct-run` script for all CTs (no hardcoded IPs)
|
||||
|
||||
1. **Infra Update** — Fleet-wide disk scan every 6 hours across all 19 CTs, with per-host GC strategy (Docker hosts vs standard LXC). Today's incident is logged as the baseline.
|
||||
## Access Matrix (documented 2026-07-09)
|
||||
|
||||
2. **Garbage Collection** — Tiered response: Phase 1 (`docker system prune -f`) for amber, Phase 2 (`-a --force`) for red, Phase 3 (`--volumes` + `builder prune --all`) for critical. Non-Docker CTs get `apt clean`, log rotation, journalctl vacuum.
|
||||
### CT Access (via pct-run)
|
||||
| CT | Name | Node | Status |
|
||||
|----|------|------|--------|
|
||||
| 100 | abiba | amdpve | local |
|
||||
| 102 | adguard | acerpve | ✅ reachable |
|
||||
| 104 | authentik | minipve | ✅ reachable |
|
||||
| 105 | kagentz | amdpve | ✅ reachable |
|
||||
| 106 | ra-h-os | storepve | ✅ reachable |
|
||||
| 107 | pbs | storepve | ✅ reachable |
|
||||
| 108 | media | storepve | ✅ reachable |
|
||||
| 110 | gitea | minipve | ✅ reachable |
|
||||
| 111 | tdunna | amdpve | ✅ reachable |
|
||||
| 112 | tanko | amdpve | ✅ reachable |
|
||||
| 113 | baggy | amdpve | ✅ reachable |
|
||||
| 114 | mumuni | minipve | ✅ reachable |
|
||||
| 115 | scottdenya | amdpve | ✅ reachable |
|
||||
| 116 | syslog-api | minipve | ✅ reachable |
|
||||
| 117 | zulip | storepve | ✅ reachable |
|
||||
|
||||
3. **Threat Resolution** — Five severity levels (Green → Amber → Red → Critical → Full), with escalating alerts via Zulip DM, channel, and relay to you for critical breaches. kagentz is flagged HIGH risk due to Agent Zero dev patterns.
|
||||
### GPU Bare Metal (via direct SSH)
|
||||
| Host | IP | GPU | Status |
|
||||
|------|-----|-----|--------|
|
||||
| llm-gpu | 192.168.68.8 | RTX 3090 | ✅ reachable |
|
||||
| ocu-llm | 192.168.68.110 | RTX 5070 | ✅ reachable |
|
||||
| amdpve | 192.168.68.15 | Strix Halo | ✅ reachable |
|
||||
|
||||
### The incident root cause
|
||||
Agent Zero was repeatedly rebuilding `kagentz-bridge`, generating 5 dangling images and 15 build cache layers. No automated GC existed. Recovery: **35.67GB freed** in one `docker system prune -a --force`.
|
||||
### KVM VM (via direct SSH)
|
||||
| Host | IP | Role | Status |
|
||||
|------|-----|------|--------|
|
||||
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
|
||||
|
||||
Want me to push this to the prose-contracts repo?
|
||||
> **Decommissioned:** CT 118 (jitsi) — intentionally stopped.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
||||
+42
-17
@@ -5,8 +5,9 @@ description: >
|
||||
Manages the GPU inference fleet across all hosts. Handles model deployment,
|
||||
registration, health checks, LiteLLM sync, agent key management, GPU
|
||||
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
|
||||
Current as of 2026-06-30 post-migration: nginx routes /v1 → LiteLLM,
|
||||
router runs behind LiteLLM on :9000 (internal-only).
|
||||
Current as of 2026-07-08: context reduced to 128K on NVIDIA GPUs, parallel 2
|
||||
on all GPUs, LiteLLM timeouts tuned (gemma 25→120s, qwen 40→90s), router fully
|
||||
deprecated — nginx routes /v1 → LiteLLM directly.
|
||||
agent: abiba
|
||||
triggers:
|
||||
- on model add/remove
|
||||
@@ -63,21 +64,21 @@ triggers:
|
||||
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||
│ │ │ │ │ llama-srv │ │ Watchdog │
|
||||
│ 128K ctx │ │ 128K ctx │ │ 256K ctx │ │ Watchdog │
|
||||
│ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │
|
||||
│ 27B-code │ │ :8090 │ │ :8080 │ │ exporter │
|
||||
│ :8090 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||
│ 27B-code │ │ :8080 │ │ :8080 │ │ exporter │
|
||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
||||
└──────────┘
|
||||
```
|
||||
|
||||
## Current Model Assignments
|
||||
## Current Model Assignments (2026-07-08)
|
||||
|
||||
| Model | GPU | Host | VRAM | Status |
|
||||
|-------|-----|------|------|--------|
|
||||
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 23.4/24GB | ✅ healthy |
|
||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 11.2/12.2GB | ✅ healthy |
|
||||
| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | 22.7/64GB | ✅ healthy |
|
||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 20.3/24GB (83%) | 128K | turbo4 | 2 | 512/512 | ✅ healthy |
|
||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 9.4/12.2GB (77%) | 128K | q4_0 | 2 | 2048/512 | ✅ healthy |
|
||||
| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | 24.4/64GB (35%) | 256K | q8_0 | 2 | 2048/512 | ✅ healthy |
|
||||
|
||||
## Operations
|
||||
|
||||
@@ -121,7 +122,19 @@ triggers:
|
||||
7. Document keys in knowledge graph
|
||||
|
||||
### list
|
||||
Show full fleet status: GPUs, models, VRAM, active requests, circuit breakers, keys
|
||||
Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, active requests, circuit breakers, keys
|
||||
|
||||
### health-check
|
||||
1. Check GPU hardware: nvidia-smi (.8, .110) + amdgpu sysfs (.15 via /sys/class/drm/card0/device/)
|
||||
2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts
|
||||
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
||||
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
||||
- gemma-4-12b: 120s, qwen3.6-27B: 90s, ornith-1.0-35b: 120s
|
||||
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
||||
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
|
||||
|
||||
## Agent Keys (LiteLLM DB — Current 2026-06-30)
|
||||
|
||||
@@ -168,9 +181,11 @@ If no SSH access, send Zulip DM via abiba-bot.
|
||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||
- **VRAM**: RTX 3090 at ~9.6/24GB (60% free), RTX 5070 at ~6.5/12GB (46% free). Healthy.
|
||||
- **RTX 3090 optimization**: `--parallel 2 --batch-size 4096` — doubled concurrent throughput.
|
||||
- **RTX 5070 optimization**: Removed draft model freed ~2GB VRAM for vision tasks. `--parallel 2`.
|
||||
- **VRAM (2026-07-08)**: RTX 3090 at 20.3/24GB (83%), RTX 5070 at 9.4/12.2GB (77%), Strix Halo at 24.4/64GB (35%). Context reduced from 256K→128K on NVIDIA GPUs freed ~3.3GB (.8) and ~1.5GB (.110).
|
||||
- **All GPUs at `--parallel 2` (2026-07-08)**: Fleet serves 6 concurrent requests (was 3). 2× throughput.
|
||||
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2`. No explicit batch flags (512/512 default). Service: `/home/llmuser/llama-wrapper.sh`.
|
||||
- **RTX 5070 config**: `--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 512 --parallel 2`. Ubatch fixed 4096→512 (was inverted — ubatch > batch killed prompt throughput). Service: `/home/llmuser/llama-wrapper.sh`.
|
||||
- **LiteLLM timeout tuning (2026-07-08)**: gemma-4-12b 25→120s, qwen3.6-27B-code 40→90s, syslog-auto (qwen route) 40→90s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. Config at `/opt/inference-harness/litellm_config.yaml`.
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `ornith-server.service` on port 8080, 256K context, flash-attn + q8 KV.
|
||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||
- **Strix Halo thermal safeguard (2026-07-02)**: `ornith-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||
@@ -185,8 +200,8 @@ If no SSH access, send Zulip DM via abiba-bot.
|
||||
|
||||
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Samples |
|
||||
|-----|-------|-----------|--------------|----------|---------|
|
||||
| RTX 3090 (.8) | gemma-4-12b | 75 | 305 | 74 | 6 |
|
||||
| RTX 5070 (.110) | qwen3.6-27B-code | 75 | 323 | 75 | 6 |
|
||||
| RTX 3090 (.8) | qwen3.6-27B-code | 75 | 305 | 74 | 6 |
|
||||
| RTX 5070 (.110) | gemma-4-12b | 75 | 323 | 75 | 6 |
|
||||
| Strix Halo (.15) | ornith-1.0-35b | 70 | 532 | 70 | 6 |
|
||||
|
||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
||||
@@ -194,3 +209,13 @@ Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
|
||||
History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
||||
|
||||
**Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration.
|
||||
|
||||
## Agent Config Implications (2026-07-08)
|
||||
|
||||
With NVIDIA GPUs at 128K context:
|
||||
- Agents using `syslog-auto` (50/50 qwen+ornith): keep `context_length: 262144` — ornith supports it, Litellm fallbacks handle qwen overflow
|
||||
- Agents using `qwen3.6-27B-code` directly: set `context_length: 131072` and `max_tokens: 4096` per thermal safety rule
|
||||
- Agents using `gemma-4-12b` directly (auxiliary tasks): set `context_length: 131072`
|
||||
- Compression threshold at 0.65: fires at ~170K for syslog-auto (262K ctx), ~85K for direct qwen/gemma (128K ctx)
|
||||
- All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap
|
||||
- Port 8080 is used on all 3 GPU hosts (not 8090 as previously documented)
|
||||
|
||||
@@ -1,11 +1,11 @@
|
||||
---
|
||||
kind: reference
|
||||
kind: template
|
||||
name: hermes-agent-baseline
|
||||
version: 1.0.0
|
||||
description: >
|
||||
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
||||
configuration state, keys, workarounds, and audit procedure. When an agent's
|
||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-05.
|
||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-08. GPU context reduced to 128K on .8/.110, parallel 2 fleet-wide.
|
||||
author: Abiba (pi agent)
|
||||
---
|
||||
|
||||
@@ -22,12 +22,12 @@ done
|
||||
|
||||
## Agent Map
|
||||
|
||||
| Agent | CT | Node | IP | LiteLLM Key | LiteLLM Alias |
|
||||
|-------|-----|------|-----|-------------|---------------|
|
||||
| Tanko | 112 | amdpve | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | `tanko` |
|
||||
| Mumuni | 114 | minipve | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | `mumuni` |
|
||||
| Tdunna | 111 | amdpve | ? | `sk-6sbCNjz2T6lTVDBdlNHXsA` | `tdunna` |
|
||||
| Baggy | 113 | amdpve | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | `baggy` |
|
||||
| Agent | CT | Node | IP | LiteLLM Key | LiteLLM Alias | Platform |
|
||||
|-------|-----|------|-----|-------------|---------------|----------|
|
||||
| Tanko | 112 | amdpve | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | `tanko` | Hermes |
|
||||
| Mumuni | 114 | minipve | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | `mumuni` | Hermes |
|
||||
| Tdunna | 111 | amdpve | srv1079750 | `sk-Qvzi4uYQBhlSK_XstEhcyQ` | `tdunna` | **pi** |
|
||||
| Baggy | 113 | amdpve | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | `baggy` | Hermes |
|
||||
|
||||
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
|
||||
|
||||
@@ -45,6 +45,8 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) →
|
||||
|
||||
## Config Pattern — Mandatory Fields
|
||||
|
||||
### For Hermes Agents (Tanko, Mumuni, Baggy)
|
||||
|
||||
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
|
||||
|
||||
### 1. Main Model
|
||||
@@ -154,6 +156,51 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
|
||||
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
|
||||
```
|
||||
|
||||
### For pi Agents (Tdunna)
|
||||
|
||||
Tdunna (CT111) runs pi 0.80.3 via PM2 with the Zulip extension (router-worker architecture).
|
||||
Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`.
|
||||
|
||||
**models.json** — Must only list models authorized for the agent's LiteLLM key:
|
||||
```json
|
||||
{
|
||||
"providers": {
|
||||
"syslog-harness": {
|
||||
"baseUrl": "http://192.168.68.116/v1",
|
||||
"api": "openai-completions",
|
||||
"apiKey": "sk-...",
|
||||
"models": [
|
||||
{ "id": "syslog-auto" },
|
||||
{ "id": "ornith-1.0-35b" },
|
||||
{ "id": "qwen3.6-27B-code" },
|
||||
{ "id": "gemma-4-12b" }
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**settings.json** — Always use `syslog-auto` as default:
|
||||
```json
|
||||
{
|
||||
"defaultProvider": "syslog-harness",
|
||||
"defaultModel": "syslog-auto"
|
||||
}
|
||||
```
|
||||
|
||||
**Validation**: Verify models match LiteLLM's authorized list:
|
||||
```bash
|
||||
curl -s http://192.168.68.116:4000/v1/models \
|
||||
-H "Authorization: Bearer $(grep apiKey ~/.pi/agent/models.json | head -1 | cut -d'"' -f4)" \
|
||||
| jq '.data[].id'
|
||||
```
|
||||
|
||||
**Stuck worker detection**: In PM2 logs, `workers=[<id>:busy:N]` with growing N indicates
|
||||
a stuck worker (model error, no `agent_end` emitted). Fix: correct models.json, delete
|
||||
stale sessions from `~/.pi/agent/sessions/zulip/`, restart PM2.
|
||||
|
||||
**Service**: `pm2 restart koby-zulip`, health at `:9201/health`.
|
||||
|
||||
## Systemd Pattern
|
||||
|
||||
For agents where the gateway runs as a user service:
|
||||
@@ -210,5 +257,6 @@ If they differ → ghost detected → kill ghost → start fresh.
|
||||
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-07-08 | Tdunna: fixed model mismatch (qwen3.6-35B-A3B→syslog-auto), added pi-specific config section. Key updated to sk-Qvzi4uYQBhlSK_XstEhcyQ. Added Failure Mode #11 to zulip-adapter-lessons. |
|
||||
| 2026-07-06 | Port conflict detection added to all 3 GPU wrappers. Consolidated health check script deployed. Zulip streaming edit_message enabled for Tanko/Mumuni. |
|
||||
| 2026-07-05 | Baseline created. All 4 agents audited, master key removed, api_key workaround applied |
|
||||
|
||||
@@ -5,7 +5,8 @@ description: >
|
||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||
Updated 2026-06-30: new LiteLLM keys, nginx routing, router back online.
|
||||
Updated 2026-07-08: GPU context reduced to 128K on NVIDIA (.8, .110),
|
||||
parallel 2 on all GPUs, LiteLLM timeouts tuned, context_length guidance added.
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -92,7 +93,8 @@ model:
|
||||
base_url: http://192.168.68.116/v1
|
||||
api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env
|
||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||
context_length: 262144 # Must match model's max context window
|
||||
context_length: 262144 # For syslog-auto (ornith route supports 256K).
|
||||
# Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly.
|
||||
|
||||
fallback_providers:
|
||||
provider: deepseek
|
||||
@@ -123,8 +125,9 @@ compression:
|
||||
enabled: true
|
||||
model: gemma-4-12b # ⚠️ Must match auxiliary.compression.model
|
||||
provider: harness
|
||||
max_context_window: 262144 # Must match model's actual capacity
|
||||
threshold: 0.65 # Fires at ~170K for 262K window (not 0.25!)
|
||||
max_context_window: 262144 # For syslog-auto (ornith supports 256K).
|
||||
# Set 131072 if using qwen or gemma directly.
|
||||
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
||||
target_ratio: 0.30
|
||||
protect_last_n: 40
|
||||
hygiene_hard_message_limit: 350
|
||||
@@ -236,6 +239,28 @@ The following MUST be identical across ALL profiles:
|
||||
- `max_context_window: 262144` MUST match the model's actual capacity
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 9: Default Model Must Be `syslog-auto` (All Agents)
|
||||
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
||||
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
||||
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
|
||||
and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
|
||||
- Model name typos that cause 403 errors and silent worker failures
|
||||
- Single GPU downtime (routing falls back automatically)
|
||||
- Key/model authorization mismatches
|
||||
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
|
||||
for specialized tasks, but MUST validate those models exist in the key's authorized list
|
||||
|
||||
### Rule 10: Validate Model IDs Before Deployment (pi Agents)
|
||||
- After configuring a pi agent's `models.json`, verify every model ID:
|
||||
```bash
|
||||
curl -s http://192.168.68.116:4000/v1/models \
|
||||
-H "Authorization: Bearer <AGENT_KEY>" | jq '.data[].id'
|
||||
```
|
||||
- All model IDs in `models.json` MUST appear in the LiteLLM response
|
||||
- The agent's API key may have a SUBSET of the full model catalog — check per-key
|
||||
- A non-existent model ID causes 403 errors that silently break the pi RPC worker
|
||||
(no `agent_end` emitted, worker stays "busy", Zulip messages pile up unprocessed)
|
||||
|
||||
## Execution
|
||||
|
||||
1. **Check current config** — Read the target agent's config.yaml
|
||||
|
||||
@@ -0,0 +1,237 @@
|
||||
---
|
||||
kind: function
|
||||
name: hermes-zulip-plugin
|
||||
description: >
|
||||
Installs or updates the Zulip platform plugin for any Hermes agent from the
|
||||
canonical zulip-platform-plugins repo (master branch). Ensures the agent runs
|
||||
the latest adapter with all Zulip chat fixes (_strip_html, streaming, event
|
||||
recovery). Verifies the installation, restarts the gateway, and sends a relay
|
||||
success signal.
|
||||
agent: abiba
|
||||
version: 1.0.0
|
||||
status: active
|
||||
runtime_contract: 2
|
||||
---
|
||||
|
||||
# Hermes Zulip Plugin — Install & Repair
|
||||
|
||||
Single-shot function that pulls the latest zulip-platform plugin from the
|
||||
canonical git repo, installs it to the correct Hermes bundled plugin path,
|
||||
verifies the installation, and signals completion.
|
||||
|
||||
Distinct from `hermes-zulip-restore`: this contract is focused on the plugin
|
||||
layer and a targeted gateway restart. It does NOT verify env credentials or
|
||||
run a live Zulip connection test. Use `hermes-zulip-restore` for full
|
||||
connectivity recovery including end-to-end DM validation.
|
||||
|
||||
## Parameters
|
||||
|
||||
| Param | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, or `koby` |
|
||||
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
|
||||
|
||||
## Maintains
|
||||
|
||||
- plugin_installed: bool — Whether all three adapter files exist at the bundled path
|
||||
- plugin_version: string — Git commit SHA of the installed version
|
||||
- strip_html_present: bool — Whether `_strip_html` fix is in the installed adapter
|
||||
- signal_sent: bool — Whether relay success message was dispatched
|
||||
|
||||
### Postconditions
|
||||
|
||||
- All three adapter files (`__init__.py`, `adapter.py`, `plugin.yaml`) present in `<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/`
|
||||
- `_strip_html` function exists in `adapter.py` (slash-command fix, commit `55ca15d`+)
|
||||
- Installed version matches HEAD of the requested branch
|
||||
- Relay success signal sent to Hermes agent's inbox
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to target host (direct or via amdpve for CTs)
|
||||
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
|
||||
- Python 3 with `httpx` installed on target
|
||||
|
||||
## Live-State Fields
|
||||
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Mumuni | CT114 | — | 192.168.68.123 | /root/.hermes | root |
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||
|
||||
| Field | Value | Trust |
|
||||
|-------|-------|-------|
|
||||
| Git repo | `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git` | ✅ Verified |
|
||||
| Default branch | `master` (contains all merged fixes including `feat/zulip-streaming`) | ✅ Verified |
|
||||
| Bundled adapter path | `<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/` | ✅ Verified |
|
||||
| Adapter files | `__init__.py`, `adapter.py`, `plugin.yaml` | ✅ Verified |
|
||||
| Strip-html commit | `55ca15d` (minimum) | ✅ Verified |
|
||||
|
||||
## Execution
|
||||
|
||||
### Step 1: Resolve Target
|
||||
|
||||
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
|
||||
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
|
||||
|
||||
### Step 2: Pull Latest Plugin Source
|
||||
|
||||
On the target host:
|
||||
|
||||
```bash
|
||||
# Ensure deploy scratch space
|
||||
mkdir -p /tmp/zulip-deploy
|
||||
cd /tmp/zulip-deploy
|
||||
|
||||
# Clone or pull
|
||||
if [ -d zulip-platform-plugins ]; then
|
||||
cd zulip-platform-plugins
|
||||
git fetch origin
|
||||
git checkout {{branch}}
|
||||
git pull origin {{branch}}
|
||||
else
|
||||
git clone --branch {{branch}} \
|
||||
https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git
|
||||
cd zulip-platform-plugins
|
||||
fi
|
||||
|
||||
# Capture installed version
|
||||
INSTALLED_SHA=$(git rev-parse HEAD)
|
||||
echo "Installed SHA: $INSTALLED_SHA"
|
||||
|
||||
# Verify we're at HEAD
|
||||
HEAD_SHA=$(git rev-parse origin/{{branch}})
|
||||
if [ "$INSTALLED_SHA" = "$HEAD_SHA" ]; then
|
||||
echo "At latest commit on {{branch}}"
|
||||
else
|
||||
echo "WARNING: not at HEAD — $INSTALLED_SHA vs $HEAD_SHA"
|
||||
fi
|
||||
```
|
||||
|
||||
### Step 3: Install Plugin Files
|
||||
|
||||
```bash
|
||||
# Ensure target directory exists
|
||||
mkdir -p {{hermes_home}}/hermes-agent/plugins/platforms/zulip
|
||||
|
||||
# Copy adapter files
|
||||
cp plugins/platforms/zulip/adapter.py \
|
||||
plugins/platforms/zulip/__init__.py \
|
||||
plugins/platforms/zulip/plugin.yaml \
|
||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
# Fix ownership (Tanko only — runs as jerome user)
|
||||
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
|
||||
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
echo "Plugin files installed"
|
||||
```
|
||||
|
||||
### Step 4: Verify Installation
|
||||
|
||||
```bash
|
||||
# Check all three files exist
|
||||
for f in __init__.py adapter.py plugin.yaml; do
|
||||
if [ -f "{{hermes_home}}/hermes-agent/plugins/platforms/zulip/$f" ]; then
|
||||
echo "✅ $f present"
|
||||
else
|
||||
echo "❌ $f MISSING"
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
|
||||
# Verify _strip_html fix
|
||||
grep -q "_strip_html" {{hermes_home}}/hermes-agent/plugins/platforms/zulip/adapter.py \
|
||||
&& echo "✅ _strip_html fix present" \
|
||||
|| echo "❌ _strip_html MISSING — plugin may be stale"
|
||||
|
||||
# Show installed plugin.yaml version
|
||||
grep "^version:" {{hermes_home}}/hermes-agent/plugins/platforms/zulip/plugin.yaml || true
|
||||
```
|
||||
|
||||
### Step 5: Restart Gateway
|
||||
|
||||
Plugin changes require a gateway restart to take effect:
|
||||
|
||||
```bash
|
||||
cd {{hermes_home}}/hermes-agent
|
||||
# Use venv if available
|
||||
python3 -m hermes_cli.main gateway restart 2>&1 || \
|
||||
venv/bin/python -m hermes_cli.main gateway restart 2>&1
|
||||
```
|
||||
|
||||
Wait for restart to complete (up to 45s), then confirm:
|
||||
|
||||
```bash
|
||||
grep "Gateway running" {{hermes_home}}/logs/gateway.log | tail -1
|
||||
```
|
||||
|
||||
Expected: `Gateway running with N platform(s)` where N > 1 (includes zulip).
|
||||
|
||||
Quick smoke check — confirm zulip platform loaded:
|
||||
|
||||
```bash
|
||||
grep -E "zulip.*loaded|zulip.*registered" {{hermes_home}}/logs/gateway.log | tail -3
|
||||
```
|
||||
|
||||
If gateway fails to restart, check logs for the crash cause before proceeding.
|
||||
|
||||
### Step 6: Send Relay Success Signal
|
||||
|
||||
On the Abiba host (local), dispatch a relay message to the target agent:
|
||||
|
||||
```
|
||||
ra-h-os-createRelayNode:
|
||||
title: "Zulip plugin updated — {{target}}"
|
||||
source: |
|
||||
Plugin installed from {{branch}} @ {{INSTALLED_SHA}}
|
||||
All 3 adapter files verified at {{hermes_home}}/hermes-agent/plugins/platforms/zulip/
|
||||
_strip_html fix: PRESENT
|
||||
Timestamp: {{timestamp}}
|
||||
|
||||
description: "hermes-zulip-plugin completed for {{target}} — plugin layer healthy"
|
||||
```
|
||||
|
||||
### Step 7: Report
|
||||
|
||||
Compile results into a single status block:
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| Target | `{{target}}` |
|
||||
| Branch | `{{branch}}` |
|
||||
| Commit SHA | `{{INSTALLED_SHA}}` |
|
||||
| Files installed | `__init__.py`, `adapter.py`, `plugin.yaml` |
|
||||
| `_strip_html` | `{{present|missing}}` |
|
||||
| Gateway restarted | `{{yes|no}}` |
|
||||
| Signal sent | `{{yes|no}}` |
|
||||
|
||||
## Known Failure Modes
|
||||
|
||||
| Symptom | Root Cause | Recovery |
|
||||
|---------|-----------|----------|
|
||||
| Git clone fails | No network or repo unreachable | Check VPN/network, verify repo URL |
|
||||
| Permission denied on copy | Wrong user for target | Use correct user (jerome for Tanko, root for others) |
|
||||
| `_strip_html` missing after install | Branch doesn't include commit `55ca15d` | Switch to `feat/zulip-streaming` branch |
|
||||
| Plugin files missing after copy | Target directory doesn't exist | Ensure `mkdir -p` ran successfully |
|
||||
| Relay signal fails | MCP bridge unreachable | Signal manually via `ra-h-os-createRelayNode` |
|
||||
|
||||
## Edge Differences from hermes-zulip-restore
|
||||
|
||||
| Concern | hermes-zulip-restore | hermes-zulip-plugin |
|
||||
|---------|---------------------|---------------------|
|
||||
| Env credential check | ✅ Full ZULIP_* verification | ❌ Out of scope |
|
||||
| Gateway restart | ✅ Full restart + state validation | ✅ Targeted restart + smoke check |
|
||||
| Live connection test | ✅ Validates `zulip.state = connected` | ❌ Out of scope |
|
||||
| Plugin deploy | ✅ Includes deploy as one step | ✅ Primary purpose |
|
||||
| Version tracking | ❌ Implicit | ✅ Explicit SHA capture |
|
||||
| Signal dispatch | ❌ None | ✅ Relay message to target |
|
||||
|
||||
For full connectivity recovery after a plugin install, chain this contract
|
||||
with `hermes-zulip-restore` (skip its Step 2 to avoid redundant deploy).
|
||||
|
||||
---
|
||||
|
||||
**Last updated**: 2026-07-08 — Switched default branch to `master`; added
|
||||
gateway restart step. If outstanding unmerged feature branches exist, address
|
||||
them in a follow-up merge after this contract completes.
|
||||
@@ -0,0 +1,188 @@
|
||||
---
|
||||
kind: function
|
||||
name: hermes-zulip-restore
|
||||
description: >
|
||||
Restores Zulip connectivity for any Hermes agent (Mumuni CT114, Tanko CT112,
|
||||
Koby CT111). Deploys the zulip-platform adapter to the correct bundled plugin
|
||||
path, verifies env credentials, restarts the gateway, and confirms Zulip
|
||||
connects. Run this whenever a Hermes agent stops responding on Zulip or after
|
||||
a fresh agent deployment.
|
||||
agent: abiba
|
||||
version: 1.0.0
|
||||
status: active
|
||||
runtime_contract: 2
|
||||
---
|
||||
|
||||
# Hermes Zulip Restore — Bring Any Agent Back to Good State
|
||||
|
||||
Single-shot function that restores full Zulip connectivity for a Hermes agent.
|
||||
Covers adapter deployment, HTML stripping (slash command fix), env verification,
|
||||
gateway restart, and connection validation.
|
||||
|
||||
## Parameters
|
||||
|
||||
| Param | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, or `koby` |
|
||||
|
||||
## Maintains
|
||||
|
||||
- adapter_deployed: bool — Whether `_strip_html` adapter is at correct bundled path
|
||||
- zulip_connected: bool — Whether gateway_state shows zulip.state = "connected"
|
||||
- env_valid: bool — Whether .env has ZULIP_SITE, ZULIP_EMAIL, ZULIP_API_KEY
|
||||
- gateway_running: bool — Whether gateway process is running
|
||||
|
||||
### Postconditions
|
||||
|
||||
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
|
||||
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
|
||||
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko)
|
||||
- Gateway restarted and zulip platform reports state `connected`
|
||||
- HTML stripping enabled for `/approve` and `/deny` slash command support
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to target host (direct or via amdpve for CTs)
|
||||
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
|
||||
- Python 3 with `httpx` installed on target
|
||||
- Zulip server accessible at `https://chat.sysloggh.net`
|
||||
|
||||
## Live-State Fields
|
||||
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Mumuni | CT114 | — | 192.168.68.123 | /root/.hermes | root |
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
|
||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||
|
||||
| Field | Value | Trust |
|
||||
|-------|-------|-------|
|
||||
| Zulip server | https://chat.sysloggh.net | ✅ Verified |
|
||||
| Git repo (zulip-platform) | `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git` | ✅ Verified |
|
||||
| Bundled adapter path | `<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/` | ✅ Verified |
|
||||
| Git branch | `feat/zulip-streaming` | ✅ Verified (contains _strip_html fix) |
|
||||
|
||||
## Execution
|
||||
|
||||
### Step 1: Locate Target
|
||||
|
||||
Map `target` to connectivity parameters from the live-state table above.
|
||||
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
|
||||
|
||||
### Step 2: Deploy Zulip Adapter
|
||||
|
||||
On the target host:
|
||||
|
||||
```bash
|
||||
# Clone or update the plugin repo
|
||||
mkdir -p /tmp/zulip-deploy
|
||||
cd /tmp/zulip-deploy
|
||||
if [ -d zulip-platform-plugins ]; then
|
||||
cd zulip-platform-plugins && git pull origin feat/zulip-streaming
|
||||
else
|
||||
git clone --branch feat/zulip-streaming \
|
||||
https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git
|
||||
fi
|
||||
|
||||
# Ensure bundled plugin directory exists
|
||||
mkdir -p <HERMES_HOME>/hermes-agent/plugins/platforms/zulip
|
||||
|
||||
# Copy adapter files
|
||||
cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
|
||||
zulip-platform-plugins/plugins/platforms/zulip/__init__.py \
|
||||
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
|
||||
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
|
||||
|
||||
# Fix ownership (Tanko only)
|
||||
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only
|
||||
|
||||
# Clean up
|
||||
rm -rf /tmp/zulip-deploy
|
||||
```
|
||||
|
||||
### Step 3: Verify _strip_html is Present
|
||||
|
||||
```bash
|
||||
grep -q "_strip_html" <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/adapter.py
|
||||
```
|
||||
Expected: exit code 0. If not found → adapter is stale, re-run Step 2 with fresh clone.
|
||||
|
||||
### Step 4: Verify Env Credentials
|
||||
|
||||
```bash
|
||||
grep -E "ZULIP_SITE|ZULIP_EMAIL|ZULIP_API_KEY" <HERMES_HOME>/.env
|
||||
```
|
||||
|
||||
Expected: all three variables set with non-empty values. If any missing:
|
||||
- ZULIP_SITE: `https://chat.sysloggh.net`
|
||||
- ZULIP_EMAIL: `<agent>-bot@chat.sysloggh.net`
|
||||
- ZULIP_API_KEY: obtain from Zulip admin panel (Bots → show API key)
|
||||
|
||||
### Step 5: Restart Gateway
|
||||
|
||||
```bash
|
||||
cd <HERMES_HOME>/hermes-agent
|
||||
# Use venv if available
|
||||
python3 -m hermes_cli.main gateway restart # or: venv/bin/python -m hermes_cli.main gateway restart
|
||||
```
|
||||
|
||||
Wait for the restart to complete (up to 45s). Check:
|
||||
|
||||
```bash
|
||||
grep "Gateway running" <HERMES_HOME>/logs/gateway.log | tail -1
|
||||
```
|
||||
|
||||
Expected: "Gateway running with N platform(s)" where N > 1 (includes zulip).
|
||||
|
||||
### Step 6: Validate Zulip Connection
|
||||
|
||||
```bash
|
||||
python3 -c "
|
||||
import json
|
||||
d = json.load(open('$HERMES_HOME/gateway_state.json'))
|
||||
print('zulip:', d.get('platforms', {}).get('zulip', {}).get('state', 'NOT FOUND'))
|
||||
"
|
||||
```
|
||||
|
||||
Expected: `zulip: connected`. If not connected, check gateway log:
|
||||
|
||||
```bash
|
||||
grep -E "zulip|Zulip|ZULIP" <HERMES_HOME>/logs/gateway.log | tail -10
|
||||
```
|
||||
|
||||
### Step 7: Report
|
||||
|
||||
Compile results: `{ adapter_deployed, zulip_connected, env_valid, gateway_running }`.
|
||||
|
||||
| State | Action |
|
||||
|-------|--------|
|
||||
| All true | ✅ Agent restored — relay success to user |
|
||||
| `adapter_deployed: false` | Re-run Step 2 |
|
||||
| `env_valid: false` | Prompt for missing credentials |
|
||||
| `zulip_connected: false` | Check Zulip server reachability, verify API key |
|
||||
| `gateway_running: false` | Check process logs for crash cause |
|
||||
|
||||
## Known Failure Modes
|
||||
|
||||
| Symptom | Root Cause | Recovery |
|
||||
|---------|-----------|----------|
|
||||
| Gateway running with 1 platform(s) | Adapter at wrong path (user plugins vs bundled) | Deploy to `<hermes-agent>/plugins/platforms/zulip/` not `~/.hermes/plugins/` |
|
||||
| Queue expired / BAD_EVENT_QUEUE_ID | Idle for 10+ minutes → normal | Auto-reconnects — no action needed |
|
||||
| No events received for N seconds | No DMs or @mentions sent to this bot | Normal if nobody messaged the agent |
|
||||
| `httpx` not found | Missing dependency | `pip install httpx` in the Hermes venv or system Python |
|
||||
| Slash commands not matching | Missing `_strip_html` — Zulip sends `<p>/approve</p>` | Verify `_strip_html` in adapter (Step 3) |
|
||||
| Permission denied on gateway restart | Running as wrong user | Use `su - jerome` for Tanko; root for others |
|
||||
|
||||
## Git Branch Reference
|
||||
|
||||
The `_strip_html` fix lives on `feat/zulip-streaming` branch:
|
||||
```
|
||||
https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins/src/branch/feat/zulip-streaming
|
||||
```
|
||||
|
||||
Commit `55ca15d` — `fix(zulip): add _strip_html for slash command matching`
|
||||
Pull request #33 is the primary integration branch.
|
||||
|
||||
---
|
||||
|
||||
**Last verified good state**: 2026-07-08 — Mumuni, Tanko, Koby all connected with `_strip_html` applied.
|
||||
@@ -144,7 +144,7 @@ description: >
|
||||
|
||||
### Ecosystem A: docker-vm (192.168.68.7)
|
||||
|
||||
11 containers across 4 compose stacks:
|
||||
16 containers across 4 compose stacks + trove agents:
|
||||
|
||||
| Stack | Path | Containers |
|
||||
|-------|------|-----------|
|
||||
@@ -152,6 +152,9 @@ description: >
|
||||
| **SearXNG** | `/opt/search-stack/searxng/` | searxng, valkey |
|
||||
| **Home stack** | `/opt/home_stack/` | jdownloader, stirling-pdf, pulse |
|
||||
| **Audiobookshelf** | `/opt/audiobookshelf/` | audiobookshelf |
|
||||
| **Trove agents** | docker run (standalone) | trove-agent-proxmox, trove-test-agent-1, trove-test-server-1, docker-stats |
|
||||
|
||||
**Last verified:** 2026-07-09 — 16/16 containers running healthy.
|
||||
|
||||
### Home Stack Services
|
||||
|
||||
|
||||
@@ -7,10 +7,14 @@ description: >
|
||||
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
|
||||
via existing /metrics Prometheus endpoint.
|
||||
|
||||
STATUS: Core stack (Prometheus + Grafana + pve/node/docker exporters)
|
||||
already deployed by proxmox-monitor contract. GPU exporters (.8/.110/.15)
|
||||
NOT yet live — gpu-exporter crash-loops on .15, sidecars never deployed.
|
||||
This contract defines the target state; proxmox-monitor is the as-built.
|
||||
DEPLOYMENT STATUS (2026-07-09):
|
||||
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
|
||||
(via proxmox-monitor contract). Grafana at :3001, 5 scrape targets active.
|
||||
❌ GPU exporters NOT deployed: gpu-exporter crash-loops on .15,
|
||||
NVIDIA sidecar exporters (.8/.110:9400) never installed.
|
||||
Router falls back to direct GPU /health probes.
|
||||
⚠️ This contract is target-state aspirational — not as-built.
|
||||
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
||||
version: 1.0.0
|
||||
---
|
||||
|
||||
|
||||
@@ -64,7 +64,7 @@ Before ANY update wave:
|
||||
**Verify after Wave 2:**
|
||||
- All VMs/CTs running: check via Proxmox API
|
||||
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
|
||||
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router
|
||||
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
|
||||
- Zulip agents connected: check Mumuni/Tanko gateway state
|
||||
- Abiba PM2 processes online: `pm2 status`
|
||||
|
||||
@@ -120,6 +120,8 @@ Before Wave 1, snapshot these files:
|
||||
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
||||
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
||||
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
||||
/etc/systemd/system/ornith-server.service (amdpve .15)
|
||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||
```
|
||||
|
||||
Run: `mkdir -p /tmp/infra-update-backup-$(date +%Y%m%d) && rsync -av ...`
|
||||
|
||||
@@ -30,6 +30,12 @@ declare -A CT_NODES=(
|
||||
# acerpve (192.168.68.9)
|
||||
[102]=acerpve # adguard
|
||||
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
|
||||
#
|
||||
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
|
||||
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090)
|
||||
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070)
|
||||
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH)
|
||||
# 118 jitsi → stopped, not in service
|
||||
)
|
||||
|
||||
# Each node must be root-accessible via SSH hostname
|
||||
|
||||
@@ -73,6 +73,25 @@ description: >
|
||||
- **Fix**: If edit fails, send the response as a new message instead
|
||||
- **Applies to**: pi extension (fixed), Hermes adapter (verify edit fallback exists)
|
||||
|
||||
### 11. Model-ID Mismatch Causes Silent Worker Failure (pi-specific)
|
||||
- **Symptom**: User sends DM, sees placeholder, but never gets response. Worker stays
|
||||
"busy" forever, accumulating pending replies (10+). Health endpoint shows green but
|
||||
`messages_processed` stalls.
|
||||
- **Root cause**: Agent's model config (`models.json` or `settings.json`) references a
|
||||
model ID that doesn't exist in LiteLLM's authorized model list. Example: `qwen3.6-35B-A3B`
|
||||
configured but LiteLLM only exposes `ornith-1.0-35b` under that key. pi's session
|
||||
workers emit 403 on first prompt, then never recover because the error doesn't trigger
|
||||
`agent_end` — worker stays `busy` and all subsequent messages pile up in the steer queue.
|
||||
- **Detection**: Compare `~/.pi/agent/models.json` model IDs against `curl -H "Authorization: Bearer <KEY>" http://192.168.68.116/v1/models` output. A stuck worker shows
|
||||
`workers=[<id>:busy:N]` with growing N in extension logs.
|
||||
- **Fix**: (1) Update `models.json` to only include models from the authorized list.
|
||||
(2) Set `defaultModel` to `syslog-auto` (safe routing model). (3) Delete stale session
|
||||
JSONL files from `~/.pi/agent/sessions/zulip/`. (4) Restart PM2 process.
|
||||
- **Prevention**: Use `syslog-auto` as default model for all agents — it handles model
|
||||
routing and fallback automatically. Direct model IDs (`ornith-1.0-35b`, etc.) should
|
||||
only be used when explicitly requested. Validate model IDs at agent setup time.
|
||||
- **Applies to**: pi extension (Tdunna CT111, fixed 2026-07-08), any agent using `syslog-harness` provider
|
||||
|
||||
## Deployment Checklist
|
||||
|
||||
When deploying a new Zulip adapter, verify:
|
||||
@@ -86,3 +105,4 @@ When deploying a new Zulip adapter, verify:
|
||||
- [ ] `@all-bots` detected via configurable user_id
|
||||
- [ ] Poll uses long-poll (not `dont_block=true` polling)
|
||||
- [ ] Stuck/idle detection accounts for quiet periods
|
||||
- [ ] Model IDs in config validated against `GET /v1/models` with actual API key
|
||||
|
||||
@@ -0,0 +1,92 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-oidc-redirect-fix
|
||||
status: active
|
||||
description: >
|
||||
Fixes Zulip OIDC authentication when the redirect_uri sent to Authentik
|
||||
uses the internal IP (192.168.68.19) instead of the public domain
|
||||
(chat.sysloggh.net). Applied via monkey-patch in ZULIP_CUSTOM_SETTINGS.
|
||||
Survives container restarts through compose.override.yaml.
|
||||
agent: abiba
|
||||
triggers:
|
||||
- Zulip OIDC login returns "Redirect URI Error" from Authentik
|
||||
- redirect_uri in OAuth URL shows 192.168.68.19 instead of chat.sysloggh.net
|
||||
- After Zulip server restart, Authentik SSO login broken
|
||||
---
|
||||
|
||||
## Maintains
|
||||
|
||||
- zulip-oidc: { redirect_uri: "https://chat.sysloggh.net/complete/oidc/", scheme: "https", host: "chat.sysloggh.net" }
|
||||
- authentik-acceptance: { status: "accepted" | "rejected" }
|
||||
- patched-strategy: { module: "social_core.strategy.BaseStrategy", method: "absolute_uri", root: "ROOT_DOMAIN_URI" }
|
||||
|
||||
## Detection
|
||||
|
||||
### Rule 1: Wrong redirect_uri host
|
||||
- **Detect**: `curl -sk -L "https://chat.sysloggh.net/accounts/login/social/oidc/authentik" 2>&1 | grep "redirect_uri=https://192.168.68.19"`
|
||||
- **Status**: CRITICAL — Authentik will reject
|
||||
- **Trigger fix** → Execute self-heal
|
||||
|
||||
### Rule 2: OIDC flow broken
|
||||
- **Detect**: `curl -sk -o /dev/null -w "%{http_code}" "https://chat.sysloggh.net/accounts/login/social/oidc/authentik"` → chain ends at Authentik 400
|
||||
- **Diagnose**: Check redirect_uri in the 302 Location header chain
|
||||
|
||||
## Fix (Self-Heal)
|
||||
|
||||
Two layers applied:
|
||||
|
||||
### Layer 1: Live patch (inside container, immediate)
|
||||
```bash
|
||||
# Add to /home/zulip/deployments/current/zproject/computed_settings.py:
|
||||
SOCIAL_AUTH_REDIRECT_IS_HTTPS = True
|
||||
|
||||
import urllib.parse
|
||||
from social_core.strategy import BaseStrategy
|
||||
_original_absolute_uri = BaseStrategy.absolute_uri
|
||||
def _patched_absolute_uri(self, path=None):
|
||||
from django.conf import settings
|
||||
root = getattr(settings, "ROOT_DOMAIN_URI", "https://chat.sysloggh.net")
|
||||
if path is not None:
|
||||
return urllib.parse.urljoin(root, path)
|
||||
return root
|
||||
BaseStrategy.absolute_uri = _patched_absolute_uri
|
||||
|
||||
# Restart Django
|
||||
supervisorctl restart zulip-django
|
||||
```
|
||||
|
||||
### Layer 2: Persistent fix (compose.override.yaml)
|
||||
The patch is baked into the `ZULIP_CUSTOM_SETTINGS` env var in
|
||||
`/opt/zulip/compose.override.yaml`. Survives Docker container restarts.
|
||||
|
||||
### Verification
|
||||
```bash
|
||||
STEP1=$(curl -sk -w "%{redirect_url}" \
|
||||
"https://chat.sysloggh.net/accounts/login/social/oidc/authentik" -o /dev/null)
|
||||
curl -sk -D- "$STEP1" -o /dev/null 2>&1 | grep "redirect_uri="
|
||||
# Expected: redirect_uri=https://chat.sysloggh.net/complete/oidc/
|
||||
# Wrong: redirect_uri=https://192.168.68.19/complete/oidc/
|
||||
```
|
||||
|
||||
### Rollback
|
||||
Remove the patch block from `compose.override.yaml` and restart the container:
|
||||
```bash
|
||||
docker compose -f /opt/zulip/compose.yaml -f /opt/zulip/compose.override.yaml up -d zulip
|
||||
```
|
||||
|
||||
## Root Cause
|
||||
|
||||
After Zulip restart, `social-auth-core` computes the OIDC `redirect_uri` via
|
||||
Django's `request.build_absolute_uri()` → `request.get_host()`. The upstream
|
||||
Netbird/Traefik proxy (72.61.0.17) forwards `Host: 192.168.68.19` instead of
|
||||
`Host: chat.sysloggh.net`, and without `HTTP_HOST` in nginx's `uwsgi_params`,
|
||||
Django falls back to the server's IP.
|
||||
|
||||
The monkey-patch overrides `BaseStrategy.absolute_uri()` to always use
|
||||
`ROOT_DOMAIN_URI` (`https://chat.sysloggh.net`) regardless of the request's
|
||||
Host header.
|
||||
|
||||
## Related Contracts
|
||||
|
||||
- `zulip-health.prose.md` — General Zulip health monitoring
|
||||
- `zulip-self-heal.prose.md` — RETIRED (pi extension removed)
|
||||
Reference in New Issue
Block a user